1. What happened
Authorized model builders bring models into Harell Data's CoreWeave-powered environment for training, fine-tuning, validation, and inference. Raw datasets remain in place and cannot be viewed or extracted. Builders receive trained models and model IP; data owners receive revenue share from each metered training run, while downstream models may be listed as Models as a Service on Harell's marketplace.
2. Buyer / seller
Authorized Harell Data model builders and scientific research teams is the AI buyer or model-development side. A-Alpha Bio, Adaptive Biotechnologies owns or contributes the proprietary corpus.
3. Economics
Harell states that data owners receive a revenue share on every training job executed against their dataset; the percentage, minimum commitment, and cash amount are not disclosed. Model builders can also collect a fee when users run marketplace-listed models. Economics scope: Per-training-run data-owner revenue share and usage fees for downstream Models as a Service; CoreWeave contract value and exact revenue-share percentages are not disclosed.. Strategic Signal has not estimated undisclosed consideration.
4. Proprietary data
Proprietary scientific datasets hosted in Harell Data's controlled environment, initially including A-Alpha Bio antibody-antigen affinity and structural data and Adaptive Biotechnologies T-cell-receptor binding data, plus hidden validation questions retained by data partners. Historical depth: The data partners generated substantial proprietary experimental datasets, but the reviewed sources do not quantify collection start dates or longitudinal depth.. Refreshability: Harell describes a growing partner and dataset ecosystem and per-run marketplace model; the cadence and exact update process for each initial dataset are not disclosed..
5. Rights and restrictions
Training, fine-tuning, validation, and inference execute inside Harell's controlled platform; raw datasets cannot be viewed or extracted, and compute is attributable and metered per run. Ownership: No dataset ownership transfer is disclosed. Raw datasets remain with data providers and do not leave the Harell environment.. Training: Conditional and permitted inside Harell's controlled environment for authorized builders; raw data may not be viewed or extracted.. Post-training: Fine-tuning is expressly permitted inside the controlled platform; other post-training rights are not disclosed..
6. Why the data matters for AI
Training, fine-tuning, validating, and serving protein-language, structure-prediction, antibody-engineering, and immune-receptor models without transferring raw proprietary records. The initial datasets link biological sequences or structures to measured antibody-antigen affinity and T-cell-receptor binding outcomes, supporting model learning and blinded validation against experimental results.
7. Asset Score — 84/100
Strategic Signal analysis. The named scientific corpora combine scarce proprietary experimental measurements with high-value biological outcomes and direct model-training utility; exact historical depth and corpus scale remain partly undisclosed. Score coverage: 92%.
| Factor | Weight | Rating / 5 | Points |
|---|---|---|---|
| Uniqueness / scarcity | 12 | 4 | 9.6 |
| Historical depth | 8 | Unknown | — |
| Decision → outcome richness | 13 | 4 | 10.4 |
| Decision Value Density | 13 | 4 | 10.4 |
| Real-world grounding | 10 | 5 | 10.0 |
| Domain economic value | 9 | 5 | 9.0 |
| Refreshability | 8 | 3 | 4.8 |
| Proprietary advantage | 9 | 4.5 | 8.1 |
| Model-learning usefulness | 10 | 4.5 | 9.0 |
| Non-replicability | 4 | 4 | 3.2 |
| Rights usability | 4 | 4 | 3.2 |
8. Transaction Signal Score — 71/100
Strategic Signal analysis. The arrangement cleanly separates controlled data use from raw-data transfer and ties data-owner compensation to training runs, but provides no dollar value, share percentage, bid process, or independent uplift result. Score coverage: 85%.
| Factor | Weight | Rating / 5 | Points |
|---|---|---|---|
| Disclosed economics | 16 | 2.5 | 8.0 |
| Clean price discovery | 16 | 1.5 | 4.8 |
| Separability of data consideration | 15 | 4.5 | 13.5 |
| Explicit training / model-improvement use | 14 | 5 | 14.0 |
| Clarity of rights purchased | 12 | 4.5 | 10.8 |
| Competing bids | 10 | Unknown | — |
| Strategic buyer quality | 6 | 3 | 3.6 |
| Repeatability / precedent value | 6 | 5 | 6.0 |
| Independent model-uplift evidence | 5 | Unknown | — |
9. M&A / valuation implications
Establishes a repeatable data-market structure in which scientific asset owners monetize controlled learning runs without selling or delivering the corpus, potentially increasing strategic value for experimental-data generators and secure model-training platforms.
10. Evidence boundary
The named data partners, initial dataset categories, controlled in-platform training and inference, raw-data non-extraction, builder model ownership, multi-year infrastructure agreement, and per-run data-owner revenue share are disclosed. Exact corpus sizes, dataset-specific update cadence, access prices, revenue-share percentages, exclusivity, geography, retention periods, patient-level content, downstream memorization controls, and measured model uplift are not disclosed. Chronology correction: the founder post is dated September 15, 2026 and already describes the same named datasets and learning-rights architecture; the September 23 release concerns supporting infrastructure. The first contract execution date, original provider agreement dates and historical page publication/version timestamps are not independently established. This record is excluded from the September 17, 2026 forecast cohort under the conservative first-disclosure treatment; the correction does not create or remove a registry event.