Merck’s $510M Protillion deal shows AI drug discovery IP ownership splits into three assets: the dataset, the model, and the candidates it surfaces.

On June 16, 2026, Merck committed up to $510M to Protillion Biosciences: not for a drug compound, not for a clinical program, but for access to a data platform called Prot-MaP.
Prot-MaP tests up to one million protein variants simultaneously to build the training datasets that protein-design AI models run on, and that function means the Merck collaboration created three distinct IP assets at signing: the generated dataset, any model trained on it, and the drug candidates that model surfaces.
Standard pharma deal templates don’t address who owns any of the three. That gap is not unique to this deal: 114 AI drug discovery partnerships closed in 2025 alone, each carrying the same three-layer exposure.
A clause assigning “all model outputs” to the pharma partner leaves the underlying dataset exposed; a reported dispute between BioNTech and Nucleai shows a data provider can assert trade secret rights in the dataset itself regardless.
Fine-tuned model weights carry their own exposure: their ownership must be resolved before training starts, not after a candidate reaches patent filing. U.S. patent law goes further: a co-owner of a drug patent can independently license it to a competitor without telling the other party.
All three must be contractualized before signing, or default rules gift the core asset to the partner.

Want to map training data, model weights, and drug candidate ownership in your next AI collaboration before default rules decide who gets what? Talk to ipCapital Group about allocating IP across all three layers before you sign.
Work with ipCapital Group
From invention to monetization, our team has guided 2,000+ engagements across the full IP lifecycle. Start with a free 30-minute discovery call.
Written by
John Cronin