rDNA.ai
Colored cups arranged in a chain-link fence beneath a blue sky with white clouds.
8 min readBy Mark Edwards

Whose Genus Is It?

The seventeenth article in the rDNA.ai biopharma BD&L series reads the menin–MLL field through chemical genus membership, separating competitive boundaries from same-sponsor comparisons by patent family and ownership, without treating structural verdicts as FTO clearance.

On November 13 of last year the FDA approved ziftomenib for relapsed or refractory acute myeloid leukemia, 363 days after it approved revumenib for a neighboring indication. Two menin–MLL inhibitors, two days short of a year apart, on a protein–protein interaction that was called undruggable for most of my career. Behind them sits a dense field of small-molecule filings from a dozen or more sponsors, each claiming a swath of chemical space around the same binding pocket.

For a BD&L professional the question that follows an approval pair like that is not whether the mechanism works. It is where the room is. When the next menin asset comes up for license — and in a field this crowded, one will — the first structural question is deceptively simple to state and famously expensive to answer: whose claimed space is this compound actually in?

The previous article in this series took that question to an antibody target, CD38, where scope is written in sequences. Menin is the small-molecule twin, where scope is written in Markush genera — the generic chemical formulae, with their branching lists of permitted substituents, that a composition-of-matter claim uses to fence a family of related structures. Different chemistry, different instrument, and, as it turned out, the same trap.

Reading a compound against a genus has historically been slow, lawyer-only work, which produces a particular and expensive habit: if you cannot afford to resolve every boundary, you assume every boundary is contested. On a target with a dozen overlapping estates, that assumption is both costly and wrong, because the boundaries are not uniformly contested. Some resolve cleanly, and the ones that do not are a minority you can name.

The way to see this is to take one foundational genus and read it against the rest of the field. Our anchor is a Janssen filing, US-11220517-B2, a spiro-bicyclic thienopyrimidine claim that compiles cleanly: precision 0.995 across 507 enumerated member structures, with a recall lower bound of 0.706. Precision of that order is what makes an OUT worth leaning on — when this genus says a structure sits outside the claim, it is rarely wrong about that. Recall stated as a floor is the honest counterpart: the genus captures at least that share of the space it should, and possibly more, but never claims to have found all of it.

The reading returns one of three answers, and the discipline is in keeping all three honest. IN means the structure plausibly falls within the anchor’s claim. OUT means it resolves outside. Neither is a clearance and neither ends a diligence process, but both let a professional narrow a question quickly, before anyone bills for an opinion. The third answer is the one worth dwelling on. INDETERMINATE does not mean we do not know, in the shrug-of-the-shoulders sense. It means membership could not be resolved from the chemistry alone — which, on a contested composition claim, is not a failure of the instrument but a localization of the problem. It marks the stretch of boundary that genuinely has to answer to counsel. The most valuable output is not the cell the tool resolves but the cell it flags.

That much survived. What did not survive was the tally.

Read across the menin field, the anchor produced a headline that was easy to write and easy to sell: 19 field genera, 36 directed comparisons, and a split of one IN, six OUT, and 29 INDETERMINATE. Stated for a deal conversation, that becomes a sentence about a foundational genus resolving cleanly outside a handful of rivals, inside one, and landing in the contested band against the large majority. The one is what a reader’s eye goes to. A rival’s chemistry found inside a Janssen fence is the kind of finding that moves a negotiation.

It was not a rival. Applying the same ownership discipline the CD38 read had forced on us — resolve every publication to its assignee and, critically, to its patent family before counting anything — the field changed shape. The 19 publications are 11 distinct INPADOC families. Four of the 11 are Janssen, the anchor’s own sponsor. Of the 22 materialized directed edges, 12 are Janssen against Janssen. And the single IN is one of those 12. Read only against genuine third parties, the anchor reads IN against nothing at all. Cross-sponsor IN is zero.

Family-level attribution is what does the work, and one member shows why publication-level attribution fails. CN-110248946-B carries no assignee of its own, so a publication-level read leaves it unattributed and, by default, counts it as a rival. Its family resolves to Janssen. Nine of the 19 had been drawn as unattributed and one as China-origin; all 19 are now named. The deduplication matters just as much, because 19 publications counted as 19 genera inflate the denominator: several families were being counted more than once, which is how 36 directed cells became 22.

Partitioned properly, the corrected picture is two tallies rather than one, because the two answer different questions and must not share a denominator. Against genuine competitors, ten directed edges: three OUT, zero IN, seven INDETERMINATE. Against the anchor’s own sponsor, 12 directed edges: two OUT, one IN, nine INDETERMINATE. The competitive field is smaller than it looked and more uniformly unresolved, and the clean OUTs — Kyowa Kirin, Medshine and HelioEast, read anchor-as-query — are all cross-sponsor. They were the three most defensible readings in the whole exercise, and the original tally buried them among the anchor’s self-comparisons.

The figure below carries both blocks, competitive boundaries above and same-sponsor below, deliberately kept apart on the page. Fills are categorical — IN, OUT, INDETERMINATE, and an explicit not-established — never a gradient, because a gradient would manufacture a precision this reading does not have.

Menin–MLL genus membership against the Janssen anchor, grouped by patent family and ownership: competitive comparisons above and same-sponsor comparisons below, with IN, OUT, INDETERMINATE, and not-established cells.
The anchor genus read against the rest of the field, 11 distinct INPADOC families in all, in two separately tallied blocks: competitive boundaries above, the anchor’s own sponsor below. Rows carry the resolved owner and family identifier, and rows previously drawn as unattributed or China-origin are marked with the label they replace. One row per publication is kept inside each family group rather than collapsing a family to a single verdict, because members disagree. Not-established means unread against the anchor — neither outside nor open. Candidate generation, not a freedom-to-operate determination.

Three things fall out of the correction, and the first is the one a BD&L professional can act on immediately. The named third parties are the map worth having: Ventyx Biosciences with three separate families, Scripps, Origenis and Kyowa Kirin with one each, and Medshine and HelioEast sharing a single family under two spellings of the assignee. Unattributed is not a sponsor list anyone can act on. Second, the single IN should not be priced as competitive risk, because it is a portfolio-shape fact rather than a boundary a third party must clear — though same-sponsor is not free ground either, since Janssen holds both genera and a newcomer clears neither. Third, direction still matters, and members disagree with each other: family 60935807 alone carries IN, OUT and INDETERMINATE across its five members. Collapsing each family to one verdict would have forced a pick and silently overwritten the conflict. Anchor-as-host answers whether their chemistry falls in your genus; anchor-as-query answers whether yours falls in theirs. Those are two different deal questions with two different answers.

The limits are worth stating plainly, because most of them cut against the headline rather than for it. Not-established cells are unread, not open, and reading them as whitespace would invent room that was never measured. Attribution is assignee of record, which is exactly the thing the CD38 read showed can be stale — licenses, options and unrecorded assignments are invisible to it, and Ventyx’s three families are three filings rather than a verified single program. This is one anchor’s star read, not an all-pairs matrix, and a full matrix is not currently available: across the August corpus only 13 of 176 on-pathway families have a readable Markush fence, with 66 degraded and 97 carrying no genus headline at all. Genus compile quality, not corpus size, is the binding constraint on this target.

None of this is a freedom-to-operate opinion. The reading is candidate generation: it surfaces boundaries worth examining and chemotypes that resolve cleanly, with recall stated as a lower bound and INDETERMINATE stated as a flag rather than a finding. It does not clear a molecule, and OUT is never free to operate. What it does is make the expensive next step cheaper and sharper — pricing the structural risk of a candidate in-license from where it sits in the resolved-versus-contested split, scoping the counsel question to particular contested boundaries rather than to a whole landscape, and telling a team going in whether an asset sits in open structural room or in the crowded band. That last fact is often the most decision-relevant thing about a crowded-target in-license, and no deal announcement will ever contain it. The supporting analysis for this read is posted at rDNA.ai, alongside the earlier articles in this series, including the attribution and partition method that produced the correction.

In real estate a land surveyor and a title lawyer do different work, and the good ones know it. The survey does not settle a boundary dispute; that is the lawyer’s job, and no competent surveyor pretends otherwise. But the survey tells you which fence lines run clean and which cross onto the neighbor’s parcel, so that when you do bring the lawyer, you bring her to the contested stretch and not to the whole perimeter. The previous article made the point that structural analysis will show you where the fence lines run but not whose fence it is. This one is what that costs when you skip it: more than half of what looked like a perimeter dispute turned out to be one owner’s internal fencing, and the single most quotable finding in the file was a company standing on both sides of its own line.