A ranked database of real human promoters — DNA that already exists in the genome. A promoter is the stretch of DNA in front of a gene that decides how strongly that gene is switched on. Every protein-coding promoter in the human genome was scored twice, once with the target cell type described to the model and once with the cell type we want it to stay quiet in, and ranked by the difference between the two — the margin. It answers one question: which promoter that already exists is the most selective one for my pair of cell types?
140 records with full provenance, across 5 on/off cell pairs — the top 25 per pair plus every reference promoter, whether or not it passes. Each carries the window it was measured in, both context strings, its noise floor and its verdict. Behind them sit 4,122 candidates that clear both gates, and the windows the sequence filters removed — both published in full, with the records below.
There is a second database, and it is strong exactly where this one is weak. The promoters here already exist, so other people have measured them in real cells, and that measured evidence backs this ranking up. But the model was trained on essentially every gene’s promoter paired with its measured expression, so a promoter that ranks well here may be partly recalled rather than predicted — how big that effect is, and what survives it, is measured below. The designed promoters have the opposite shape: they exist in no genome, so no rank there can be recall — and for exactly the same reason, no laboratory has ever measured one.
The two are separate releases with separate file names and separate limitation blocks. Do not merge them into one table or one ranking -- they carry different caveats. The designs are 600 bp modules scored in one named scaffold and their margins are only comparable within that scaffold; the natural candidates are 9,198 bp genomic windows. A ranking that mixed them would compare two different measurements. Exactly one quantity is allowed to cross between them — a design’s margin against the best natural promoter in the same window — and it has its own comparison view, labelled as the only one.
The designed release carries no section like this one, and cannot. This question can only be asked of DNA that already exists in a genome and that other people have measured in a laboratory. Nothing in the designed release exists in any genome, so there is no field standard there to recover — and, for exactly the same reason, no rank there can be recall of training data.
Genome-wide ranking: 20,107 protein-coding TSS windows x 11 context strings = 221,177 predictions, ranked over the 19,987 windows that pass the sequence filters. No gene was prompted for, weighted, or pre-selected. The question asked before the run was whether the promoters the field already uses come back near the top. NOTE what 'blind' means: the ranking was blind to the answer, not to the genes. The model was trained on the expression of essentially every gene ranked here, including all five promoters in the headline below. Read training_split_caveat with this.
hSYN1, one of the three promoters the go/no-go named explicitly, is the softest arm: SYN1 ranks 1,669 in A4 and 626 in A5. Both margins are reportable and the direction is right, so it is recovered -- but it is not near the top, and it fails the on-target expression floor outright. STMN2 and ENO2 outrank it and we cannot show the model is wrong about that. Two further field standards are missed in the all-vs-all view and recover only in their own pair: SERPINA7 (TBG), the standard liver promoter, ranks 1,994; CAMK2A ranks 11,866 all-vs-all and 331 in A4.
How big it is. 118 of the 125 provenanced top-25 rows (25 per pair) are genes in the model's training split -- 94.4%, against a held-out base rate of 21.0%. (4,053 of the 19,987 rankable windows are held out, test plus validation; 21.0% is that share among the windows carrying a split label.) And ten of the eleven promoters highlighted in the recall evidence above are training genes.
What survives it. Re-ranking ONLY the 4,053 held-out genes, with the identical gates and no score changed, four of the five field-standard promoters that are held out are recovered under both gates: ENO2 (NSE) at held-out rank 17 of 4,053 in A4 and 11 in A5, MYL1 at 3 in A2, DES (desmin) at 60 in A2, SNAP25 at 101 in A4. The fifth is SYN1, which fails its on-target floor in the all-genes ranking too and is neither rescued nor worsened by excluding memorisation. MYBPC3 -- a held-out cardiac gene -- is the top held-out hit in both cardiac pairs. These cannot be memorisation: the model never saw their expression.
| pair | held-out generalisation | |
|---|---|---|
| A2 | strong | held-out genes clear both gates at 1.17x the training rate |
| A4 | strong | 1.09x the training rate; all ten of its held-out top 10 are recognised neuronal genes |
| A1 | clean | the raw 0.69x looks like a deficit but is pool composition: the held-out pool clears the cardiomyocyte on-target floor at 0.62x the training rate. Adjusted, A1 observes 95 held-out survivors against 85.1 expected (p = 0.87) -- no deficit |
| A3 | not demonstrated | exactly ONE of 4,053 held-out genes clears both gates, a ~6x deficit after the same pool adjustment (1 observed vs 6.1 expected, p = 0.016). Separately, none of A3's four reference promoters is held out, so no field standard can test A3 on this split at all. A3 IS NOT VALIDATED -- treat its shortlist as a hypothesis list only |
| A5 | not demonstrated | ten of 4,053 held-out genes clear both gates, a ~2x deficit after adjustment (10 observed vs 19.8 expected, p = 0.012) |
How to quote this. Quote a rank together with its training status. 'The screen rediscovered cTnT at rank 4' is not usable without 'and cTnT was in the training data'. 'Four of five held-out field standards are recovered under both gates' is usable, with its n -- four promoters, three pairs -- and never for A3, which has no held-out field standard and one held-out survivor. Prefer held-out genes when re-running this check.
What is still confounded. The split is locus-blocked -- assigned over contiguous genomic intervals -- so held-out genes are concentrated on some chromosomes and nearly absent from others, and the tissue programmes encoded in tandem arrays travel with them. The pool-availability adjustment above removes the measurable part of that; composition within the available set is not removable on this split.
This method recovers the promoters the field already uses, in five of the six A-tier pairs it was checked on -- A1 through A5 pass; the sixth, A6, failed and is stated as an exclusion under Withdrawn below rather than dropped -- from a genome-wide ranking. It is not evidence that any individual novel candidate works. It is also not, on its own, evidence that the model generalises rather than remembers: the promoters in the table above were in its training data. The claim that survives that objection is the held-out one above -- four of the five field-standard promoters that are held out are recovered under both gates, which holds for A1, A2 and A4 and not for A3 or A5.
Every figure in this section, per pair and per promoter, is in recall_check.json (JSON download).
| pair | on-target context | off-target context | clears both gatesgenome-wide, not just the rows below | provisionalinside the noise band of its own floor | on-target floor | median genewhat a RANDOM gene achieves in this pair |
|---|---|---|---|---|---|---|
| A1 | cardiomyocyte_ventricular | hepatocyte_primary | 618 | 190 | 4.5 | +0.7446 |
| A2 | skeletal_myofiber | hepatocyte_primary | 1,920 | 1,501 | 1.8203 | +0.2471 |
| A3 | cardiomyocyte_ventricular | skeletal_myofiber | 38 | 31 | 4.5 | +0.457 |
| A4 | cns_neuron | hepatocyte_primary | 1,450 | 554 | 3.5 | +0.5 |
| A5 | cns_neuron | astrocyte | 96 | 142 | 3.5 | +0.1719 |
The last column is the score an ordinary gene already gets — the level a result has to beat before it means anything. It is the median gene of the whole genome, measured in that pair, and it is not the same number twice: contexts sit at different genome-wide baselines, so raw margins are not comparable across pairs. In A1 the median gene already scores +0.7446 — a raw margin of 0.8 there is what a random gene achieves. Every record below carries that number as the leftmost mark on its own axis, and percentile_margin is the cross-pair quantity.
A6 — Retinal pigment epithelium ON / Hepatocyte OFF. Failed its own go/no-go. RPE65's own promoter does not clear the reportable margin threshold against liver -- and RPE65 is the gene whose loss causes the disease treated by the only approved RPE-directed gene therapy, so if any RPE promoter should have been easy to recover, it was this one. (That therapy delivers RPE65 as the transgene under a general-purpose viral promoter; it does not use the RPE65 promoter. The point here is the gene's standing in this cell type, not its construct.) One of five RPE markers is reportable. A coarse cross-organ contrast that should have been easy.
B2 — Photoreceptor ON / Retinal pigment epithelium OFF. 0 of 9 RPE marker genes separate RPE from photoreceptor. The two retinal contexts are not resolvable from each other on this checkpoint, so no ranked list in either direction is defensible.
These pairs did not fail. They are ranked internally and do not yet carry the per-pair evidence the statements above assume, so every count and verdict here excludes them.
A7 — Astrocyte ON / Hepatocyte (liver parenchymal cell) OFF. Ranked 2026-08-14 and held out of this release until it carries what every shipped pair carries: a row in the recall check, a per-pair verdict in the held-out re-analysis (which covers A1-A5 only), and a stated rationale. Its ranking exists internally; nothing about it is published here, and no number in this release includes it.
The top 25 per pair plus every reference promoter, whether or not it passes, each with full provenance. The full database — every candidate clearing both gates in every pair — and the windows the sequence filters threw out, with their scores, are published as tab-separated files. Downloads: candidates.tsv — the records below, flat; shortlists.tsv — every candidate clearing both gates in every pair; excluded.tsv — the windows the sequence filters removed, with their scores.
assembly GRCh38 annotation GENCODE release 50 (Ensembl 116) MANE NCBI MANE v1.4 window 9198 bp, TSS at 0-based offset 4599, gene-sense model g0-expression rev v1, checkpoint 20260523 unit ln(quantile-normalised TPM + 1) endpoint POST https://api.genomicintelligence.ai/v1/tasks/expression/predict
The 5 context strings behind the pairs in this release — the ones whose noise floor and marker controls we measured — are in contexts.json, with complete SHA-256 digests. They are reproduced verbatim and must not be corrected — why, on the overview, along with how the window above is built and what silently goes wrong when it is not.
Every row in candidates.jsonl carries a reproduce block with the exact coordinates, request body and expected values for that row. The release as a whole — every field above, the file inventory and the limitations block — is in MANIFEST.json (JSON download).