Natural promoters

A ranked database of real human promoters — DNA that already exists in the genome. A promoter is the stretch of DNA in front of a gene that decides how strongly that gene is switched on. Every protein-coding promoter in the human genome was scored twice, once with the target cell type described to the model and once with the cell type we want it to stay quiet in, and ranked by the difference between the two — the margin. It answers one question: which promoter that already exists is the most selective one for my pair of cell types?

140 records with full provenance, across 5 on/off cell pairs — the top 25 per pair plus every reference promoter, whether or not it passes. Each carries the window it was measured in, both context strings, its noise floor and its verdict. Behind them sit 4,122 candidates that clear both gates, and the windows the sequence filters removed — both published in full, with the records below.

There is a second database, and it is strong exactly where this one is weak. The promoters here already exist, so other people have measured them in real cells, and that measured evidence backs this ranking up. But the model was trained on essentially every gene’s promoter paired with its measured expression, so a promoter that ranks well here may be partly recalled rather than predicted — how big that effect is, and what survives it, is measured below. The designed promoters have the opposite shape: they exist in no genome, so no rank there can be recall — and for exactly the same reason, no laboratory has ever measured one.

The two are separate releases with separate file names and separate limitation blocks. Do not merge them into one table or one ranking -- they carry different caveats. The designs are 600 bp modules scored in one named scaffold and their margins are only comparable within that scaffold; the natural candidates are 9,198 bp genomic windows. A ranking that mixed them would compare two different measurements. Exactly one quantity is allowed to cross between them — a design’s margin against the best natural promoter in the same window — and it has its own comparison view, labelled as the only one.

Limitations

RESEARCH USE ONLY. Every number in this release is a prediction from the g0-expression model on checkpoint 20260523. Nothing here has been validated in a wet lab, in any cell type, at any length. The margins are model outputs on a checkpoint with a measured compressed dynamic range, computed on natural genomic TSS windows -- not on AAV cassettes, which are out of distribution for this model. Do not use these sequences or rankings in a clinical, diagnostic or therapeutic decision.
not_wet_lab_validated · criticalNOTHING IN THIS RELEASE IS WET-LAB VALIDATED. Every margin, rank and cassette recommendation is a model prediction. No candidate has been cloned, transfected, transduced or measured. Treat the whole database as a hypothesis generator.
training_set_membership · criticalTHE RANKING IS PARTLY RECALL OF TRAINING DATA. The model was trained on the TSS window of essentially every human gene paired with its measured expression, over a declared train/test/validation split -- and this release ranks those same windows. The provenanced top-25 tables are 118 of 125 training genes (94.4%) against a 21.0% held-out base rate, and ten of the eleven promoters the recall evidence highlights are training genes. A high rank here is therefore NOT on its own evidence that the model generalises rather than remembers. What is: re-ranking only the 4,053 held-out genes, with the identical gates and no score changed, recovers four of the five field-standard promoters that are held out -- ENO2 (held-out rank 17 in A4, 11 in A5), MYL1 (3, A2), DES (60, A2) and SNAP25 (101, A4), all clearing both gates -- and MYBPC3, a held-out cardiac gene, is the top held-out hit in both cardiac pairs. Per pair: held-out genes clear both gates at the training rate in A2 (1.17x) and A4 (1.09x), and in A1 once the held-out pool's depletion of cardiac genes is adjusted for (95 observed vs 85.1 expected, p = 0.87); in A3 and A5 they do not (~6x and ~2x deficits after the same adjustment). A3 IS NOT VALIDATED: one held-out gene of 4,053 clears both gates, and none of A3's four reference promoters is held out, so no field standard can test it on this split. Both statements hold at once: the limitation is real and so is the generalisation result. Generated sequences cannot have been memorised at all; they are packaged separately, see relationship_to_the_generated_design_release.
compressed_dynamic_range · criticalThe 20260523 checkpoint compresses magnitudes. Most real genes occupy 0.03-0.9 ln(TPM+1) and the smallest reportable margin threshold, 0.80, is close to the full width of that band. In that band a selectivity claim is barely resolvable from prompting noise alone. A larger margin is better evidence than a smaller one; neither is a measurement.
baselines_differ_across_pairs · criticalContexts sit at different genome-wide baselines, so RAW MARGINS ARE NOT COMPARABLE ACROSS PAIRS. Median score over the 19,987 rankable windows: cardiomyocyte 2.17, CNS neuron 2.05, astrocyte 1.67, myofiber 1.38, hepatocyte 0.90. A margin against hepatocyte therefore starts with a ~1.3 log-unit head start that has nothing to do with the candidate. In A1 THE MEDIAN GENE ALREADY SCORES +0.745 -- a raw margin of 0.8 in A1 is what a random gene achieves. Use percentile_margin when comparing across pairs, and read every margin against the pair's genome_wide_median_margin.
cassettes_out_of_distribution · criticalAAV CASSETTES ARE OUT OF DISTRIBUTION FOR THIS MODEL. It was trained on genomic TSS windows of real genes in real chromatin context. A [promoter]->[payload] construct inside an AAV backbone, episomal and unchromatinised, is not that. The trimmed fragments in this release are natural genomic fragments scored as bare DNA -- a step toward a cassette, not a cassette. The go/no-go on cassette encoding has not been run.
hsyn1_fails_its_own_floor · highhSYN1, the standard neuron-specific AAV promoter, FAILS THE ON-TARGET EXPRESSION FLOOR in both neuronal pairs. SYN1 scores 2.469 under the cns_neuron context against a floor of 3.50. Its A4 margin (+2.419) is reportable and its direction is right, but the field's standard neuronal promoter is not in our shortlist and the model ranks STMN2 and ENO2 above it. We cannot show the model is wrong. The gate was not moved to accommodate it; SYN1 is carried in both pair files so that this is visible rather than absent.
trimming_destroys_selectivity · highTRIMMING TO A DELIVERABLE CASSETTE LENGTH DESTROYS SELECTIVITY FOR A SUBSTANTIAL FRACTION OF CANDIDATES. Per-pair counts of the shortlisted candidates where no deliverable-length fragment clears its own margin floor (specificity_lost_on_trim): A1 1/25, A2 1/25, A3 4/25, A4 4/25, A5 11/25. Separately, the best fragment retains under half the full-window margin for A1 4/25, A2 8/25, A3 4/25, A4 22/25, A5 12/25 (margin_substantially_reduced_on_trim). Read both flags before quoting a recommended_cassette. In A4 the selectivity localises to the first 1,533 bp DOWNSTREAM of the TSS for 24 of 25 candidates, which is exactly the region a conventional upstream cassette discards.
margin_is_not_evidence · highA margin is not evidence on its own, at any size. ACTB scores +2.031 HepG2-K562 where the measured truth is -0.568 -- wrong sign, and a spurious margin larger than most real hits. The control panel checks a handful of genes; it cannot tell you the model is right about the next one.
The other 8 limitations
single_off_target · mediumOff-target coverage is one cell type per pair. A candidate quiet against hepatocyte may be loud in a tissue not screened here. Macrophage, vascular endothelium, microglia and oligodendrocyte are not in these numbers.
fragment_scores_within_gene_only · mediumFragment scores are within-gene, within-context comparisons only. A 600 bp input is shorter than anything the model was trained on and the length effect is large: the median on-target score drops 1.6 (2,500 bp) to 5.6 (500 bp) log units purely from shortening the input. All fragment gates are therefore re-derived at the fragment's own length.
rank_is_not_biological_quality · mediumRank is a ranking of predicted margin under two specific context strings, not a ranking of biological quality. Change the strings -- even to a defensible paraphrase -- and low scores move by ~60% of their own value. The exact strings used are in every record and in contexts.json.
contexts_are_pseudobulk_profiles · mediumThe context strings are pseudobulk cell-type expression profiles flattened to '<field> is <value>.' clauses, not promoter assays. A prediction under cardiomyocyte_ventricular is the expected profile of that annotated cell population, not a measurement of a promoter.
effective_window_is_asymmetric · mediumDNA is tokenized at ~6.25 bp/token, so a 9,198 bp window is ~1,400-1,500 tokens against ~1,022 usable slots. The model sees all 4,599 bp upstream of the TSS but only roughly 1,600-2,160 bp downstream, and how much survives varies with the sequence's compressibility. A tile scored beyond that boundary shows the model sequence the full-window score never saw.
not_localised_means_unknown · mediumselectivity_locus = 'not_localised' means no single 1,533 bp tile reproduced the contrast, i.e. we could not locate the active element -- not that it is proximal. It is the commonest verdict in A3 (14/25) and A5 (12/25). Do not read a recommended_cassette on such a candidate as 'the active element is in there'.
excluded_windows · lowThe model reads N as real sequence. 120 of 20,107 windows are excluded -- 104 duplicate sequence (mostly pseudoautosomal genes annotated twice, a double-counting hazard in any ranked list), 11 assembly gap, 5 chrM contig-edge padding. All 120 are listed in excluded.tsv with their scores rather than dropped.
single_model · mediumEvery number comes from one model. There is no cross-model concordance and no independent measured evidence (ENCODE/SCREEN cCRE overlap, FANTOM5 CAGE, lentiMPRA) attached to any entry in this release. Model agreement would in any case be largely correlated error; only measured data can overrule a model.

Recall check: does the screen recover the promoters the field already uses?

The designed release carries no section like this one, and cannot. This question can only be asked of DNA that already exists in a genome and that other people have measured in a laboratory. Nothing in the designed release exists in any genome, so there is no field standard there to recover — and, for exactly the same reason, no rank there can be recall of training data.

Genome-wide ranking: 20,107 protein-coding TSS windows x 11 context strings = 221,177 predictions, ranked over the 19,987 windows that pass the sequence filters. No gene was prompted for, weighted, or pre-selected. The question asked before the run was whether the promoters the field already uses come back near the top. NOTE what 'blind' means: the ranking was blind to the answer, not to the genes. The model was trained on the expression of essentially every gene ranked here, including all five promoters in the headline below. Read training_split_caveat with this.

TNNT2 (cTnT)
#4
of 19,987 · cardiomyocyte vs skeletal myofiber · training gene
CKM (MCK)
#24
of 19,987 · myofiber vs hepatocyte · training gene
ALB
#1
of 19,987 · SERPINA1 rank 2 · training gene
GFAP
#2
of 19,987 · S100B rank 1 · training gene

hSYN1, one of the three promoters the go/no-go named explicitly, is the softest arm: SYN1 ranks 1,669 in A4 and 626 in A5. Both margins are reportable and the direction is right, so it is recovered -- but it is not near the top, and it fails the on-target expression floor outright. STMN2 and ENO2 outrank it and we cannot show the model is wrong about that. Two further field standards are missed in the all-vs-all view and recover only in their own pair: SERPINA7 (TBG), the standard liver promoter, ranks 1,994; CAMK2A ranks 11,866 all-vs-all and 331 in A4.

READ THIS BEFORE QUOTING THE NUMBERS ABOVE. Every promoter in those tiles was in the model's training data, so a high rank here is partly recall and not only prediction. How big that is, and what survives it, is the rest of this section.

How big it is. 118 of the 125 provenanced top-25 rows (25 per pair) are genes in the model's training split -- 94.4%, against a held-out base rate of 21.0%. (4,053 of the 19,987 rankable windows are held out, test plus validation; 21.0% is that share among the windows carrying a split label.) And ten of the eleven promoters highlighted in the recall evidence above are training genes.

What survives it. Re-ranking ONLY the 4,053 held-out genes, with the identical gates and no score changed, four of the five field-standard promoters that are held out are recovered under both gates: ENO2 (NSE) at held-out rank 17 of 4,053 in A4 and 11 in A5, MYL1 at 3 in A2, DES (desmin) at 60 in A2, SNAP25 at 101 in A4. The fifth is SYN1, which fails its on-target floor in the all-genes ranking too and is neither rescued nor worsened by excluding memorisation. MYBPC3 -- a held-out cardiac gene -- is the top held-out hit in both cardiac pairs. These cannot be memorisation: the model never saw their expression.

pairheld-out generalisation
A2strongheld-out genes clear both gates at 1.17x the training rate
A4strong1.09x the training rate; all ten of its held-out top 10 are recognised neuronal genes
A1cleanthe raw 0.69x looks like a deficit but is pool composition: the held-out pool clears the cardiomyocyte on-target floor at 0.62x the training rate. Adjusted, A1 observes 95 held-out survivors against 85.1 expected (p = 0.87) -- no deficit
A3not demonstratedexactly ONE of 4,053 held-out genes clears both gates, a ~6x deficit after the same pool adjustment (1 observed vs 6.1 expected, p = 0.016). Separately, none of A3's four reference promoters is held out, so no field standard can test A3 on this split at all. A3 IS NOT VALIDATED -- treat its shortlist as a hypothesis list only
A5not demonstratedten of 4,053 held-out genes clear both gates, a ~2x deficit after adjustment (10 observed vs 19.8 expected, p = 0.012)

How to quote this. Quote a rank together with its training status. 'The screen rediscovered cTnT at rank 4' is not usable without 'and cTnT was in the training data'. 'Four of five held-out field standards are recovered under both gates' is usable, with its n -- four promoters, three pairs -- and never for A3, which has no held-out field standard and one held-out survivor. Prefer held-out genes when re-running this check.

What is still confounded. The split is locus-blocked -- assigned over contiguous genomic intervals -- so held-out genes are concentrated on some chromosomes and nearly absent from others, and the tissue programmes encoded in tandem arrays travel with them. The pool-availability adjustment above removes the measurable part of that; composition within the available set is not removable on this split.

This method recovers the promoters the field already uses, in five of the six A-tier pairs it was checked on -- A1 through A5 pass; the sixth, A6, failed and is stated as an exclusion under Withdrawn below rather than dropped -- from a genome-wide ranking. It is not evidence that any individual novel candidate works. It is also not, on its own, evidence that the model generalises rather than remembers: the promoters in the table above were in its training data. The claim that survives that objection is the held-out one above -- four of the five field-standard promoters that are held out are recovered under both gates, which holds for A1, A2 and A4 and not for A3 or A5.

Every figure in this section, per pair and per promoter, is in recall_check.json (JSON download).

The cell pairs

pairon-target contextoff-target contextclears both gatesgenome-wide, not just the rows belowprovisionalinside the noise band of its own flooron-target floormedian genewhat a RANDOM gene achieves in this pair
A1cardiomyocyte_ventricularhepatocyte_primary6181904.5+0.7446
A2skeletal_myofiberhepatocyte_primary1,9201,5011.8203+0.2471
A3cardiomyocyte_ventricularskeletal_myofiber38314.5+0.457
A4cns_neuronhepatocyte_primary1,4505543.5+0.5
A5cns_neuronastrocyte961423.5+0.1719

The last column is the score an ordinary gene already gets — the level a result has to beat before it means anything. It is the median gene of the whole genome, measured in that pair, and it is not the same number twice: contexts sit at different genome-wide baselines, so raw margins are not comparable across pairs. In A1 the median gene already scores +0.7446 — a raw margin of 0.8 there is what a random gene achieves. Every record below carries that number as the leftmost mark on its own axis, and percentile_margin is the cross-pair quantity.

What was attempted and not shipped

Withdrawn

A6 — Retinal pigment epithelium ON / Hepatocyte OFF. Failed its own go/no-go. RPE65's own promoter does not clear the reportable margin threshold against liver -- and RPE65 is the gene whose loss causes the disease treated by the only approved RPE-directed gene therapy, so if any RPE promoter should have been easy to recover, it was this one. (That therapy delivers RPE65 as the transgene under a general-purpose viral promoter; it does not use the RPE65 promoter. The point here is the gene's standing in this cell type, not its construct.) One of five RPE markers is reportable. A coarse cross-organ contrast that should have been easy.

B2 — Photoreceptor ON / Retinal pigment epithelium OFF. 0 of 9 RPE marker genes separate RPE from photoreceptor. The two retinal contexts are not resolvable from each other on this checkpoint, so no ranked list in either direction is defensible.

Ranked, not in this release

These pairs did not fail. They are ranked internally and do not yet carry the per-pair evidence the statements above assume, so every count and verdict here excludes them.

A7 — Astrocyte ON / Hepatocyte (liver parenchymal cell) OFF. Ranked 2026-08-14 and held out of this release until it carries what every shipped pair carries: a row in the recall check, a per-pair verdict in the held-out re-analysis (which covers A1-A5 only), and a stated rationale. Its ranking exists internally; nothing about it is published here, and no number in this release includes it.

Natural promoters

The top 25 per pair plus every reference promoter, whether or not it passes, each with full provenance. The full database — every candidate clearing both gates in every pair — and the windows the sequence filters threw out, with their scores, are published as tab-separated files. Downloads: candidates.tsv — the records below, flat; shortlists.tsv — every candidate clearing both gates in every pair; excluded.tsv — the windows the sequence filters removed, with their scores.

How to read a record
the check ran and this record cleared itthe check ran and this record is marginal, or clears it with a caveatthe check ran and this record did not clear itno measurement exists — not a failure, an absence

What this release was built from

assembly     GRCh38
annotation   GENCODE release 50 (Ensembl 116)
MANE         NCBI MANE v1.4
window       9198 bp, TSS at 0-based offset 4599, gene-sense
model        g0-expression rev v1, checkpoint 20260523
unit         ln(quantile-normalised TPM + 1)
endpoint     POST https://api.genomicintelligence.ai/v1/tasks/expression/predict

The 5 context strings behind the pairs in this release — the ones whose noise floor and marker controls we measured — are in contexts.json, with complete SHA-256 digests. They are reproduced verbatim and must not be corrected — why, on the overview, along with how the window above is built and what silently goes wrong when it is not.

Every row in candidates.jsonl carries a reproduce block with the exact coordinates, request body and expected values for that row. The release as a whole — every field above, the file inventory and the limitations block — is in MANIFEST.json (JSON download).