{
 "release_id": "gi-promoter-atlas-designs-2026-08-05",
 "use_rule": "Use these strings VERBATIM, byte for byte, including any apparent internal inconsistency. Tidying a real experimental-context string is not a paraphrase: making one HepG2 row's `mapped run type` agree with its own description clause moves ALB from 8.625 to 2.922 and drops the model into a collapsed regime that returns a near-population-average profile for every sequence while still looking entirely plausible.",
 "contexts": [
  {
   "arm": "off",
   "context_id": "astrocyte",
   "label": "Astrocyte (cellxgene snRNA pseudobulk, human striatum)",
   "cell_ontology_id": "CL:0000127",
   "biosample_class": "primary_cell",
   "description_string": "organism ontology term id is NCBITaxon:9606. tissue ontology term id is UBERON:0001873;UBERON:0001874. assay ontology term id is EFO:0009922. disease ontology term id is PATO:0000461. cell type ontology term id is CL:0000127. author cell type is Astrocytes. self reported ethnicity ontology term id is HANCESTRO:0005;HANCESTRO:0568. sex ontology term id is PATO:0000383;PATO:0000384. suspension type is nucleus. is primary data is True. development stage ontology term id is HsapDv:0000117;HsapDv:0000148;HsapDv:0000149;HsapDv:0000133;HsapDv:0000131;HsapDv:0000154. tissue type is tissue. cell type is astrocyte. assay is 10x 3' v3. disease is normal. organism is Homo sapiens. sex is female;male. tissue is caudate nucleus;putamen. self reported ethnicity is European;African American. development stage is 23-year-old stage;54-year-old stage;55-year-old stage;39-year-old stage;37-year-old stage;60-year-old stage. dataset id is c893ddc3-f25b-45e2-8c9e-155918b4261c. id is 3b5a65b240dfccb049b58d9becf893ab. genome is unknown.",
   "sha256": "7f78e3c9ea85cd67fbd3235c0133488b64d5d0f9b75a6f35c4eb0f87571f7329",
   "length_chars": 1027,
   "verbatim": true,
   "use_rule": "Use these strings VERBATIM, byte for byte, including any apparent internal inconsistency. Tidying a real experimental-context string is not a paraphrase: making one HepG2 row's `mapped run type` agree with its own description clause moves ALB from 8.625 to 2.922 and drops the model into a collapsed regime that returns a near-population-average profile for every sequence while still looking entirely plausible.",
   "marker_genes": [
    "S100B",
    "GFAP",
    "MLC1",
    "AQP4",
    "GJA1"
   ],
   "prompt_noise_floor": 0.0996,
   "prompt_noise_floor_method": "p90 of |score - baseline| over the core paraphrase families, per panel gene; the headline floor restricts to genes scoring below 1.0. Kit is src/gi_promoters/cellxgene_context.py, not the ENCODE kit in gi_promoters.context — a cellxgene row has no `description`, `life stage age`, `strand specificity` or `mapped run type` clause, and the ENCODE kit would have *added* a `life stage age` clause that does not exist in this vocabulary.",
   "admission_status": "admitted",
   "biosample_class_caveat": "`biosample_class: primary_cell` in a context row describes where the material came from, not whether it still behaves like the cell type on its label. These six contexts are cellxgene tissue-derived pseudobulk rather than expanded culture, which is why they pass their marker checks -- but that is a fact verified per context, never assumed."
  },
  {
   "arm": "on",
   "context_id": "cardiomyocyte_ventricular",
   "label": "Ventricular cardiomyocyte (cellxgene snRNA pseudobulk)",
   "cell_ontology_id": "CL:0002131",
   "biosample_class": "primary_cell",
   "description_string": "NRP is No;Yes. cell source is Harvard-Nuclei;Sanger-Nuclei. donor id is H5;H6;H3;H2;H7;H4;D11. source is Nuclei. type is DBD;DCD. cell states is vCM1;vCM2;vCM3;vCM4;vCM5. Used is Yes. disease ontology term id is PATO:0000461. assay ontology term id is EFO:0009922. cell type original is Ventricular_Cardiomyocyte. tissue ontology term id is UBERON:0002098;UBERON:0002084;UBERON:0002080;UBERON:0002094. development stage ontology term id is HsapDv:0000240;HsapDv:0000239;HsapDv:0000241. cell type ontology term id is CL:0002131. suspension type is nucleus. self reported ethnicity ontology term id is HANCESTRO:0005;HANCESTRO:0008. sex ontology term id is PATO:0000383;PATO:0000384. is primary data is True. organism ontology term id is NCBITaxon:9606. tissue type is tissue. cell type is regular ventricular cardiac myocyte. assay is 10x 3' v3. disease is normal. organism is Homo sapiens. sex is female;male. tissue is apex of heart;heart left ventricle;heart right ventricle;interventricular septum. self reported ethnicity is European;Asian. development stage is sixth decade stage;fifth decade stage;seventh decade stage. dataset id is d4e69e01-3ba2-4d6b-a15d-e7048f78f22e. id is 9538b2cdcc9f2492f980aece7866b335. genome is unknown. n obs is 73066. total reads is 69718896.0.",
   "sha256": "f0274863ae594d39a712c19f3a1382ed8d8dca6f5431088e9e429dab4549cadb",
   "length_chars": 1279,
   "verbatim": true,
   "use_rule": "Use these strings VERBATIM, byte for byte, including any apparent internal inconsistency. Tidying a real experimental-context string is not a paraphrase: making one HepG2 row's `mapped run type` agree with its own description clause moves ALB from 8.625 to 2.922 and drops the model into a collapsed regime that returns a near-population-average profile for every sequence while still looking entirely plausible.",
   "marker_genes": [
    "MYH6",
    "TNNT2",
    "MYBPC3",
    "ACTC1",
    "NPPA",
    "NPPB",
    "MYL2",
    "MYH7"
   ],
   "prompt_noise_floor": 0.166,
   "prompt_noise_floor_method": "p90 of |score - baseline| over the core paraphrase families, per panel gene; the headline floor restricts to genes scoring below 1.0. Kit is src/gi_promoters/cellxgene_context.py, not the ENCODE kit in gi_promoters.context — a cellxgene row has no `description`, `life stage age`, `strand specificity` or `mapped run type` clause, and the ENCODE kit would have *added* a `life stage age` clause that does not exist in this vocabulary.",
   "admission_status": "admitted",
   "biosample_class_caveat": "`biosample_class: primary_cell` in a context row describes where the material came from, not whether it still behaves like the cell type on its label. These six contexts are cellxgene tissue-derived pseudobulk rather than expanded culture, which is why they pass their marker checks -- but that is a fact verified per context, never assumed."
  },
  {
   "arm": "on",
   "context_id": "cns_neuron",
   "label": "CNS neuron (cellxgene snRNA pseudobulk, human thalamus)",
   "cell_ontology_id": "CL:0000540",
   "biosample_class": "primary_cell",
   "description_string": "roi is Human ANC. organism ontology term id is NCBITaxon:9606. disease ontology term id is PATO:0000461. self reported ethnicity ontology term id is HANCESTRO:0005. assay ontology term id is EFO:0009922. sex ontology term id is PATO:0000384. development stage ontology term id is HsapDv:0000136;HsapDv:0000144;HsapDv:0000123. donor id is H19.30.001;H18.30.002;H19.30.002. suspension type is nucleus. dissection is Thalamus (THM) - Anterior nuclear complex - ANC. sample id is 10X354_5;10X190_3;10X190_4;10X348_3;10X354_6;10X348_4. cell type ontology term id is CL:0000540. tissue ontology term id is UBERON:0010225. is primary data is True. tissue type is tissue. cell type is neuron. assay is 10x 3' v3. disease is normal. organism is Homo sapiens. sex is male. tissue is thalamic complex. self reported ethnicity is European. development stage is 42-year-old stage;50-year-old stage;29-year-old stage. dataset id is e8681d74-ac9e-4be5-be14-1cf1bbd54dd7. id is 15ab0118ee8683ff8782be951f4613ba. genome is unknown. n obs is 24985. total reads is 860872832.0.",
   "sha256": "dba9daef8fe2381000f61e45f841239c48a571da865af1f6cfc7cc2d31477247",
   "length_chars": 1058,
   "verbatim": true,
   "use_rule": "Use these strings VERBATIM, byte for byte, including any apparent internal inconsistency. Tidying a real experimental-context string is not a paraphrase: making one HepG2 row's `mapped run type` agree with its own description clause moves ALB from 8.625 to 2.922 and drops the model into a collapsed regime that returns a near-population-average profile for every sequence while still looking entirely plausible.",
   "marker_genes": [
    "STMN2",
    "SYT1",
    "GRIN1",
    "NRGN",
    "GAD1",
    "ENO2",
    "TUBB3",
    "RBFOX3",
    "SNAP25"
   ],
   "prompt_noise_floor": 0.2812,
   "prompt_noise_floor_method": "p90 of |score - baseline| over the core paraphrase families, per panel gene; the headline floor restricts to genes scoring below 1.0. Kit is src/gi_promoters/cellxgene_context.py, not the ENCODE kit in gi_promoters.context — a cellxgene row has no `description`, `life stage age`, `strand specificity` or `mapped run type` clause, and the ENCODE kit would have *added* a `life stage age` clause that does not exist in this vocabulary.",
   "admission_status": "admitted",
   "biosample_class_caveat": "`biosample_class: primary_cell` in a context row describes where the material came from, not whether it still behaves like the cell type on its label. These six contexts are cellxgene tissue-derived pseudobulk rather than expanded culture, which is why they pass their marker checks -- but that is a fact verified per context, never assumed."
  },
  {
   "arm": "off",
   "context_id": "hepatocyte_primary",
   "label": "Primary hepatocyte (cellxgene scRNA pseudobulk, Tabula Sapiens liver)",
   "cell_ontology_id": "CL:0000182",
   "biosample_class": "primary_cell",
   "description_string": "tissue in publication is Liver. anatomical position is NA. method is 10X. assay ontology term id is EFO:0009922. cell type ontology term id is CL:0000182. compartment is Epithelium. broad cell class is hepatocyte. free annotation is hepatocyte. organism ontology term id is NCBITaxon:9606. suspension type is cell. tissue type is tissue. tissue ontology term id is UBERON:0002107. disease ontology term id is PATO:0000461. is primary data is True. sex ontology term id is PATO:0000384. self reported ethnicity ontology term id is HANCESTRO:0005. development stage ontology term id is HsapDv:0000154. cell type is hepatocyte. assay is 10x 3' v3. disease is normal. organism is Homo sapiens. sex is male. tissue is liver. self reported ethnicity is European. development stage is 60-year-old stage. dataset id is 53d208b0-2cfd-4366-9866-c3c6114081bc. id is fc4e0d9da79519bc9db281af9631869d. genome is unknown.",
   "sha256": "350f5ccd7f1e91435428145b600a879dcba3aed966586b446b69ed82782b3e12",
   "length_chars": 907,
   "verbatim": true,
   "use_rule": "Use these strings VERBATIM, byte for byte, including any apparent internal inconsistency. Tidying a real experimental-context string is not a paraphrase: making one HepG2 row's `mapped run type` agree with its own description clause moves ALB from 8.625 to 2.922 and drops the model into a collapsed regime that returns a near-population-average profile for every sequence while still looking entirely plausible.",
   "marker_genes": [
    "ALB",
    "SERPINA1",
    "APOA1",
    "TTR",
    "APOB",
    "FGA",
    "CYP3A4"
   ],
   "prompt_noise_floor": 0.2695,
   "prompt_noise_floor_method": "p90 of |score - baseline| over the core paraphrase families, per panel gene; the headline floor restricts to genes scoring below 1.0. Kit is src/gi_promoters/cellxgene_context.py, not the ENCODE kit in gi_promoters.context — a cellxgene row has no `description`, `life stage age`, `strand specificity` or `mapped run type` clause, and the ENCODE kit would have *added* a `life stage age` clause that does not exist in this vocabulary.",
   "admission_status": "admitted",
   "biosample_class_caveat": "`biosample_class: primary_cell` in a context row describes where the material came from, not whether it still behaves like the cell type on its label. These six contexts are cellxgene tissue-derived pseudobulk rather than expanded culture, which is why they pass their marker checks -- but that is a fact verified per context, never assumed."
  },
  {
   "arm": "on",
   "context_id": "skeletal_myofiber",
   "label": "Skeletal muscle fiber (cellxgene snRNA pseudobulk, hamstring tendon)",
   "cell_ontology_id": "CL:0008002",
   "biosample_class": "primary_cell",
   "description_string": "cell type is skeletal muscle fiber. tissue is tendon of semitendinosus. genome is GENCODE 33 hg38. organism is Homo sapiens. author cell type is Transitional skeletal muscle cells. development stage is 18-year-old stage;24-year-old stage;26-year-old stage;37-year-old stage. disease is normal. sex is female;male. mapped reference assembly is GRCh38. mapped reference annotation is GENCODE 33. alignment software is kallisto bustools. donor id is MSK0782;MSK1139;MSK1144;MSK1216. self reported ethnicity ontology term id is HANCESTRO:0462;unknown. donor living at sample collection is True. organism ontology term id is NCBITaxon:9606. sample uuid is 5eb3139b-6e0b-48c5-9bce-ddabcf3a518d;c1d37851-34b4-40dc-8729-051de8361e31;2cd105b1-6c73-4f7a-a5aa-9058773657f0;d398562a-4b4d-428c-8fc8-78c19ac19fb0. sample preservation method is flash-freezing. tissue ontology term id is UBERON:8480009. development stage ontology term id is HsapDv:0000112;HsapDv:0000118;HsapDv:0000120;HsapDv:0000131. sample derivation process is resection. sample source is Oxford. donor BMI at collection is 29.73;27.78;nan;24.38. tissue type is tissue. suspension derivation process is mechanical dissociation,detergent solubilization. suspension dissociation reagent is 0.5% CHAPS. suspension dissociation time is 10 minute. suspension uuid is e99a2f31-8ad8-4103-9b8e-d85c30529908;18717cdc-cc83-40f5-a499-49f58ec28113;df6a56a4-46c1-4660-91e8-53b6ceb012f5;473e3da0-fcca-4808-bd47-614237d76293. suspension type is nucleus. tissue handling interval is <2 hours. library uuid is 4733f679-f608-4b2e-a404-e09917bb4c87;93a7ac60-62a9-4332-99c6-f06d91b39693;ae0d8c72-b5b3-420f-bc0f-f6970050de23;7b00afb9-2727-44e0-97c5-918c20c03836. assay ontology term id is EFO:0009922. library starting quantity is 200-1000 nuclei. sequencing platform is Illumina NovaSeq 6000. is primary data is True. cell type ontology term id is CL:0008002. disease ontology term id is PATO:0000461. sex ontology term id is PATO:0000383;PATO:0000384. assay is 10x 3' v3. self reported ethnicity is British;unknown. dataset id is 06ef6b36-6c9b-4e10-8a94-d0baf274276e. id is 7cd66380206b0b7c32be330200b43a6a. n obs is 1684. total reads is 4529096.0.",
   "sha256": "8560e1d77aa30c015f1b5b9bc4f025311c7d98fed357011993e9d3c21367d01d",
   "length_chars": 2185,
   "verbatim": true,
   "use_rule": "Use these strings VERBATIM, byte for byte, including any apparent internal inconsistency. Tidying a real experimental-context string is not a paraphrase: making one HepG2 row's `mapped run type` agree with its own description clause moves ALB from 8.625 to 2.922 and drops the model into a collapsed regime that returns a near-population-average profile for every sequence while still looking entirely plausible.",
   "marker_genes": [
    "NEB",
    "MYBPC1",
    "MYL1",
    "MYF6",
    "MYOG",
    "MYH1",
    "MYBPH",
    "TNNI1"
   ],
   "prompt_noise_floor": 0.3418,
   "prompt_noise_floor_method": "p90 of |score - baseline| over the core paraphrase families, per panel gene; the headline floor restricts to genes scoring below 1.0. Kit is src/gi_promoters/cellxgene_context.py, not the ENCODE kit in gi_promoters.context — a cellxgene row has no `description`, `life stage age`, `strand specificity` or `mapped run type` clause, and the ENCODE kit would have *added* a `life stage age` clause that does not exist in this vocabulary.",
   "admission_status": "admitted",
   "biosample_class_caveat": "`biosample_class: primary_cell` in a context row describes where the material came from, not whether it still behaves like the cell type on its label. These six contexts are cellxgene tissue-derived pseudobulk rather than expanded culture, which is why they pass their marker checks -- but that is a fact verified per context, never assumed."
  }
 ]
}
