Syntropa

Tier 0 · open

Corpus atlas

The honest inventory: not only what the corpus can do, but what it cannot express. Coverage gaps are reported as findings, because a gap is a research target rather than an embarrassment.

Sources

Models come from eight sources: published collections, automated reconstructions from genome databases, and the APOLLO-595 pan-archaeal set. Source mix is declared because it sets the floor on what any discovery result can claim.

SourceModelsShare
embl_gems5,57435%
AGORA24,67230%
CarveMe4,32628%
AGORA6014%
APOLLO-5953662%
BiGG911%
PARADIGM770%
SysBioChalmers90%

What the corpus can do

Capability is measured from model exchange reactions, not asserted from literature. The counts below show how many models can grow on a minimal defined feedstock without organic growth factors — the rest require rich media (gap-filled).

Feedstock (defined medium)Models that growShare
sucrose3,38922%
fructose3,34821%
cellobiose3,33621%
arabinose3,25221%
glucose3,13120%
sulfate+glc3,05819%
formate2,97319%
syngas2,96219%

Anaerobic-capable:11,601 models (74%) can grow under anaerobic conditions.

Oxygen niche

NicheModelsShare
facultative2,29415%
anaerobe1,4289%
aerobe8716%
microaerophile4353%
nanaerobe2241%
aerotolerant40%

O2 annotation covers 5,256 of 15,717 models; the remainder have no annotation in the source metadata.

Hazard degradation capability

Models with a confirmed metabolic route to consume known environmental hazards — measured by exchange-reaction presence, not by literature curation.

HazardModels with degradation routeShare
nitrite10,89069%
benzoate10,09664%
catechol7,58448%
toluene7,54148%
formaldehyde7,22346%
cyanide5,29534%
phenol2452%

Quality and coverage

Quality is graded by canonical-identifier coverage — how many exchange metabolites map to a MetaNetX universal ID. Full coverage (>99%) enables cross-model comparisons with no disambiguation step.

Coverage tierModelsShare
Full canonical coverage (>99%)10,08964%
High coverage (>90%)2,86318%
Mid coverage (>80%)2,59016%
Low coverage (<80%)4593%
QC gateModels
Passed (discovery-ready)15,636
Exploratory (draft / unverified)79

Biosafety

BSL annotation is from DSMZ / ATCC metadata where available. Models without a verified annotation are listed as unverified — a missing annotation is not an implicit BSL-1 clearance.

Biosafety tierModels
unverified7,724
BSL15,133
BSL22,860
BSL38

Of 15,717 models, 1,737 carry a human-pathogen annotation. Discovery runs flag these automatically and exclude them from unsupervised proposals.

Licensing

License status is tracked at the model level. Redistributable models may be published as part of a derived corpus; link-only models require fetching from the source repository. The Data page has the full breakdown by source with citation links.

License classModelsShare
Openly redistributable10,27665%
Link-only (access from source)5,44135%

What this corpus cannot express

Gap-filled reconstructions. Most models are auto-reconstructed with gap-filling, which adds transporters and reactions to make the model biomass-viable on rich media. This inflates apparent metabolic capacity: a gap-filled transporter is a computational assumption, not a measured fact. Discovery results are candidates, not confirmations.

Eukaryotes. The current corpus is bacteria and archaea. Plant, algal, and fungal models are in development. Rhizobia–legume nitrogen fixation is the near-term priority.

Defined-medium viability. Most models require organic growth factors that do not exist in a minimal mineral medium. Obligate cross-feeding dependencies that rely on one partner supplying these factors cannot be confirmed in situ with gap-filled models; they are confirmed instead by the thermodynamic gate on specific energy reactions.

Corpus version