ONGA

Ontology for Genomic Annotations

11 vocabularies + 5 schemas for describing genomic data files

Vocabularies

Closed value sets — the permissible values you choose from ("which one?").

Schemas

Descriptor properties you assert about a track ("describe this track"); slots draw on the vocabularies.

How they relate

The three core vocabularies are closed value sets, complemented by eight small facet vocabularies. The five schemas are sets of properties you assert about a track, with slots whose values come from those vocabularies or from primitives. Track Format describes how a track is encoded (file format, BED columns). Track Interpretation describes what it is / means, split into an algorithmic output_type (drawing on DataType) and a biological feature_type (drawing on FeatureType). Track Provenance describes what was done to it — the processing/derivation operations applied (read selection, QC filtering, scaling normalization, thresholding, and observed-vs-predicted derivation). Track Geometry describes its structural shape (point vs interval vs signal vs graph) — measured from the bytes, orthogonal to the others. Reference Genome describes the reference assembly a track is defined against (its identifier and build sex). Properties that describe the sample or assay rather than the file's content — anatomy, sample sex, method — are deliberately kept out of scope and delegated to external ontologies.

162
DataTypes
75
FeatureTypes
12
Formats
14
Geometry Properties
86
EDAM Mappings
DataTypes FeatureTypes Formats Strand ReadMultiplicity FilterStatus Normalization Thresholding Derivation ReferenceBuildSex HaplotypeResolution Track Format Track Interpretation Track Provenance Track Geometry Reference Genome Scope Boundary EDAM Mappings About ONGA