DataType

162 terms

Computational output types — the data structure or processing result a file contains. Describes HOW data was produced or represented.

alignment (11)

annotation (5)

chromatin accessibility (9)

contact matrix (3)

count matrix (8)

crispr screen (10)

deep learning (18)

bias models

Trained models capturing sequencing bias patterns based on sequence composition.

counts sequence contribution scores

Per-nucleotide importance scores explaining sequence contribution to predicted counts (e.g., DeepLIFT, integrated gradients).

DNN-MPRA contribution scores

Nucleotide contribution scores from a deep neural network trained on MPRA data.

DNN-MPRA predicted signal

Regulatory activity signal predicted by a DNN trained on MPRA data.

model performance metrics

Evaluation metrics (AUC, correlation, etc.) assessing predictive model performance.

models

Trained computational or machine learning models saved for prediction or interpretation.

motif model

Sequence motif model (e.g., convolutional filter weights) from deep learning or motif discovery.

EDAM

profile sequence contribution scores

Per-nucleotide importance scores explaining sequence contribution to predicted signal profiles.

promoter prediction model

Computational model trained to predict promoter activity from sequence.

selected regions for bias-corrected predicted signal profile

Genomic regions selected for bias-corrected signal profile interpretation.

selected regions for count sequence contribution scores

Genomic regions selected for count contribution score analysis.

selected regions for predicted bias profile

Genomic regions selected for predicted bias profile interpretation.

selected regions for predicted signal and sequence contribution scores

Genomic regions selected for combined signal and contribution score analysis.

selected regions for predicted signal profile

Genomic regions selected for predicted signal profile interpretation.

selected regions for profile sequence contribution scores

Genomic regions selected for profile contribution score analysis.

TF binding prediction model

Deep learning model predicting transcription factor binding from DNA sequence.

training and test regions

Genomic regions designated for model training and held-out evaluation.

training set

Data used to train a computational or machine learning model.

peak set (16)

bidirectional peaks

Peaks from bidirectional transcription signal, characteristic of active enhancers and promoters.

→ enhancers, promoters

EDAM

conservative IDR thresholded peaks

Peaks using a conservative (stricter) IDR cutoff, yielding high-confidence but smaller peak set.

→ TF binding sites, open chromatin regions

EDAM

distal peaks

Peaks located distal (>2-3kb) from transcription start sites, often representing enhancers.

EDAM

divergent peaks

Peaks from divergent transcription where initiation occurs in both directions from a central point.

EDAM

IDR ranked peaks

Peaks ranked by IDR score, with lower IDR indicating higher reproducibility across replicates.

EDAM

optimal IDR thresholded peaks

Peaks using the optimal IDR cutoff balancing sensitivity and reproducibility.

→ TF binding sites, open chromatin regions

EDAM

peaks

Discrete genomic regions of statistically significant enrichment from peak calling. The fundamental unit of ChIP-seq and ATAC-seq analysis.

→ TF binding sites, open chromatin regions, histone modifications, enhancers

EDAM

peaks and background as input for IDR

Combined peak and background signal data formatted as input for IDR analysis.

EDAM

proximal peaks

Peaks located proximal (<2-3kb) to transcription start sites, often representing promoters.

EDAM

pseudoreplicated IDR thresholded peaks

IDR-thresholded peaks from pseudoreplicates (subsampled reads) when true replicates unavailable.

EDAM

pseudoreplicated peaks

Peak calls from pseudoreplicates created by subsampling reads from a single experiment.

EDAM

replicated peaks

Peaks reproducibly called across biological or technical replicates.

→ TF binding sites, open chromatin regions, histone modifications

EDAM

representative DNase hypersensitivity sites

A curated representative set of DNase hypersensitivity sites for reference.

EDAM

representative IDR thresholded peaks

A representative set of IDR-thresholded peaks selected for downstream analysis.

EDAM

unidirectional peaks

Peaks from unidirectional transcription signal, typically associated with gene bodies.

EDAM

valleys

Local minima in signal tracks used in footprint detection or nucleosome positioning analysis.

quantification (20)

differential expression quantifications

Statistical results from differential expression analysis comparing conditions.

EDAM

differential splicing quantifications

Statistical results from differential splicing analysis comparing conditions.

EDAM

element quantifications

Quantification values for regulatory elements.

exon quantifications

Read counts or expression values quantified at individual exons.

→ exon usage, splicing

EDAM

gene quantifications

Expression quantifications at gene level as read counts, TPM, or FPKM values.

→ gene expression

EDAM

gene stabilities

Measurements of mRNA or gene expression stability over time.

EDAM

genic features quantifications

Quantifications across various genic features (exons, introns, UTRs).

EDAM

genic regions quantifications

Read count quantifications over defined genic regions.

EDAM

merged transcription segment quantifications

Quantifications from merged transcription segments.

EDAM

microRNA quantifications

Expression quantifications of microRNAs (miRNAs).

EDAM

mRNA stabilities

Measurements of mRNA half-life or decay rates.

EDAM

novel peptides

Peptides identified that are absent from reference databases.

EDAM

peptide quantifications

Abundance measurements of peptides from mass spectrometry proteomics.

EDAM

protein expression quantifications

Abundance measurements of proteins from proteomics data.

EDAM

scaled RNA stability

RNA stability measurements scaled across samples.

EDAM

transcript quantifications

Expression quantifications at transcript isoform level.

→ transcript expression, splicing

EDAM

transcribed region quantifications

Quantifications over transcribed genomic regions.

EDAM

transcription segment quantifications

Quantifications over discrete transcription segments.

EDAM

UV enriched segment quantifications

Quantifications from UV-crosslinking enriched RNA segments.

EDAM

modified peptide quantification

Quantification of post-translationally modified peptides.

reference (17)

regulatory element (2)

signal track (13)

single cell (4)

technical (24)

capture targets

Genomic regions targeted for enrichment in capture-based sequencing (exome, panels).

exclusion list regions

Genomic blacklist regions excluded due to mapping artifacts or technical issues.

filtered regions

Genomic regions removed from analysis after filtering.

fragments

DNA or RNA fragment data prior to alignment.

idat green channel

Green channel intensity data from Illumina IDAT microarray files.

idat red channel

Red channel intensity data from Illumina IDAT microarray files.

inclusion list

Allowlist of genomic regions or barcodes included in analysis.

index reads

Index read sequences for sample demultiplexing.

intensity values

Raw intensity measurements from microarray or imaging experiments.

kmer weights

Frequency or weight values for k-mer sequences.

library fraction

Proportion of sequencing library represented by a sample or subset.

mitochondrial exclusion list regions

Mitochondrial regions excluded from nuclear genome analysis.

nanopore signal

Raw ionic current signal from nanopore sequencing.

negative control regions

Genomic regions used as negative controls in experiments.

positive control regions

Genomic regions used as positive controls in experiments.

primer sequence

Oligonucleotide primer sequences used in PCR or sequencing.

R2C2 subreads

Rolling circle amplification sub-reads from R2C2 long-read sequencing.

raw data

Unprocessed experimental data in original format before computational processing.

raw imaging signal

Unprocessed signal from imaging-based experiments.

sequence adapters

Adapter sequences ligated to library fragments for sequencing platform compatibility.

sequence barcodes

Short DNA sequences labeling samples (multiplexing) or individual cells (single-cell).

spike-ins

Exogenous sequences of known concentration added for normalization (e.g., ERCC RNA spike-ins).

subreads

Sub-read data from long-read sequencing platforms (PacBio).

validation

Data generated for experimental validation purposes.

variant (2)