DataType
162 terms
Computational output types — the data structure or processing result a file contains. Describes HOW data was produced or represented.
alignment (11)
alignments
Sequencing reads mapped to positions in a reference genome, typically in BAM/CRAM format with mapping quality scores and alignment coordinates.
→ sequencing reads
alignments with modifications
Aligned reads preserving base modification information (e.g., methylation from bisulfite-seq or direct detection) encoded in BAM auxiliary fields.
diploid personal genome alignments
Reads aligned to a diploid personal genome reference including both parental haplotypes, enabling allele-specific analysis.
gene alignments
Reads or sequences aligned at the gene level.
preprocessed alignments
Alignments after preprocessing: duplicate marking, base quality score recalibration, or indel realignment.
reads
Raw or minimally processed sequencing reads in FASTQ format, including quality scores and read identifiers.
redacted alignments
Alignments with sensitive genomic positions masked or removed for privacy protection in controlled-access data sharing.
redacted transcriptome alignments
Transcriptome-level alignments with sensitive positions redacted.
rejected reads
Reads that failed quality control filters and were excluded from downstream analysis.
spike-in alignments
Reads aligned to exogenous spike-in control sequences (e.g., ERCC) for normalization and quality assessment.
transcriptome alignments
Reads aligned to a transcriptome reference (cDNA sequences) rather than the genome.
→ gene expression, splicing
annotation (5)
HMM predicted chromatin state
Chromatin state annotations (e.g., ChromHMM) predicted by hidden Markov model from histone marks.
read annotations
Per-read annotations of alignment features or classifications.
semi-automated genome annotation
Genome annotation produced by a combination of automated and manual methods.
sequence alignability
Track indicating mappability of sequences to the reference genome.
sequence uniqueness
Track indicating uniqueness of k-mer sequences across the genome.
chromatin accessibility (9)
consensus DNase hypersensitivity sites
DNase I hypersensitivity sites consistently identified across multiple samples or cell types.
DHS peaks
Peak calls from DNase I hypersensitivity sequencing (DNase-seq), indicating open chromatin regions.
→ open chromatin regions, footprints
DHS regions reference
Reference set of DNase I hypersensitivity regions.
FDR cut rate
False discovery rate-controlled cut rate signal from DNase-seq analysis.
hotspots
Broad regions of elevated DNase cleavage activity representing domains of chromatin accessibility.
→ open chromatin regions
hotspots1 reference
Reference hotspot calls at lenient threshold (hotspot1 algorithm).
hotspots2 reference
Reference hotspot calls at stringent threshold (hotspot2 algorithm).
nuclease cleavage frequency
Per-base frequency of DNase I or Tn5 transposase cleavage across the genome.
nuclease cleavage corrected frequency
Bias-corrected per-base nuclease cleavage frequency.
contact matrix (3)
contact matrix
Genome-wide matrix of chromatin interaction frequencies from Hi-C or similar chromosome conformation capture experiments.
→ chromatin loops, TADs, compartments
pairs
Raw read pair data from Hi-C or proximity ligation before matrix construction.
variants contact matrix
Contact matrix incorporating variant information.
count matrix (8)
fold over change matrix
Matrix of fold-change values relative to control across features.
sparse gene count matrix
Sparse matrix of gene-level read counts across cells or samples, standard for single-cell RNA-seq.
sparse peak count matrix
Sparse matrix of peak accessibility counts across cells, used in single-cell ATAC-seq.
sparse transcript count matrix
Sparse matrix of transcript-level counts across cells or samples.
TF peaks matrix
Matrix of transcription factor peak counts across samples.
z scores matrix
Matrix of z-score normalized values across features and samples.
sparse splice junction count matrix
Sparse count matrix of splice junctions.
signals matrix
Matrix of signal values across features and samples.
crispr screen (10)
element barcode mapping
Mapping of element barcodes used in CRISPR screen experiments.
gRNAs
Guide RNA sequences used in CRISPR screens.
guide locations
Genomic locations targeted by guide RNAs.
guide quantifications
Abundance quantifications of guide RNAs from screen data.
non-targeting gRNAs
Guide RNAs designed as non-targeting negative controls.
perturbation signal
Phenotypic signal associated with CRISPR perturbations.
ranked gRNAs
Guide RNAs ranked by screen performance or activity.
reporter code counts
Count data from reporter codes in CRISPR screen readouts.
safe-targeting gRNAs
Guide RNAs targeting genomic safe-harbor loci as controls.
sparse gRNA count matrix
Sparse count matrix of guide RNA abundances across cells.
deep learning (18)
bias models
Trained models capturing sequencing bias patterns based on sequence composition.
counts sequence contribution scores
Per-nucleotide importance scores explaining sequence contribution to predicted counts (e.g., DeepLIFT, integrated gradients).
DNN-MPRA contribution scores
Nucleotide contribution scores from a deep neural network trained on MPRA data.
DNN-MPRA predicted signal
Regulatory activity signal predicted by a DNN trained on MPRA data.
model performance metrics
Evaluation metrics (AUC, correlation, etc.) assessing predictive model performance.
models
Trained computational or machine learning models saved for prediction or interpretation.
motif model
Sequence motif model (e.g., convolutional filter weights) from deep learning or motif discovery.
profile sequence contribution scores
Per-nucleotide importance scores explaining sequence contribution to predicted signal profiles.
promoter prediction model
Computational model trained to predict promoter activity from sequence.
selected regions for bias-corrected predicted signal profile
Genomic regions selected for bias-corrected signal profile interpretation.
selected regions for count sequence contribution scores
Genomic regions selected for count contribution score analysis.
selected regions for predicted bias profile
Genomic regions selected for predicted bias profile interpretation.
selected regions for predicted signal and sequence contribution scores
Genomic regions selected for combined signal and contribution score analysis.
selected regions for predicted signal profile
Genomic regions selected for predicted signal profile interpretation.
selected regions for profile sequence contribution scores
Genomic regions selected for profile contribution score analysis.
TF binding prediction model
Deep learning model predicting transcription factor binding from DNA sequence.
training and test regions
Genomic regions designated for model training and held-out evaluation.
training set
Data used to train a computational or machine learning model.
peak set (16)
bidirectional peaks
Peaks from bidirectional transcription signal, characteristic of active enhancers and promoters.
→ enhancers, promoters
conservative IDR thresholded peaks
Peaks using a conservative (stricter) IDR cutoff, yielding high-confidence but smaller peak set.
→ TF binding sites, open chromatin regions
distal peaks
Peaks located distal (>2-3kb) from transcription start sites, often representing enhancers.
divergent peaks
Peaks from divergent transcription where initiation occurs in both directions from a central point.
IDR ranked peaks
Peaks ranked by IDR score, with lower IDR indicating higher reproducibility across replicates.
optimal IDR thresholded peaks
Peaks using the optimal IDR cutoff balancing sensitivity and reproducibility.
→ TF binding sites, open chromatin regions
peaks
Discrete genomic regions of statistically significant enrichment from peak calling. The fundamental unit of ChIP-seq and ATAC-seq analysis.
→ TF binding sites, open chromatin regions, histone modifications, enhancers
peaks and background as input for IDR
Combined peak and background signal data formatted as input for IDR analysis.
proximal peaks
Peaks located proximal (<2-3kb) to transcription start sites, often representing promoters.
pseudoreplicated IDR thresholded peaks
IDR-thresholded peaks from pseudoreplicates (subsampled reads) when true replicates unavailable.
pseudoreplicated peaks
Peak calls from pseudoreplicates created by subsampling reads from a single experiment.
replicated peaks
Peaks reproducibly called across biological or technical replicates.
→ TF binding sites, open chromatin regions, histone modifications
representative DNase hypersensitivity sites
A curated representative set of DNase hypersensitivity sites for reference.
representative IDR thresholded peaks
A representative set of IDR-thresholded peaks selected for downstream analysis.
unidirectional peaks
Peaks from unidirectional transcription signal, typically associated with gene bodies.
valleys
Local minima in signal tracks used in footprint detection or nucleosome positioning analysis.
quantification (20)
differential expression quantifications
Statistical results from differential expression analysis comparing conditions.
differential splicing quantifications
Statistical results from differential splicing analysis comparing conditions.
element quantifications
Quantification values for regulatory elements.
exon quantifications
Read counts or expression values quantified at individual exons.
→ exon usage, splicing
gene quantifications
Expression quantifications at gene level as read counts, TPM, or FPKM values.
→ gene expression
gene stabilities
Measurements of mRNA or gene expression stability over time.
genic features quantifications
Quantifications across various genic features (exons, introns, UTRs).
genic regions quantifications
Read count quantifications over defined genic regions.
merged transcription segment quantifications
Quantifications from merged transcription segments.
microRNA quantifications
Expression quantifications of microRNAs (miRNAs).
mRNA stabilities
Measurements of mRNA half-life or decay rates.
novel peptides
Peptides identified that are absent from reference databases.
peptide quantifications
Abundance measurements of peptides from mass spectrometry proteomics.
protein expression quantifications
Abundance measurements of proteins from proteomics data.
scaled RNA stability
RNA stability measurements scaled across samples.
transcript quantifications
Expression quantifications at transcript isoform level.
→ transcript expression, splicing
transcribed region quantifications
Quantifications over transcribed genomic regions.
transcription segment quantifications
Quantifications over discrete transcription segments.
UV enriched segment quantifications
Quantifications from UV-crosslinking enriched RNA segments.
modified peptide quantification
Quantification of post-translationally modified peptides.
reference (17)
chromosome sizes
File listing chromosome/contig names and lengths, required by many genomics tools.
chromosomes reference
Reference sequences for individual chromosomes.
elements reference
Reference set of annotated genomic elements.
genome index
Precomputed index enabling rapid sequence alignment to a reference genome.
genome reference
Reference genome sequence assembly used for alignment and annotation.
miRNA reference
Reference sequences and annotations for microRNAs.
mitochondrial genome index
Alignment index for the mitochondrial genome.
mitochondrial genome reference
Reference sequence for the mitochondrial genome.
motif clusters reference
Reference set of clustered sequence motifs.
phastcons score reference
Reference phastCons conservation scores across the genome.
reference
Generic reference data file.
repeats reference
Reference annotations of repetitive elements.
rRNA reference
Reference sequences for ribosomal RNA.
snRNA reference
Reference sequences for small nuclear RNA.
transcriptome index
Index for rapid alignment to transcriptome reference sequences.
transcriptome reference
Reference transcript sequences for a species, used for RNA-seq alignment.
tRNA reference
Reference sequences for transfer RNA.
regulatory element (2)
signal track (13)
base overlap signal
Signal computed from base-level read overlap counts at each genomic position.
bias-corrected predicted signal profile
Model-predicted signal after correction for sequence-composition bias.
control normalized signal
Signal normalized against a matched control experiment to remove background and technical artifacts.
enrichment
Quantitative enrichment score over background, measuring signal above expected noise level.
fold change over control
Signal expressed as the ratio of experimental signal to input/control, highlighting enrichment over background.
→ TF binding, histone modifications
signal
Quantitative signal track showing per-base or per-bin values across the genome, typically in bigWig format.
→ gene expression, chromatin accessibility, TF binding
signal p-value
Statistical significance track showing -log10(p-value) of enrichment at each position.
summed densities signal
Signal computed as the sum of per-base read densities across the region.
wavelet-smoothed signal
Signal smoothed using wavelet transform to reduce noise while preserving peak structure.
end position signal
Signal of read end positions.
signal profile
Signal profile across genomic positions.
bias profile
Sequencing bias profile across genomic positions.
control profile
Control signal profile.
single cell (4)
archr project
ArchR software project file with processed single-cell ATAC-seq data and analyses.
cell coordinates
Low-dimensional coordinates (UMAP, t-SNE, PCA) for single-cell visualization.
cell topic participation
Cell-level participation scores in latent topics from topic modeling.
clusters
Cell cluster assignments from unsupervised clustering of single-cell data.
technical (24)
capture targets
Genomic regions targeted for enrichment in capture-based sequencing (exome, panels).
exclusion list regions
Genomic blacklist regions excluded due to mapping artifacts or technical issues.
filtered regions
Genomic regions removed from analysis after filtering.
fragments
DNA or RNA fragment data prior to alignment.
idat green channel
Green channel intensity data from Illumina IDAT microarray files.
idat red channel
Red channel intensity data from Illumina IDAT microarray files.
inclusion list
Allowlist of genomic regions or barcodes included in analysis.
index reads
Index read sequences for sample demultiplexing.
intensity values
Raw intensity measurements from microarray or imaging experiments.
kmer weights
Frequency or weight values for k-mer sequences.
library fraction
Proportion of sequencing library represented by a sample or subset.
mitochondrial exclusion list regions
Mitochondrial regions excluded from nuclear genome analysis.
nanopore signal
Raw ionic current signal from nanopore sequencing.
negative control regions
Genomic regions used as negative controls in experiments.
positive control regions
Genomic regions used as positive controls in experiments.
primer sequence
Oligonucleotide primer sequences used in PCR or sequencing.
R2C2 subreads
Rolling circle amplification sub-reads from R2C2 long-read sequencing.
raw data
Unprocessed experimental data in original format before computational processing.
raw imaging signal
Unprocessed signal from imaging-based experiments.
sequence adapters
Adapter sequences ligated to library fragments for sequencing platform compatibility.
sequence barcodes
Short DNA sequences labeling samples (multiplexing) or individual cells (single-cell).
spike-ins
Exogenous sequences of known concentration added for normalization (e.g., ERCC RNA spike-ins).
subreads
Sub-read data from long-read sequencing platforms (PacBio).
validation
Data generated for experimental validation purposes.