Variation ID

Information

This page provides a variant-centered integrative annotation framework, offering a comprehensive view of a single genetic variant by combining population genetics, functional consequence annotation, regulatory effect prediction, association evidence, and molecular impact modeling. By integrating both traditional annotation methods and modern deep learning–based predictions, this module supports systematic interpretation of variant effects at the DNA, RNA, and protein levels.

Annotation Framework: Variant annotations in this module are organized into two complementary layers: (i) classical coding variant annotations, which focus on protein-coding consequences, and (ii) deep learning–based predictive models, which extend interpretation to regulatory, splicing, and protein-level effects.

Classical Coding Variant Annotations: Classical annotation methods primarily target variants located in protein-coding regions and infer functional consequences based on gene models, evolutionary conservation, and protein sequence or structural features.

  • SnpEff (Cingolani et al., 2012): annotates variants by mapping them to gene models and transcript structures, classifying consequences such as synonymous, missense, or stop-gain mutations and assigning impact levels.
  • Coovar (Vergara et al., 2012): provides complementary classification of coding variants by evaluating codon changes and their potential effects on protein sequences.
  • SIFT (Ng and Henikoff, 2003): predicts whether amino acid substitutions are likely to affect protein function based on evolutionary conservation across homologous sequences.
  • PolyPhen-2 (Adzhubei et al., 2010): evaluates the potential impact of amino acid substitutions by integrating sequence conservation, structural context, and physicochemical properties.

Deep Learning–based Variant Effect Predictions: To extend variant interpretation beyond classical coding annotations, this module integrates deep learning models trained on large-scale functional genomics and molecular datasets.

Regulatory Effect Prediction (Non-coding Variants): Regulatory effects of non-coding variants are predicted using the Basenji (Kelley et al., 2018), a deep learning sequence-to-signal model trained to predict chromatin accessibility from genomic sequence. Values represent log-ratio differences between the alternative (ALT) and reference (REF) alleles. Positive values indicate increased chromatin accessibility associated with the alternative allele, whereas negative values indicate reduced accessibility relative to the reference allele. Larger absolute values correspond to stronger predicted regulatory effects.

High-impact regulatory variants (HEVs): HEV cutoffs depend on both the reference genome and the variant set used for calibration. Each q90 cutoff is calculated from the maximum absolute predicted effect across the available tissues for every variant-allele record, before alternate alleles at the same genomic position are collapsed.

  • Nipponbare (IRGSP-1.0; vg) SNP and non-SNP variants: A light-red denotes an HEV at 0.052 (exact value 0.051688551903), a cutoff calibrated from the SNP-only q90 and applied to both SNP and non-SNP scores. A red denotes a strong HEV at the combined SNP and non-SNP q90 cutoff of 0.106 (exact value 0.10640232). When a score reaches the strong-HEV cutoff, the heatmap displays only rather than adding both symbols.
  • MH63RS3 (vm) SNPs: A light-red denotes an HEV at the SNP-only q90 cutoff of 0.052 (exact value 0.052074253559). No strong-HEV tier is used.
  • ZS97RS3 (vz) SNPs: A light-red denotes an HEV at the SNP-only q90 cutoff of 0.047 (exact value 0.047439455986). No strong-HEV tier is used.

All classifications use the absolute unrounded tissue score. A marker is shown in a tissue only when that tissue's score reaches the applicable cutoff. The term strong HEV describes a higher predicted effect-magnitude tier; it does not indicate greater statistical confidence or experimentally validated causality.

RNA Splicing Impact Prediction: Potential effects of variants on RNA splicing are predicted using the DRANetSplicer (Liu et al., 2024), which assesses whether variants may disrupt donor or acceptor splice sites.

Protein-level Impact Prediction:

  • Subcellular localization effects are predicted using DeepLoc 2.0 (Thumuluri et al., 2022), which assesses whether variants may alter protein targeting or intracellular distribution.

    Genome-wide percentile rank (0–100%): the percentage of all genome-wide DeepLoc variant–transcript prediction records with a Wasserstein Distance less than or equal to this result. A higher percentile indicates a larger predicted overall change in protein subcellular localization.

    The paired plots use the same subcellular-localization categories. The left plot shows the wild-type/reference (WT/REF) score, and the right plot shows the variant effect (ALT − REF). Red rightward bars indicate an increased predicted localization score, whereas blue leftward bars indicate a reduced score. Values are labelled directly; hovering also shows the WT/REF score, alternative score and exact change.

  • Protein stability and folding effects are evaluated using DDGun (Montanucci et al., 2022), PROSTATA (Umerenkov et al., 2023), and THPLM (Gong et al., 2023), which estimate variant-induced changes in protein stability.

GWAS Associations and External Evidence

  • Trait associations: GWAS results linking the variant to agronomic or metabolic traits are summarized, including associated populations and p-values.
  • Lead Variant identification: Variants are flagged when they represent the strongest association signal within a GWAS locus.
  • Literature evidence: Links to relevant publications are provided to support reported associations.
TOC

Navigation