Ontology Propagation for Interpretable Omics Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current transcriptomics datasets are high-dimensional, making it difficult to identify genes associated with biological effects, and machine learning models often fail to integrate biological annotations or explain the importance of genes in predicting tasks.
Innovation Solution
A method involving the construction of a combined graph that integrates omics data with ontology data using a propagation algorithm to assign values to ontology terms, allowing for a functional interpretation of omics data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If machine learning models are used to analyze high-dimensional transcriptomics data, then classification performance is improved, but interpretability and explainability of gene associations deteriorate
Solution Approach 1:
The patent introduces ontology terms as intermediary concepts that bridge raw gene expression data and biological interpretation. Instead of directly mapping genes to disease states, the model uses ontology terms (GO terms) as mediators that encode biological knowledge about gene functions, processes, and pathways. This intermediary layer preserves interpretability by maintaining the biological semantic meaning while enabling complex pattern recognition in high-dimensional data.
Solution Approach 2:
The patent transforms the analysis from gene-centric high-dimensional space to ontology term space, effectively changing the dimensional representation. By aggregating gene expression data through ontology hierarchies, the method projects thousands of gene dimensions into a smaller set of biologically meaningful ontology dimensions, reducing complexity while preserving essential biological relationships.
2Productivity
If simple machine learning models are used, then computational efficiency is improved, but the ability to learn complex gene interactions deteriorates
Solution Approach 1:
The patent performs preliminary organization of gene data into ontology-based groups before feeding data to the machine learning model. By pre-aggregating gene expressions according to their functional annotations in the ontology hierarchy, the method simplifies the input structure and reduces the complexity of interactions the model must learn, enabling efficient processing while capturing complex biological relationships through the pre-structured ontology framework.
3Quantity of substance
If high-dimensional gene expression data is used directly, then data completeness is improved, but difficulty in identifying associated genes increases
Solution Approach 1:
The patent segments the high-dimensional gene expression data into smaller, biologically meaningful groups based on ontology annotations. Instead of analyzing all genes simultaneously in a single high-dimensional space, the method divides genes into functional categories (segments) according to their GO term associations, making the detection of associated genes more manageable and interpretable within each segmented group.
Solution Approach 2:
The patent merges information from multiple genes that share common ontology annotations into aggregated ontology term representations. By combining gene expression signals within ontology-defined groups, the method enhances the statistical power to detect associations while reducing the dimensionality from individual gene level to functional category level, making identification of biologically relevant patterns more efficient.
Data Source
AI summary
A medical information processing apparatus comprising processing circuitry configured to receive omics data comprising a plurality of biomolecules and a plurality of associated measured values; receive a first plurality of associations mapping a respective biomolecule to another respective biomolecule; receive ontology data based on the omics data, the ontology data comprising a plurality of ontology terms associated with at least one other ontology term and/or at least one other biomolecule; and assign a value to each of the plurality ontology terms based on the omics data, the ontology data, and the associations between them. A value can be assigned to each of the ontology terms based on a propagation algorithm.


