Gene Selection via Annotation Frequency Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for predicting protein functions from databases require supervised machine-learning, which necessitates both positive and negative examples, limiting their ability to predict functions without such examples.
Innovation Solution
A device and method that selects genes or proteins relevant to a subject by gathering and choosing annotations linked more frequently than control genes or proteins, using a data warehouse to identify statistically significant associations without the need for positive and negative examples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If supervised machine-learning is used to predict protein functions from databases, then prediction accuracy can be improved, but the method cannot predict functions for proteins without positive and negative examples
Solution Approach 1:
The system uses the protein's own annotations and their frequency of association with genes to predict function, without requiring external positive or negative examples. The annotation frequency itself serves as the predictive signal, allowing the system to self-determine functional relevance based on inherent data patterns rather than trained examples.
Solution Approach 2:
The patent introduces an intermediary approach by using annotation frequency as a mediator between the protein and its function. Instead of directly comparing proteins with known examples, the system uses the frequency of annotations linking proteins to genes as an intermediate metric to infer functional relationships, enabling prediction without direct example matching.
2Quantity of substance
If the number of candidate genes is increased to improve coverage, then comprehensiveness is improved, but the difficulty of narrowing down relevant genes increases
Solution Approach 1:
The system changes the parameter for gene relevance assessment from binary presence/absence to frequency-based annotation counting. By transforming the evaluation metric to annotation frequency, the system can effectively rank and narrow down genes from large candidate sets, making the detection of relevant genes feasible even when the candidate list is comprehensive and large.
Data Source
AI summary
The present invention provides a device, method and program for selecting genes or proteins from a set of candidate genes or proteins so that the selected genes or proteins have a stronger relevance to a specific subject. The device of the present invention contains a storage device, an input device and a processor. The storage device stores a data warehouse that contains a data about a collection of genes or proteins, with which annotations are associated. The input device receives an input of the set of candidate genes or proteins. The processor (a) gathers annotations that are associated with the candidate genes or proteins, (b) chooses annotations that are associated with the candidate genes or proteins more than a threshold number of times or frequencies, and (c) selects genes or proteins, with which at least one of the chosen annotations is associated.


