Gene Selection via Annotation Frequency Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting protein functions from databases require supervised machine-learning, which necessitates both positive and negative examples, limiting their ability to predict functions without such examples.

Innovation Solution

A device and method that selects genes or proteins relevant to a subject by gathering and choosing annotations linked more frequently than control genes or proteins, using a data warehouse to identify statistically significant associations without the need for positive and negative examples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If supervised machine-learning is used to predict protein functions from databases, then prediction accuracy can be improved, but the method cannot predict functions for proteins without positive and negative examples

Engineering Contradiction:
Improveprediction accuracyVSAvoidapplicability to proteins without examples
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system uses the protein's own annotations and their frequency of association with genes to predict function, without requiring external positive or negative examples. The annotation frequency itself serves as the predictive signal, allowing the system to self-determine functional relevance based on inherent data patterns rather than trained examples.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary approach by using annotation frequency as a mediator between the protein and its function. Instead of directly comparing proteins with known examples, the system uses the frequency of annotations linking proteins to genes as an intermediate metric to infer functional relationships, enabling prediction without direct example matching.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If the number of candidate genes is increased to improve coverage, then comprehensiveness is improved, but the difficulty of narrowing down relevant genes increases

Engineering Contradiction:
Improvenumber of candidate genesVSAvoiddifficulty of narrowing down
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The system changes the parameter for gene relevance assessment from binary presence/absence to frequency-based annotation counting. By transforming the evaluation metric to annotation frequency, the system can effectively rank and narrow down genes from large candidate sets, making the detection of relevant genes feasible even when the candidate list is comprehensive and large.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9141755B2Device and method for selecting genes and proteins
Publication Date: 2015.09.22 NAT INST OF BIOMEDICAL INNOVATION HEALTH & NUTRITION
  • US9141755B2 patent drawing
  • US9141755B2 patent drawing
  • US9141755B2 patent drawing

AI summary

The present invention provides a device, method and program for selecting genes or proteins from a set of candidate genes or proteins so that the selected genes or proteins have a stronger relevance to a specific subject. The device of the present invention contains a storage device, an input device and a processor. The storage device stores a data warehouse that contains a data about a collection of genes or proteins, with which annotations are associated. The input device receives an input of the set of candidate genes or proteins. The processor (a) gathers annotations that are associated with the candidate genes or proteins, (b) chooses annotations that are associated with the candidate genes or proteins more than a threshold number of times or frequencies, and (c) selects genes or proteins, with which at least one of the chosen annotations is associated.