Genetic Sequence Classification via Machine Learning Causal Feature Sets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for classifying genetic sequences associated with gene expression patterns are limited in accuracy and require experimental data, making it difficult to effectively identify and manipulate circadian and non-circadian gene expression profiles.

Innovation Solution

A machine learning-based approach that utilizes time-series transcriptome data and k-mer analysis to classify genetic sequences, employing a trained model to distinguish between circadian and non-circadian patterns without experimental data, and allows for the identification of causal features that can be used for gene editing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If experimental genetic expression data and prior knowledge of regulatory elements are used for classification, then measurement precision is improved, but device complexity and loss of time increase due to required experimental procedures

Engineering Contradiction:
Improveclassification accuracyVSAvoidexperimental procedure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses computational models to create virtual copies of experimental classification processes. Instead of performing actual experiments, the system uses machine learning models trained on existing data to predict gene expression patterns, thereby avoiding the complexity and time requirements of wet-lab experiments while maintaining classification accuracy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary classification and feature identification through computational analysis before any experimental validation. By using trained machine learning models to pre-identify causal features and predict expression patterns, the system reduces the need for extensive experimental procedures, saving time and reducing operational complexity

Inventive Principle:
Principle #10Preliminary action

2Reliability

If experimental data collection methods are used, then reliability of genetic expression profiles is improved, but loss of time and productivity decrease due to lengthy experimental procedures

Engineering Contradiction:
Improveexpression profile accuracyVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system creates computational replicas of experimental data collection processes through machine learning models. These models have been trained on existing experimental data and can rapidly generate predictions without requiring actual experiments, thereby maintaining reliability while dramatically reducing time loss

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces mechanical/wet-lab experimental systems with computational information processing systems. Instead of physically collecting experimental data through laboratory procedures, the system uses algorithmic processing of sequence data to predict expression profiles, eliminating the time-consuming experimental phase while maintaining predictive reliability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If comprehensive feature sets are used for classification, then measurement precision is improved, but device complexity and computational requirements increase

Engineering Contradiction:
Improvefeature classification accuracyVSAvoidfeature set complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and identifies only the most relevant causal features from comprehensive feature sets using machine learning analysis. Instead of using all possible features for classification, the system identifies and focuses on the specific features that actually drive gene expression patterns, reducing complexity while maintaining or improving classification precision

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different levels of feature analysis to different regions or types of genetic sequences. Rather than uniformly applying comprehensive feature sets to all sequences, the system adapts the feature analysis to local characteristics of the genetic data, improving precision where needed while reducing unnecessary complexity in other regions

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20220156632A1Identifying genetic sequence expression profiles according to classification feature sets
Publication Date: 2022.05.19 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20220156632A1 patent drawing
  • US20220156632A1 patent drawing
  • US20220156632A1 patent drawing

AI summary

Classifying genetic sequences by receiving genetic sequence data according to sequence features associated with gene expression, determining a genetic sequence feature set, determining a first classification for the genetic sequence feature set according to a machine learning model, defining a causal feature set associated with the first classification for the genetic sequence according to the machine learning model, altering the causal feature set for the genetic sequence, yielding an altered causal feature set, determining a second classification for the altered causal feature set according to the machine learning model, wherein the second classification differs from the first classification, and defining a set of target features, wherein the target features include causal features the altered causal feature set.