Computational Model for Rapid Disease Diagnosis Using Random Genomic Subsets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for diagnosing and classifying diseases, particularly brain tumors, are time-consuming and labor-intensive due to the need for genome-wide assessment of CpG methylation and prognostic molecular features, which cannot provide diagnostic feedback within the timeframe of a routine neurosurgical procedure, and existing classification methods do not support real-time interpretation of nanopore sequencing data.

Innovation Solution

A method involving the use of a computational model, preferably a linear classifier with independent feature sampling, that processes genetic and epigenetic information from a random subset of genomic positions independently, allowing for rapid diagnosis and classification by assigning samples to classes based on trained data from pre-classified samples, even with sparse and shallow coverage data from nanopore sequencing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If genome-wide assessment of CpG methylation is performed using conventional methods, then diagnostic accuracy is improved, but diagnostic time and labor requirements increase significantly

Engineering Contradiction:
Improvediagnostic accuracyVSAvoiddiagnostic time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the genome-wide assessment into two parts: (1) a random subset of genomic positions is sampled for rapid initial classification, and (2) only if needed, full genome-wide assessment is performed. This segmentation allows most diagnoses to be made quickly using only the random subset, while maintaining the option for comprehensive analysis when required.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by performing assessment on only a random subset of genomic positions rather than the complete genome-wide set. This partial assessment is sufficient for accurate classification in most cases, eliminating the need for time-consuming complete genome-wide analysis while maintaining diagnostic accuracy.

Inventive Principle:
Principle #16Partial or excessive action

2Adaptability or versatility

If conventional classification methods are used, then comprehensive disease classification is achieved, but real-time diagnostic feedback cannot be provided

Engineering Contradiction:
Improveclassification comprehensivenessVSAvoiddiagnostic speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent segments the classification process into a rapid first stage using a random subset of genomic positions with a simplified computational model, and a second stage using the complete genome-wide data if more comprehensive classification is needed. This enables real-time feedback in the first stage while preserving comprehensive classification capability for complex cases.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a dynamic classification approach where the level of analysis is adjusted based on the case requirements. The system starts with rapid classification using a random subset and can dynamically escalate to full genome-wide analysis when the initial classification is insufficient or when comprehensive assessment is clinically indicated.

Inventive Principle:
Principle #15Dynamics

3Productivity

If nanopore sequencing is used for rapid data acquisition, then diagnostic time is reduced, but data sparsity and shallow coverage reduce classification reliability

Engineering Contradiction:
Improvedata acquisition speedVSAvoidclassification reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent accepts partial data from the random subset of genomic positions sampled by nanopore sequencing and develops classification methods that are robust to this sparsity. By designing classifiers that can accurately distinguish sample classes even with limited data points, the patent achieves reliable classification without requiring complete genome-wide coverage.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes the parameters of the classification approach by developing computational models specifically optimized for sparse, shallow coverage data from random subsets. These models use different mathematical formulations and assumptions that are appropriate for the noisy, incomplete data regime produced by rapid nanopore sequencing, thereby maintaining reliability despite data sparsity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240379236A1Method for the diagnosis and/or classification of a disease in a subject
Publication Date: 2024.11.14 FREE UNIV OF BERLIN
  • US20240379236A1 patent drawing
  • US20240379236A1 patent drawing
  • US20240379236A1 patent drawing

AI summary

The present invention relates to a method for the diagnosis and/or classification of a disease in a subject based on the genetic and/or epigenetic information of a sample obtained from the subject, the method comprising the steps of: a) providing data from said sample, wherein said data comprises genetic and/or epigenetic information of a random subset of genomic positions: b) assigning said sample to a sample class based on genetic and/or epigenetic information of said random subset of genomic positions by employing a computational model, which discriminates a plurality of sample classes based on genetic and/or epigenetic information of a set of genomic positions comprising said random subset, wherein the computational model has been trained with pre-determined genetic and/or epigenetic information obtained from a plurality of pre-classified samples of known diseases and wherein said computational model processes the genetic and/or epigenetic information of a genomic position of said random subset independently of the genetic and/or epigenetic information of another genomic position of said random subset, wherein said computational model is preferably in the form of a linear classifier with independent feature sampling.