Machine Learning Classifiers for COVID-19 Phenotype Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods lack effective tools for accurately identifying specific phenotypes associated with COVID-19 disease states or conditions from diverse biological data sources, such as nucleic acid sequencing and gene expression data, which are crucial for diagnosis, prognosis, and therapeutic intervention.

Innovation Solution

A machine learning-based method employing classifiers like elastic generalized linear models, k-nearest neighbors, and random forest classifiers to analyze nucleic acid sequencing, transcriptome, and gene expression data, enabling the identification of specific phenotypes with high accuracy and sensitivity, even from non-overlapping data sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional analysis methods are used on diverse biological data sources, then the analysis process is simpler, but the accuracy of identifying specific phenotypes associated with COVID-19 disease states is insufficient

Engineering Contradiction:
Improveaccuracy of identifying specific phenotypesVSAvoidcomplexity of machine learning-based analysis system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex analysis task into distinct phases: data reception from multiple sources, application of machine learning algorithms to determine classifiers, and subsequent application of classifiers to identify phenotypes. This segmentation manages complexity by breaking down the overall system into manageable functional components that can be developed and validated independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces machine learning algorithms and classifiers as intermediary computational tools between the raw biological data and the final phenotype identification. These intermediaries process the diverse data sources (nucleic acid sequencing, transcriptome, gene expression) and transform them into actionable diagnostic information, bridging the gap between complex data and clinical interpretation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If diverse biological data sources are processed, then the comprehensiveness of COVID-19 phenotype identification is improved, but the data processing complexity and computational requirements increase

Engineering Contradiction:
Improveability to process diverse biological data sourcesVSAvoidcomplexity of data processing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal machine learning framework that can process multiple types of biological data sources including nucleic acid sequencing data, transcriptome data, and gene expression data. The same computational infrastructure and algorithmic approaches are applied across different data types, enabling the system to handle diverse inputs without requiring separate specialized processing pipelines for each data source.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If machine learning algorithms are applied to determine classifiers, then the sensitivity of phenotype identification is improved, but the computational time and resources required increase

Engineering Contradiction:
Improvesensitivity of phenotype identificationVSAvoidcomputational time for analysis
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies machine learning algorithms in advance to train and determine classifiers from initial datasets before these classifiers are deployed for actual phenotype identification. This preliminary training phase enables the system to quickly and accurately classify new data without requiring complex real-time computations during diagnostic applications, thereby reducing computational time while maintaining high sensitivity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230220470A1Methods and systems for analyzing targetable pathologic processes in covid-19 via gene expression analysis
Publication Date: 2023.07.13 AMPEL BIOSOLUTIONS LLC
  • US20230220470A1 patent drawing
  • US20230220470A1 patent drawing
  • US20230220470A1 patent drawing

AI summary

The present disclosure provides systems and methods for machine learning classification and assessment of COVID-19 disease based on gene expression data. In an aspect, a method for determining a COVID-19 disease state of a subject may comprise: (a) assaying a biological sample obtained or derived from the subject to produce a data set comprising gene expression measurements of the biological sample at each of a plurality of COVID-19 disease-associated genomic loci; (b) computer processing the data set to determine the COVID-19 disease state of the subject; and (c) electronically outputting a report indicative of the COVID-19 disease state of the subject.