Automated Genetic Variant Classification via NLP and Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge in clinical settings is the overwhelming amount of information regarding genetic variants, which complicates the interpretation of their pathogenicity and impact on clinical care decisions, due to the vast number of resources and continuously updating evidence.

Innovation Solution

A data processing system utilizing machine learning and natural language processing to automatically extract, interpret, and classify genetic variants from multiple sources, including publications and genomic databases, by identifying focal genetic variants, assessing their pathogenicity, and generating summaries of their findings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual interpretation of genetic variant information is performed, then accuracy of pathogenicity assessment can be maintained, but time consumption and workload increase significantly

Engineering Contradiction:
Improvepathogenicity assessment accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automated self-service processing of genetic variant information through machine learning models that automatically extract, classify, and assess pathogenicity of variants from electronic files, eliminating the need for manual interpretation while maintaining assessment accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual interpretation process with automated computational systems including NLP models, machine learning classifiers, and data processing algorithms that automatically analyze genetic variant data and generate pathogenicity assessments

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If comprehensive evidence from multiple sources is collected, then reliability of variant classification improves, but system complexity and data processing burden increase

Engineering Contradiction:
Improvevariant classification reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements a universal multi-functional platform that handles multiple types of electronic files (VCF, CSV, text formats), processes various genetic variant data types, and performs multiple analysis functions including extraction, classification, and pathogenicity assessment through integrated machine learning models

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces intermediary components including NLP models that mediate between raw electronic file data and structured variant information, and machine learning classifiers that serve as intermediaries between extracted features and final pathogenicity classifications, simplifying the overall system architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated processing is implemented, then productivity increases, but measurement precision of pathogenicity assessment may decrease

Engineering Contradiction:
Improveprocessing throughputVSAvoidpathogenicity assessment accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements feedback mechanisms where machine learning models are trained on labeled genetic variant data, with performance continuously evaluated and models retrained to improve accuracy, ensuring that automated processing maintains high measurement precision while achieving high productivity

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary actions by pre-training machine learning models on extensive labeled datasets before deployment, and by implementing preliminary data validation and quality checks on input electronic files to ensure accurate automated processing from the outset

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12014281B2Automatic processing of electronic files to identify genetic variants
Publication Date: 2024.06.18 MERATIVE US LP
  • US12014281B2 patent drawing
  • US12014281B2 patent drawing
  • US12014281B2 patent drawing

AI summary

A mechanism is provided for processing electronic files to identify genetic variants of a gene. Evidence of one or more genetic variants of the gene and corresponding information is extracted from a corpus of information. Each genetic variant of the one or more genetic variants is classified based on whether the genetic variant is identified as being pathogenic. Genetic variant annotation is then performed to generate a summary.