Transcriptomic Classification Confidence Scoring Under Sequencing Variance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Next-generation sequencing technologies introduce noise and artifacts that lead to indeterminate classification results, causing low-confidence or discordant patient grouping, which affects clinical decision-making and patient confidence.

Innovation Solution

A simulation-based approach using machine learning models to predict standard deviation and confidence scores for gene expression values, flagging samples with low confidence as potentially discordant to ensure accurate classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple sequencing runs are performed to reduce noise and artifacts, then measurement precision improves, but loss of time increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidsequencing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing simulations and generating confidence scores before actual clinical sequencing results are finalized. The system pre-calculates what the classification outcomes would be under various noise conditions, allowing clinicians to know in advance whether a sample is likely to produce indeterminate results. This prevents wasted sequencing runs on samples that would anyway yield discordant classifications.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If sequencing depth is increased to reduce variance, then measurement precision improves, but use of energy increases

Engineering Contradiction:
Improvegene expression measurement accuracyVSAvoidsequencing energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent changes the parameter from physical sequencing depth to computational confidence scoring. Instead of increasing sequencing depth to reduce variance, the system uses machine learning models to simulate various expression levels and calculate confidence scores. This computational approach achieves the same goal of assessing measurement reliability without the energy cost of additional sequencing.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If classification thresholds are made more stringent to improve accuracy, then measurement precision improves, but productivity decreases

Engineering Contradiction:
Improveclassification accuracyVSAvoidsample throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent introduces an intermediary confidence score between the raw sequencing data and the final classification. This confidence score acts as a mediator that filters samples before classification. Samples with low confidence scores are flagged for additional analysis or resequencing, while high confidence samples proceed directly to classification. This intermediary step maintains high accuracy without bottlenecking the overall workflow.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250285704A1Quantifying effects of sequencing variance on classification
Publication Date: 2025.09.11 TEMPUS AI INC
  • US20250285704A1 patent drawing
  • US20250285704A1 patent drawing
  • US20250285704A1 patent drawing

AI summary

A method includes receiving transcriptomic data; predicting standard deviation values for each gene expression values; for each sample: computing a simulated expression value, classifying the simulated expression values, computing confidence scores; and flagging the sample; and storing the confidence scores. A computing system includes a processor; and a memory having stored thereon instructions that when executed, cause the computing system to: receive transcriptomic data; predict standard deviation values for each gene expression values; for each sample: compute a simulated expression value, classify the simulated expression values, compute confidence scores; and flag the sample; and store the confidence scores. A computer-readable media includes non-transitory computer-readable instructions that, when executed, cause a computer to: receive transcriptomic data; predict standard deviation values for each gene expression values; for each sample: compute a simulated expression value, classify the simulated expression values, compute confidence scores; and flag the sample; and store the confidence scores.