Generative Machine Learning Models for Computational Testing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computational algorithms in personalized medicine rely heavily on expensive biologically derived datasets for testing and validation, which are not always readily available and can introduce biases, making it difficult to ensure consistent performance across diverse patient populations.

Innovation Solution

Utilize artificially generated datasets created by trained generative machine learning models, such as large and small language models, to test the performance of computational algorithms, including genomic and epigenetic data generation, and evaluate the output against predetermined criteria.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If biologically derived datasets are used for testing computational algorithms, then the datasets provide real-world biological accuracy, but the cost and availability become problematic

Engineering Contradiction:
Improvebiological accuracyVSAvoidcost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent creates synthetic biological datasets that copy the essential statistical properties and patterns of real biological data without using actual biological samples. Generative models produce artificial genomic sequences, protein structures, and clinical data that replicate the complexity and variability of real-world biological data, enabling algorithm testing at minimal cost while maintaining biological realism.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces expensive, scarce real biological datasets with inexpensive synthetic data that can be generated on-demand in unlimited quantities. These artificial datasets serve as disposable testing resources that can be created, used, and regenerated without the constraints of biological sample availability, storage, or ethical considerations.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Measurement precision

If biologically derived datasets are used for testing, then real-world biological accuracy is achieved, but biases and variability issues arise

Engineering Contradiction:
Improvebiological accuracyVSAvoidconsistency across populations
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent uses generative models to systematically vary parameters in synthetic biological data, such as mutation rates, sequence compositions, and clinical outcomes, to create diverse test scenarios. This allows controlled exploration of how algorithms perform across different population characteristics without being constrained by the fixed biases present in real-world datasets.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

Instead of accepting the biases inherent in real biological data, the patent inverts the approach by deliberately designing synthetic data with controlled characteristics. The generative models can create datasets that either replicate specific bias patterns or deliberately eliminate them, allowing researchers to test algorithm robustness against various population variations in a controlled manner.

Inventive Principle:
Principle #13The other way round (Inversion)

3Measurement precision

If manual dataset curation is performed, then data quality is improved, but labor and time requirements increase

Engineering Contradiction:
Improvedata qualityVSAvoidcuration time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service data generation through automated generative models that produce high-quality synthetic biological datasets without human intervention. The models automatically generate realistic biological data with proper statistical properties, annotations, and variations, eliminating the need for manual data collection, cleaning, validation, and curation processes that traditionally require extensive researcher time and effort.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary data generation and validation through generative models before actual algorithm testing begins. The models pre-generate large volumes of synthetic datasets with known ground truths and controlled characteristics, allowing researchers to skip the time-consuming manual curation step and directly proceed to algorithm evaluation with ready-to-use, high-quality test data.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If diverse patient population data is collected, then generalizability is improved, but data availability and quality become problematic

Engineering Contradiction:
ImprovegeneralizabilityVSAvoiddata availability
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent creates a universal synthetic data generation system that can produce diverse patient population data across multiple disease types, genetic backgrounds, and demographic characteristics from a single generative model framework. This multi-functional approach allows the same system to generate test data for various algorithm applications, ensuring broad generalizability testing capability without requiring separate data collection efforts for each population type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250336491A1Machine learning models to test computational algorithms
Publication Date: 2025.10.30 GUARDANT HEALTH INC
  • US20250336491A1 patent drawing
  • US20250336491A1 patent drawing
  • US20250336491A1 patent drawing

AI summary

Methods and systems for testing the performance of computational algorithms to avoid relying on manually curated datasets or depending on expensive biologically derived sequencing datasets with known outcomes. These methods produce ample artificial datasets for faster more efficient software testing pipelines. Generative machine learning models can be implemented to generate the artificial datasets used for computational algorithm testing and evaluation.