Machine Learning Model for Rare Disease Diagnosis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Rare diseases often take longer to diagnose due to unfamiliarity among medical practitioners, diverse symptoms, and masking by common diseases, leading to significant delays in treatment.

Innovation Solution

A computer-implemented method for generating a training dataset for a machine-learning model to diagnose rare diseases, involving unsupervised clustering to identify least representative clusters, pruning the dataset, and combining it with a control dataset to create a balanced training set.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional diagnostic algorithms are used, then healthcare providers can diagnose diseases, but the diagnosis time for rare diseases exceeds four years due to unfamiliarity and symptom diversity

Engineering Contradiction:
Improvediagnosis accuracyVSAvoiddiagnosis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary clustering and identification of representative symptoms before actual diagnosis. By pre-processing medical data to identify least representative clusters and pruning the dataset, the system prepares optimized training data in advance, enabling faster and more accurate diagnosis when the model is deployed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a machine-learning model that copies and learns from patterns in historical medical data. The model is trained on pruned datasets that capture essential disease patterns, allowing it to diagnose rare diseases without requiring healthcare providers to have extensive personal familiarity with each condition.

Inventive Principle:
Principle #26Copying

2Reliability

If the training dataset includes all individuals with rare diseases, then comprehensive coverage is achieved, but the model accuracy decreases due to inclusion of least representative cases

Engineering Contradiction:
Improvemodel accuracyVSAvoiddataset size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system extracts and removes least representative clusters from the training dataset through unsupervised clustering analysis. By identifying clusters that do not adequately represent the rare disease and removing them, the system improves model accuracy while maintaining an appropriate dataset size for effective training.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the composition parameters of the training dataset by pruning least representative cases. This parameter change in dataset composition improves the signal-to-noise ratio, enabling the machine-learning model to learn more accurate disease patterns without requiring proportionally larger datasets.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If healthcare providers verify numerous clinical characteristics including differential diagnoses, then diagnostic thoroughness is improved, but the complexity and time required increases significantly

Engineering Contradiction:
Improvediagnostic thoroughnessVSAvoiddiagnostic process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The machine-learning model performs self-service by automatically analyzing clinical characteristics and differential diagnoses without requiring manual verification by healthcare providers. The model has been trained to independently evaluate numerous clinical features and make diagnostic decisions, reducing the burden of complex manual verification while maintaining diagnostic thoroughness.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250118437A1Machine learning systems and methods to diagnose rare diseases
Publication Date: 2025.04.10 SANOFI SA(FR)
  • US20250118437A1 patent drawing
  • US20250118437A1 patent drawing
  • US20250118437A1 patent drawing

AI summary

A machine-learned model to diagnose patients with a rare disease based on medical data/records, and methods of training such a model are disclosed. A computer implemented method is disclosed for generating a training dataset for training a machine-learning model to identify individuals with a rare disease. The method comprises: receiving an initial dataset comprising medical data relating to a plurality of individuals with the rare disease; identifying a plurality of clusters of individuals in the initial dataset; identifying one or more of the clusters as being least representative of the rare disease; removing one or more of the individuals from the one or more clusters identified as being least representative based on the medical data of said one or more individuals to generate a pruned dataset; and combining the pruned dataset with a control dataset comprising a plurality of individuals without the rare disease to generate the training dataset.