ML Sample Classification with UMAP Dimensionality Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning-based classification models for physical samples, such as cellular and biological materials, face challenges due to variations in sensor data over time and across different subjects, leading to reduced performance and incorrect classification, especially when molecular markers are not specific or sufficient for training.
Innovation Solution
A system and method that includes a flow cytometer and processors to retrieve sensor data, apply it to a classifier configured based on feature data from different examples, and use dimensionality reduction techniques like UMAP to set gates for predicting classes, allowing for classification without artificial labels and controlling for data variations, enabling accurate classification of target objects even without specific biomarkers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If classification models are trained using traditional methods with molecular markers, then classification can be performed, but performance is limited by data variations over time and across subjects
Solution Approach 1:
The patent transforms raw sensor data into standardized feature representations that are invariant to temporal and subject-specific variations. By changing the parameter representation from raw sensor readings to normalized feature vectors, the system maintains classification reliability across different conditions while eliminating the negative impact of data variations.
Solution Approach 2:
The patent introduces an intermediary feature extraction layer between raw sensor data and classification. This intermediary process standardizes diverse inputs into a common feature space, acting as a mediator that preserves essential information while filtering out variations related to time and subject identity.
2Reliability
If dimensionality reduction techniques are applied to control data variations, then classification performance improves, but computational complexity increases
Solution Approach 1:
The patent segments the complex classification task into distinct stages: feature extraction, dimensionality reduction, and classification. By dividing the process into manageable segments, each with a specific function, the system achieves improved classification performance while keeping computational complexity controlled through efficient algorithms at each stage.
Solution Approach 2:
The patent applies dimensionality reduction techniques to transform high-dimensional sensor data into a lower-dimensional feature space. This dimensional transformation preserves the essential classification information while reducing computational complexity, allowing the system to handle variations effectively without excessive computational burden.
3Measurement precision
If multiple feature extraction criteria are applied to ensure reproducibility and differentiation, then classification accuracy improves, but processing time increases
Solution Approach 1:
The patent performs preliminary feature extraction and selection before the actual classification process. By pre-processing the data to extract only the most relevant features that satisfy reproducibility and differentiation criteria, the system achieves high measurement precision while minimizing processing time during the classification phase.
Solution Approach 2:
The patent applies multiple feature extraction criteria but selectively uses only the most critical features for classification. This partial application of exhaustive feature extraction ensures that the system achieves sufficient precision without the full computational cost of analyzing every possible feature, thus reducing processing time while maintaining accuracy.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The approach enhances the performance of classification models by controlling for data variations, allowing for accurate prediction of cell types and classification of samples with high throughput, even when molecular markers are not available, and improves the ability to detect unknown classifications.
Implementation Method 1
a flow cytometer configured to direct a fluid flow including an object through a field of view of a photosensor/photodetector, and to cause the photosensor/photodetector to detect sensor data regarding the object
Implementation Method 2
illuminating observed objects by light with an illumination pattern over a time period. The method can include receiving, by a sensor, electromagnetic waves irradiated from the observed objects illuminated by the illumination pattern. The method can include converting the electromagnetic waves into time-series electrical signals
Data Source
AI summary
Systems and methods are provided to implement classification of objects, based on sensor data regarding the objects, in a manner that addresses variations in the sensor data, including measurement variables among the objects. A system can include one or more processors to retrieve sensor data regarding an object that is at least one of cellular material from one or more cells, nucleic acid material, biological material, or chemical material. The one or more processors can apply the sensor data as input to a classifier to cause the classifier to determine a classification of the object, the classifier configured based on feature data from a first example of object data and a second example of object data associated with at least one of a different time of detection or a different subject than the first example.


