Subsampling Flow Cytometry Data Using t-SNE Binning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current flow cytometry methods face challenges in efficiently subsampling high-dimensional event data, particularly in preserving rare cell populations and reducing data volume without discarding valuable information.
Innovation Solution
The method involves transforming flow cytometric event data from a higher-dimensional space to a lower-dimensional space using dimensionality reduction functions like t-Distributed Stochastic Neighbor Embedding (t-SNE), allowing for binning and selective subsampling based on predefined criteria, ensuring that rare events and cells of interest are adequately represented in the subsampled dataset.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If flow cytometric event data is subsampled to reduce data volume, then data processing efficiency is improved, but rare cell populations may be lost or underrepresented
Solution Approach 1:
The patent segments the high-dimensional flow cytometric data space into multiple lower-dimensional subspaces using dimensionality reduction functions. Each subspace captures different aspects of the data, allowing selective subsampling from each subspace while collectively preserving rare cell populations across all subspaces. This segmentation enables efficient processing of each subspace independently while maintaining overall data integrity.
Solution Approach 2:
The patent transforms data from a higher-dimensional space to a lower-dimensional space using dimensionality reduction functions like t-SNE. This dimensionality change allows the data to be projected into multiple lower-dimensional subspaces where rare events can be captured more effectively, enabling subsampling that preserves rare cell populations while reducing computational burden.
2Quantity of substance
If dimensionality reduction is applied to flow cytometric data, then data volume is reduced, but measurement precision may deteriorate
Solution Approach 1:
The patent divides the dimensionality reduction process into multiple segments, creating several lower-dimensional subspaces rather than a single reduced space. Each subspace maintains specific relationships and patterns from the original high-dimensional data, allowing the system to preserve measurement precision for different event characteristics while collectively reducing overall data volume.
Solution Approach 2:
The patent changes the dimensional parameters of the data by applying dimensionality reduction functions that transform high-dimensional event data into lower-dimensional representations. This parameter change reduces data volume while the multi-subspace approach ensures that critical measurement precision is maintained by distributing information across multiple reduced subspaces.
Data Source
AI summary
Disclosed herein include systems, devices, computer readable media, and methods for subsampling flow cytometric event data. First and second flow cytometric event data can be transformed into a lower-dimensional space, associated with a plurality of bins, and assigned to a first bin and a second bin. Subsampled flow cytometric event data comprising the first flow cytometric event data can be generated. The subsampled flow cytometric event data can comprise the second flow cytometric event data if the first bin and the second bin are different. The subsampled flow cytometric event data may not comprise the second flow cytometric event data if the first bin and the second bin are identical.


