Multidimensional Event Classification via Density Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for automatic classification of events in multidimensional spaces are inefficient, particularly when dealing with large datasets from biomedical analyses, as they require prior knowledge, are subjective, and struggle with non-Gaussian distributions or unknown populations, leading to incomplete analysis and inaccuracies.
Innovation Solution
A computer-implemented method that clusters events by determining event density, connecting closest neighbors, calculating affinity between groups, and comparing with reference groups using dimensionality reduction and statistical criteria to identify populations without requiring prior knowledge of group numbers or distributions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual analysis methods are used to classify events in multidimensional space, then analysis precision can be maintained for small datasets, but productivity decreases dramatically as the number of parameters increases
Solution Approach 1:
The patent replaces manual mechanical analysis with an automated computational system that uses algorithms to cluster and classify events in multidimensional space, eliminating the need for manual inspection of numerous parameter combinations while maintaining classification precision
Solution Approach 2:
The patent transforms the analysis approach by changing from manual parameter-by-parameter examination to automated multivariate analysis, where the system automatically adjusts and optimizes parameter combinations to identify populations without manual intervention
2Productivity
If automatic clustering methods based on finite mixture models are used, then productivity increases, but reliability decreases because the methods require prior knowledge of group numbers and thresholds
Solution Approach 1:
The patent implements a self-service classification system where the algorithm automatically determines the number of groups and thresholds by analyzing the data distribution itself, without requiring pre-programmed knowledge of population characteristics or manual setting of classification criteria
Solution Approach 2:
The patent introduces dynamic adaptation where the classification system automatically adjusts its parameters and thresholds based on the specific dataset being analyzed, allowing it to adapt to different population distributions and characteristics without requiring manual reconfiguration
3Productivity
If conventional automatic classification methods are used, then productivity improves, but measurement precision deteriorates when dealing with non-Gaussian distributions or unknown populations
Solution Approach 1:
The patent changes the mathematical approach from assuming Gaussian distributions to using non-parametric methods that can handle any data distribution, allowing accurate classification of unknown populations without requiring the data to follow specific statistical patterns
Solution Approach 2:
The patent segments the classification process into distinct stages: density estimation, cluster formation, and validation, where each stage independently handles different aspects of the data, allowing the system to process complex non-Gaussian distributions through systematic decomposition of the classification task
4Measurement precision
If manual step-by-step classification is performed for each population, then measurement precision is maintained, but device complexity increases as the number of steps increases with the number of parameters
Solution Approach 1:
The patent merges multiple manual classification steps into a single integrated automated algorithm that simultaneously handles multiple parameters and populations, consolidating what would otherwise require numerous separate analysis steps into one cohesive computational process
Data Source
Figure 1~2
Figure 3~4
Figure 5(a)~6
AI summary
The invention relates to a computer-implemented method for classifying events present in a sample, wherein each event is characterized by a multidimensional set of parameters obtained by means of hardware and/or software, wherein the values of the parameters associated with each event define the position coordinates of said event in a multidimensional space. Said method comprises the following stages: a) clustering the events in groups, b) checking if within each formed group there is a connection between events exceeding a maximum distance threshold, c) calculating the affinity between each pair of previously generated sample groups, d) comparing each sample group with at least one reference group stored in at least one database, e) classifying the sample groups based on the comparisons with the reference groups.