Data Analysis Method for Locating Entities in Multivariable Datasets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for analyzing large biological datasets are ineffective in identifying the subset of data most relevant for further analysis and distinguishing between experimental groups, often providing too much or too little information, and fail to discover new molecular entities involved in biological phenomena.

Innovation Solution

The method involves obtaining data points for a biological phenomenon, designating a trend associated with it, developing a mathematical model, and testing each data point for adherence to the model to identify subsets of interest, which can include biochemical, gene expression, or protein expression profiling data, using techniques like stepwise discriminant analysis to pinpoint biomarkers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If existing statistical methods are used to analyze large biological datasets, then comprehensive information is obtained, but the information is too extensive for users to effectively identify the most relevant subsets

Engineering Contradiction:
Improveinformation completenessVSAvoiduser ability to identify relevant data
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent extracts only the most relevant data subsets from large biological datasets by applying mathematical models and trend analysis. Instead of presenting all statistical results, the system identifies and extracts specific entities (genes, compounds, proteins) that exhibit trends of interest, thereby reducing information overload while maintaining relevance.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments large datasets into meaningful subsets based on identified trends and patterns. By dividing the comprehensive data into distinct groups of entities that share common characteristics or behaviors, the system makes the data more manageable and easier for users to analyze without losing important information.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If discriminant analysis is used to choose measurements that distinguish between experimental groups, then differentiation is achieved, but too much information is provided to the user

Engineering Contradiction:
Improveability to distinguish between groupsVSAvoidamount of information provided
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the critical measurements and entities that are most important for distinguishing between experimental groups. By applying trend analysis and mathematical models, the system identifies a focused subset of data points that provide the necessary discriminatory power without presenting the full breadth of discriminant analysis results.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If manual analysis is used on complex datasets, then detailed examination is possible, but the process becomes impractical or impossible due to data volume

Engineering Contradiction:
Improvedetail of analysisVSAvoidanalysis speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual mechanical analysis with automated computational methods. Mathematical models and algorithms automatically process large datasets, identify trends, and locate entities of interest, thereby maintaining detailed analysis capabilities while dramatically increasing productivity and eliminating the impracticality of manual examination.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically analyzing datasets, identifying trends, and locating relevant entities without requiring manual intervention. The mathematical models autonomously process the data and generate results, enabling detailed analysis at scale without human effort.

Inventive Principle:
Principle #25Self-service

4Difficulty of detecting and measuring

If prior knowledge of treatment mechanism is used to look for expected perturbations, then targeted search is achieved, but new undiscovered entities are missed

Engineering Contradiction:
Improvetargeted search efficiencyVSAvoidability to discover new entities
Core Design Contradiction:
Difficulty of detecting and measuringVSAdaptability or versatility

Solution Approach 1:

The patent employs dynamic trend analysis that can adapt to both expected and unexpected patterns in the data. Rather than being constrained by fixed prior knowledge, the mathematical models identify emerging trends and anomalies, enabling the system to maintain targeted search efficiency while simultaneously discovering new entities that deviate from expected patterns.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8131473B1Data analysis methods for locating entities of interest within large, multivariable datasets
Publication Date: 2012.03.06 METABOLON INC
  • US8131473B1 patent drawing
  • US8131473B1 patent drawing
  • US8131473B1 patent drawing

AI summary

The present invention provides data analysis methods for the rapid location of subsets of large, multivariable biological datasets that are of most interest for further analysis, for the investigation of molecular modes of action of biological phenomena of interest, and for the identification of sets of data points that best distinguish between experimental groups in larger datasets as putative biomarkers. While existing methods for analyzing large biological datasets generally provide too much information to the user, or not enough, the methods of the present invention entail taking user input on what kinds of trends are of interest and then finding results that match the designated trend. In such manner, the methods of the invention allow a user to quickly pinpoint the subset of data of most interest without a concomitant loss of a large percentage of relevant information, as is typical with standard methods. The methods of the invention allow for identification of molecular entities that are involved in a biological phenomenon of interest, entities that may have otherwise gone undiscovered in a large, multivariable dataset.