Automated Learning Machine Selection for Gene Expression Pattern Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning techniques, such as support vector machines (SVMs), are limited in their ability to efficiently recognize patterns and categorize data, especially in large, varied databases, and require mathematical expertise, making them difficult for non-specialists to use effectively in biomedical research, particularly for querying gene expression data.
Innovation Solution
The development of methods and systems that optimize the selection and training of learning machines using performance functions like divergence, n-fold cross-validation, and feature reduction to identify the most effective machine for pattern recognition and classification, allowing for automated selection and use by researchers without extensive mathematical knowledge.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If support vector machines and traditional machine learning techniques are used for pattern recognition, then classification accuracy can be achieved, but the system becomes difficult to operate for non-specialists and requires extensive mathematical expertise
Solution Approach 1:
The patent introduces automated machine selection and configuration systems that act as intermediaries between the user and complex machine learning algorithms. The system automatically selects appropriate learning machines, configures their parameters, and optimizes performance without requiring users to understand underlying mathematical concepts, thereby maintaining high accuracy while improving ease of operation
Solution Approach 2:
The system enables self-service by automatically performing machine selection, parameter optimization, and performance evaluation without human intervention. The automated processes evaluate multiple learning machines, select the best performing one based on cross-validation, and configure optimal parameters, allowing non-specialists to use sophisticated pattern recognition tools without mathematical expertise
2Reliability
If multiple learning machines are trained and evaluated using cross-validation, then the selection of the most effective machine can be optimized, but the computational time and resources required increase
Solution Approach 1:
The patent applies preliminary action by performing n-fold cross-validation and performance evaluation during the machine selection phase before actual pattern recognition tasks. By pre-evaluating multiple learning machines and selecting the best one in advance, the system ensures reliable machine selection while avoiding the need for repeated evaluation during operational use, thus reducing overall computational time
Solution Approach 2:
The system uses partial action by evaluating a limited but sufficient number of learning machines and using n-fold cross-validation with optimized fold numbers. This approach provides reliable machine selection without exhaustively testing every possible configuration, balancing selection reliability with computational efficiency
3Device complexity
If feature reduction techniques are applied to large databases, then the complexity of pattern recognition can be reduced, but the amount of information available for analysis decreases
Solution Approach 1:
The patent extracts only the most relevant and informative features from large databases using automated feature selection techniques. The system identifies and extracts key features that contribute most to pattern recognition accuracy while removing redundant or less informative features, thereby reducing processing complexity without significant loss of analytical information
Solution Approach 2:
The system applies local quality by treating different features differently based on their importance and contribution to pattern recognition. Rather than uniformly reducing all features, the system identifies and retains high-value features while reducing or removing less important ones, maintaining information quality in critical areas while reducing overall complexity
Data Source
AI summary
Methods for training machines to categorize data, and/or recognize patterns in data, and machines and systems so trained. More specifically, variations of the invention relates to methods for training machines that include providing one or more training data samples encompassing one or more data classes, identifying patterns in the one or more training data samples, providing one or more data samples representing one or more unknown classes of data, identifying patterns in the one or more of the data samples of unknown class(es), and predicting one or more classes to which the data samples of unknown class(es) belong by comparing patterns identified in said one or more data samples of unknown class with patterns identified in said one or more training data samples. Also provided are tools, systems, and devices, such as support vector machines (SVMs) and other methods and features, software implementing the methods and features, and computers or other processing devices incorporating and/or running the software, where the methods and features, software, and processors utilize specialized methods to analyze data.


