Automated Learning Machine Selection for Gene Expression Pattern Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning techniques, such as support vector machines (SVMs), are limited in their ability to efficiently recognize patterns and categorize data, especially in large, varied databases, and require mathematical expertise, making them difficult for non-specialists to use effectively in biomedical research, particularly for querying gene expression data.

Innovation Solution

The development of methods and systems that optimize the selection and training of learning machines using performance functions like divergence, n-fold cross-validation, and feature reduction to identify the most effective machine for pattern recognition and classification, allowing for automated selection and use by researchers without extensive mathematical knowledge.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If support vector machines and traditional machine learning techniques are used for pattern recognition, then classification accuracy can be achieved, but the system becomes difficult to operate for non-specialists and requires extensive mathematical expertise

Engineering Contradiction:
Improvepattern recognition accuracyVSAvoidease of use for non-specialists
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces automated machine selection and configuration systems that act as intermediaries between the user and complex machine learning algorithms. The system automatically selects appropriate learning machines, configures their parameters, and optimizes performance without requiring users to understand underlying mathematical concepts, thereby maintaining high accuracy while improving ease of operation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables self-service by automatically performing machine selection, parameter optimization, and performance evaluation without human intervention. The automated processes evaluate multiple learning machines, select the best performing one based on cross-validation, and configure optimal parameters, allowing non-specialists to use sophisticated pattern recognition tools without mathematical expertise

Inventive Principle:
Principle #25Self-service

2Reliability

If multiple learning machines are trained and evaluated using cross-validation, then the selection of the most effective machine can be optimized, but the computational time and resources required increase

Engineering Contradiction:
Improvemachine selection reliabilityVSAvoidcomputational time for training and evaluation
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing n-fold cross-validation and performance evaluation during the machine selection phase before actual pattern recognition tasks. By pre-evaluating multiple learning machines and selecting the best one in advance, the system ensures reliable machine selection while avoiding the need for repeated evaluation during operational use, thus reducing overall computational time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses partial action by evaluating a limited but sufficient number of learning machines and using n-fold cross-validation with optimized fold numbers. This approach provides reliable machine selection without exhaustively testing every possible configuration, balancing selection reliability with computational efficiency

Inventive Principle:
Principle #16Partial or excessive action

3Device complexity

If feature reduction techniques are applied to large databases, then the complexity of pattern recognition can be reduced, but the amount of information available for analysis decreases

Engineering Contradiction:
Improvedata processing complexityVSAvoidinformation available for analysis
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent extracts only the most relevant and informative features from large databases using automated feature selection techniques. The system identifies and extracts key features that contribute most to pattern recognition accuracy while removing redundant or less informative features, thereby reducing processing complexity without significant loss of analytical information

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies local quality by treating different features differently based on their importance and contribution to pattern recognition. Rather than uniformly reducing all features, the system identifies and retains high-value features while reducing or removing less important ones, maintaining information quality in critical areas while reducing overall complexity

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10402748B2Machine learning methods and systems for identifying patterns in data
Publication Date: 2019.09.03 VIRKAR HEMANT V
  • US10402748B2 patent drawing
  • US10402748B2 patent drawing
  • US10402748B2 patent drawing

AI summary

Methods for training machines to categorize data, and/or recognize patterns in data, and machines and systems so trained. More specifically, variations of the invention relates to methods for training machines that include providing one or more training data samples encompassing one or more data classes, identifying patterns in the one or more training data samples, providing one or more data samples representing one or more unknown classes of data, identifying patterns in the one or more of the data samples of unknown class(es), and predicting one or more classes to which the data samples of unknown class(es) belong by comparing patterns identified in said one or more data samples of unknown class with patterns identified in said one or more training data samples. Also provided are tools, systems, and devices, such as support vector machines (SVMs) and other methods and features, software implementing the methods and features, and computers or other processing devices incorporating and/or running the software, where the methods and features, software, and processors utilize specialized methods to analyze data.