Dimension Usefulness Calculator for Machine Learning Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models for discrimination and regression analysis face overtraining issues due to unnecessary dimensions in data, particularly when the number of training samples is small, leading to poor performance in diseases like cancer diagnosis.

Innovation Solution

An analysis-data analyzing device and method that selects useful dimensions by calculating their degree of usefulness, stochastically invalidating less useful dimensions, and iteratively updating their importance, allowing for high-reliability discrimination or regression analysis even with limited training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all dimensions in multidimensional analysis data are used for machine learning, then the model can potentially capture all relevant information, but overtraining occurs due to unnecessary dimensions leading to poor discrimination performance

Engineering Contradiction:
Improvediscrimination performanceVSAvoidnumber of dimensions
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes unnecessary dimensions from the multidimensional analysis data. The machine learning model identifies dimensions that contribute little to discrimination performance and excludes them from the model construction, thereby preventing overtraining while maintaining reliable discrimination. This is achieved through dimension selection mechanisms that evaluate and filter out redundant features from the original high-dimensional dataset.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different treatment to different dimensions based on their individual quality and usefulness. Rather than uniformly using all dimensions or removing all dimensions, the system evaluates each dimension's contribution to discrimination and selectively retains only those dimensions that provide meaningful information. This local quality assessment allows the model to focus on high-value dimensions while ignoring noisy or redundant ones.

Inventive Principle:
Principle #3Local quality

2Reliability

If a large number of training samples are collected to prevent overtraining, then the discrimination performance improves, but the cost and time for data collection increase significantly

Engineering Contradiction:
Improvediscrimination performanceVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts the essential information needed for reliable discrimination from a limited dataset by identifying and focusing on the most informative dimensions. Rather than requiring large amounts of data to capture all relevant patterns, the system extracts key discriminative features from available data, reducing the need for extensive data collection while maintaining model reliability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of dimension selection to optimize model performance with limited data. By adjusting which dimensions are included in the model based on their discriminative power, the system adapts to work effectively with smaller sample sizes, eliminating the need to increase data quantity to achieve reliable performance.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If dimensions with small coefficients in discrimination functions are set to zero to prevent overtraining, then unnecessary dimensions are removed, but the accuracy of dimension selection becomes low when training data is insufficient

Engineering Contradiction:
Improvenumber of dimensionsVSAvoidaccuracy of dimension selection
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary evaluation of dimension importance before final model construction. By assessing the discriminative value of each dimension upfront using available training data, the system prepares a ranked list of dimensions that guides subsequent model building. This preliminary action allows for more accurate dimension selection even with limited data, as the evaluation framework is established before the main learning process begins.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where the model's performance on validation data is used to refine dimension selection. The system continuously evaluates whether retained dimensions actually improve discrimination performance and adjusts the selected dimensions accordingly. This feedback loop enhances the accuracy of dimension selection by verifying choices against actual model performance rather than relying solely on initial coefficient estimates.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11341404B2Analysis-data analyzing device and analysis-data analyzing method that calculates or updates a degree of usefulness of each dimension of an input in a machine-learning model
Publication Date: 2022.05.24 SHIMADZU CORP
  • US11341404B2 patent drawing
  • US11341404B2 patent drawing
  • US11341404B2 patent drawing

AI summary

Using training data, machine learning is performed to construct a learning model which is a non-linear function for discrimination or regression analysis (S2). A degree of contribution of each input dimension is calculated from a partial differential value of that function. Input dimensions to be invalidated are determined using a threshold defined by a Gaussian distribution function based on the degrees of contribution (S3-S5). Machine learning is once more performed using the training data with the partially-invalidated input dimensions (S6). A new value of the degree of contribution of each input dimension is determined from the obtained learning model, and the degree of contribution is updated using the old and new values (S7-S8). After the processes of Steps S5-S8 are iterated a specified number of times (S9), useful dimensions are determined based on the finally obtained degrees of contribution, and the machine-learning model is constructed (S10).