Sparse Estimation for Symbolic Regression Model Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional techniques for modeling physical phenomena using symbolic regression face challenges in improving accuracy due to issues with normalization of variables, leading to loss of physical meaning and improper operation of sparse estimation techniques when basis function ranges differ.
Innovation Solution
An information processing device employing a new sparse estimation technique that selects basis functions based on the magnitude of terms affecting the left-hand side of equations, rather than coefficients, and uses non-negative least squares to estimate coefficients, allowing for unnormalized variables and improving model generation accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional symbolic regression is used to model physical phenomena, then a mathematical model can be obtained from time-series data, but the accuracy of model generation cannot be further improved due to normalization issues
Solution Approach 1:
The patent changes the parameter selection criterion from coefficient magnitude to term magnitude (coefficient × basis function value). This allows the system to work with unnormalized variables directly, eliminating the normalization step while improving model accuracy by selecting basis functions based on their actual impact on the equation rather than their scaled coefficients
Solution Approach 2:
Instead of normalizing variables first and then selecting basis functions based on normalized coefficients, the patent inverts the approach: it selects basis functions based on the magnitude of their terms (coefficient × basis function value) using unnormalized variables, thereby eliminating the need for normalization while maintaining physical meaning
2Adaptability or versatility
If sparse estimation techniques are applied with normalized variables, then basis functions can be selected, but physical meaning is lost when variables have different ranges
Solution Approach 1:
The patent changes the selection parameter from normalized coefficient to term magnitude (coefficient × basis function value). This allows basis functions to be selected based on their actual contribution to the equation while working with unnormalized variables that retain their physical meaning and original units
Solution Approach 2:
The patent uses the magnitude of the term (coefficient × basis function value) as a proxy for importance rather than relying on normalized coefficients. This copying approach preserves the physical meaning of unnormalized variables while still enabling effective basis function selection through the term magnitude metric
3Ease of operation
If normalization is applied to variables with different ranges, then sparse estimation can operate, but the operation becomes improper and accuracy decreases
Solution Approach 1:
The patent changes the operation parameter from normalized coefficient to term magnitude. This allows sparse estimation to operate properly on unnormalized variables by selecting basis functions based on the actual impact (coefficient × basis function value) rather than scaled coefficients, thereby improving accuracy while maintaining ease of operation
Solution Approach 2:
The patent inverts the conventional approach by not normalizing variables at all. Instead, it directly uses unnormalized variables and selects basis functions based on term magnitude, eliminating the improper operations that arise from normalizing variables with different physical meanings and units
Data Source
AI summary
According to an embodiment, an information processing device of an embodiment includes a memory and one or more processors coupled to the memory. The memory stores therein time-series data including at least one of a dependent variable and an independent variable. The one or more processors are configured to: generate a nonlinear function based on at least one of the dependent variable and the independent variable; generate a linear regression equation in which the nonlinear function is a basis function; estimate a coefficient of the linear regression equation; calculate a product of the coefficient and a maximum value of the basis function corresponding to the coefficient, as a degree of influence; correct the coefficient based on the degree of influence; and output the linear regression equation expressed by the corrected coefficient.


