Weighted Regression Modeling for Sparse New Product Categories
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing regression models in production systems face challenges in accurately estimating regression coefficients for new product categories with limited data, leading to unstable estimates due to noise and reduced data volume, especially when integrating coefficients across categories with varying data quantities.
Innovation Solution
A method that integrates regression coefficients using a weighted regularization term, where coefficients for new categories are aligned with those of established categories based on data quantity, and a range is set to limit data values, stabilizing estimates by actively integrating coefficients between categories with sufficient and insufficient data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If regression analysis is performed using limited data for new product categories, then model construction can proceed, but the regression coefficient estimates become unstable and noisy
Solution Approach 1:
The patent combines data from multiple product categories to construct regression models. By merging datasets across categories, the system increases the effective sample size for new categories, thereby stabilizing regression coefficient estimates and reducing noise while maintaining the ability to handle diverse product types.
Solution Approach 2:
The patent develops a universal regression modeling framework that can be applied across different product categories. The system uses category identifiers to enable the same modeling approach to work for both established and new categories, providing adaptability while maintaining estimation accuracy through shared statistical power.
2Adaptability or versatility
If data quantity varies significantly between categories, then existing methods can process each category separately, but integration of regression coefficients across categories becomes difficult and unreliable
Solution Approach 1:
The patent applies local quality by allowing category-specific regression coefficients while using a unified modeling framework. Each category can have its own estimated coefficients tailored to its specific data characteristics, while the overall methodology provides consistent integration across categories with varying data quantities, ensuring reliability through localized adaptation.
Solution Approach 2:
The patent changes parameters (category identifiers) to differentiate between various product categories while maintaining a consistent modeling approach. This allows the system to adapt to varying data quantities across categories by adjusting category-specific parameters within a unified framework, enabling reliable integration of regression coefficients despite heterogeneous data volumes.
Data Source
AI summary
An information processing device includes a processing unit. The processing unit calculates the number of pieces of data, which is the number of pieces of input data for each of a plurality of categories, by using n pieces of input data (n is an integer of 2 or more) each including a plurality of explanatory variables including a category variable representing any one of the plurality of categories. The processing unit calculates, for a plurality of combinations each including two of the categories included in the plurality of categories, a weight based on the number of pieces of data between two of the categories included in a combination. The processing unit learns a first regression model that estimates an objective variable from the plurality of explanatory variables by using a loss function including a regularization term in which a strength of regularization changes according to the weight.


