Clean Dataset Generation from Noisy Customer Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
E-commerce entities face challenges in generating accurate insights and forecasts due to high-dimensional data being noisy, sparse, and fragmented, which affects the accuracy of historical trend analysis and future pattern predictions.
Innovation Solution
A computing system that generates clean datasets from high-dimensional noisy, sparse, and fragmented data by utilizing constraint data and customer profile data to score customer profiles based on their closeness to global distributions, thereby creating a representative subset of data that is less noisy and sparse.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If high-dimensional data is used for insights and forecasts, then the comprehensiveness of analysis is improved, but the accuracy deteriorates due to noise, sparsity, and fragmentation
Solution Approach 1:
The patent segments the high-dimensional data into multiple subsets, each satisfying specific constraints. By dividing the data into manageable segments that individually meet quality criteria, the system maintains comprehensiveness while improving accuracy of insights generated from each segment.
Solution Approach 2:
The patent introduces constraint data as an intermediary between the raw high-dimensional data and the insights generation process. This intermediary layer filters and structures the data to eliminate noise and sparsity while preserving the essential information needed for accurate forecasts.
2Quantity of substance
If high-dimensional noisy data is processed directly, then the data volume is maintained, but the data quality deteriorates
Solution Approach 1:
The patent extracts clean subsets of data from the high-dimensional noisy data by applying constraint-based filtering. This extraction process removes harmful elements (noise, sparsity, fragmentation) while retaining the essential data structures needed for reliable analysis.
Solution Approach 2:
The patent changes the parameters of data selection by introducing constraint satisfaction criteria. By modifying how data is selected and organized (from raw high-dimensional to constraint-satisfied subsets), the system improves data quality while maintaining sufficient data volume for analysis.
3Reliability
If constraint-based filtering is applied to generate clean subsets, then the data quality is improved, but the processing complexity increases
Solution Approach 1:
The patent segments the complex processing task into manageable steps: obtaining constraint data, generating subsets based on constraints, and verifying constraint satisfaction. This segmentation makes the overall complex process more tractable and implementable.
Solution Approach 2:
The system uses the constraint data to automatically guide the subset generation process without requiring manual intervention. The constraints themselves serve as the filtering mechanism, making the process self-regulating and reducing operational complexity.
Data Source
AI summary
In some examples, a system may, obtain constraint data and customer profile data of a plurality of customers associated with the system. Moreover, for each customer of the plurality of customers, the system may, based on the customer profile data of the customer and the constraint data, generate a score associated with one or more constraints of the plurality of constraints, based on the score of each of the one or more constraints, generate an overall score, and associate the overall score with a customer profile of the customer. Further, the system may, implement operations that generate a clean dataset based on the overall score associated with a customer profile of each of the plurality of customers.


