ML Pipeline for Data Quality in Software Recommendations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for identifying qualified leads in software recommendations, such as financial products, often result in poor target quality due to inefficiencies in integrating user data and generating personalized offer outputs, leading to high missed detections and false alarms, which in turn cause low user engagement.
Innovation Solution
An automated machine learning pipeline that integrates user data from various sources, automatically selects a suitable machine learning model based on user data properties, trains the model to predict product propensity likelihood, and uses it to define qualified customer targets, allowing for real-time adaptation and improved recommendation quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If business logic methods are used to identify qualified leads, then the process is simple and easy to implement, but the target quality is poor with high missed detections and false alarms
Solution Approach 1:
The patent replaces traditional business logic methods (mechanical rule-based systems) with machine learning models that automatically learn patterns from integrated user data. This substitution enables the system to identify qualified leads with higher accuracy while maintaining ease of implementation through automated model training and deployment.
Solution Approach 2:
The system changes the parameters used for lead identification from simple business rules to complex, multi-dimensional user data attributes. By integrating diverse data sources and using machine learning to analyze these parameters, the system achieves better target quality without compromising implementation simplicity.
2Device complexity
If traditional data integration methods are used, then the system architecture is simple, but the ability to efficiently integrate various user data is insufficient
Solution Approach 1:
The patent segments the data integration process into distinct layers: data collection from multiple sources, data preprocessing and cleaning, feature engineering, and model training. This segmentation allows each component to be optimized independently, improving overall integration efficiency while managing system complexity through modular architecture.
Solution Approach 2:
The system introduces intermediary components such as data preprocessing modules and feature engineering layers that mediate between raw user data and the machine learning models. These intermediaries efficiently transform and integrate data from various sources, improving productivity without significantly increasing perceived system complexity.
3Ease of operation
If personalized offer generation is implemented, then user engagement improves, but the complexity of generating offer outputs linked to user data increases
Solution Approach 1:
The system performs preliminary actions by pre-processing user data, pre-training machine learning models on historical data, and pre-generating offer templates before actual recommendation needs arise. This preliminary preparation reduces the complexity of real-time personalized offer generation while maintaining high user engagement through quickly delivered personalized recommendations.
Data Source
AI summary
A processor may receive user interaction data of a user for a plurality of electronically-presented offers. The processor may generate a plurality of labels, the generating comprising generating a label for each respective offer according to a comparison of the quality of the user interactions of the respective offer to the frequency of the user interactions of the respective offer. Each label may be a positive label or a negative label. The processor may determine whether the generating produced both positive and negative labels. The processor may select one of a plurality of available ML models, wherein a two-class ML model is chosen in response to determining that the generating produced both positive and negative labels and a one-class ML model is chosen in response to determining that the generating did not produce both positive and negative labels. The selected ML model may be trained and/or may be used to process user profile data and provide recommendations.


