ML Pipeline for Data Quality in Software Recommendations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for identifying qualified leads in software recommendations, such as financial products, often result in poor target quality due to inefficiencies in integrating user data and generating personalized offer outputs, leading to high missed detections and false alarms, which in turn cause low user engagement.

Innovation Solution

An automated machine learning pipeline that integrates user data from various sources, automatically selects a suitable machine learning model based on user data properties, trains the model to predict product propensity likelihood, and uses it to define qualified customer targets, allowing for real-time adaptation and improved recommendation quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If business logic methods are used to identify qualified leads, then the process is simple and easy to implement, but the target quality is poor with high missed detections and false alarms

Engineering Contradiction:
Improveease of implementationVSAvoidtarget quality
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent replaces traditional business logic methods (mechanical rule-based systems) with machine learning models that automatically learn patterns from integrated user data. This substitution enables the system to identify qualified leads with higher accuracy while maintaining ease of implementation through automated model training and deployment.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the parameters used for lead identification from simple business rules to complex, multi-dimensional user data attributes. By integrating diverse data sources and using machine learning to analyze these parameters, the system achieves better target quality without compromising implementation simplicity.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If traditional data integration methods are used, then the system architecture is simple, but the ability to efficiently integrate various user data is insufficient

Engineering Contradiction:
Improvesystem architecture complexityVSAvoiddata integration efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent segments the data integration process into distinct layers: data collection from multiple sources, data preprocessing and cleaning, feature engineering, and model training. This segmentation allows each component to be optimized independently, improving overall integration efficiency while managing system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces intermediary components such as data preprocessing modules and feature engineering layers that mediate between raw user data and the machine learning models. These intermediaries efficiently transform and integrate data from various sources, improving productivity without significantly increasing perceived system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If personalized offer generation is implemented, then user engagement improves, but the complexity of generating offer outputs linked to user data increases

Engineering Contradiction:
Improveuser engagementVSAvoidoffer generation complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing user data, pre-training machine learning models on historical data, and pre-generating offer templates before actual recommendation needs arise. This preliminary preparation reduces the complexity of real-time personalized offer generation while maintaining high user engagement through quickly delivered personalized recommendations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11861633B2Machine learning for improving mined data quality using integrated data sources
Publication Date: 2024.01.02 INTUIT INC
  • US11861633B2 patent drawing
  • US11861633B2 patent drawing
  • US11861633B2 patent drawing

AI summary

A processor may receive user interaction data of a user for a plurality of electronically-presented offers. The processor may generate a plurality of labels, the generating comprising generating a label for each respective offer according to a comparison of the quality of the user interactions of the respective offer to the frequency of the user interactions of the respective offer. Each label may be a positive label or a negative label. The processor may determine whether the generating produced both positive and negative labels. The processor may select one of a plurality of available ML models, wherein a two-class ML model is chosen in response to determining that the generating produced both positive and negative labels and a one-class ML model is chosen in response to determining that the generating did not produce both positive and negative labels. The selected ML model may be trained and/or may be used to process user profile data and provide recommendations.