Contextual Bandits Model for Multi-Objective Ecosystem Personalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional reinforcement learning models for enterprise resource management face challenges such as high computational resource requirements, slow experimentation, and difficulty in managing multiple objectives, especially when departments have disparate goals, leading to inefficiencies and limited feedback for refining models.

Innovation Solution

A contextual bandits machine learning model that optimizes multiple enterprise objectives, including contradictory ones, by leveraging lifecycle models, propensity scores, and action-dependent-feature algorithms, and incorporates a rewards system to determine optimal next actions and channels for recommendations, enabling flexible and granular personalized customer interactions across various channels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional reinforcement learning models are used for enterprise resource management, then computational resources and data requirements are significant, but deployment becomes difficult and slow experimentation blocks productivity

Engineering Contradiction:
Improveexperimentation speedVSAvoidcomputational resource requirements
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the reinforcement learning framework into modular components: contextual bandits for multi-objective optimization, lifecycle models for customer behavior prediction, and propensity scores for action prioritization. This segmentation allows independent development and deployment of each component, enabling faster experimentation without requiring complete system reconfiguration.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-training lifecycle models and calculating propensity scores before actual recommendation deployment. These pre-computed assets enable rapid experimentation with different recommendation strategies without retraining entire models, thus blocking productivity less and allowing iterative improvements.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If conventional reinforcement learning models are used, then they can allocate resources to achieve enterprise objectives, but they struggle to manage multiple disparate departmental objectives simultaneously

Engineering Contradiction:
Improvemulti-objective optimization capabilityVSAvoidconfiguration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The contextual bandits framework provides universal multi-objective optimization capability that can handle disparate departmental objectives simultaneously. The system accepts multiple objectives (e.g., revenue, customer satisfaction, resource allocation) and automatically balances them through unified optimization, making the system adaptable to different enterprise needs without requiring separate configurations for each department.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system manages multiple objectives by dynamically changing optimization parameters and weights rather than creating separate models for each objective. By adjusting the objective function parameters and propensity score thresholds, the system can adapt to different departmental priorities while maintaining a single unified framework, reducing configuration complexity.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If conventional models are used, then they can provide recommendations, but feedback loops are limited and calculations suffer from static optimization functions based on limited data

Engineering Contradiction:
Improvemodel refinement capabilityVSAvoiddata quantity requirement
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system implements enhanced feedback loops where user interactions with recommendations are continuously fed back into the lifecycle models and contextual bandits. This feedback mechanism enables iterative model refinement, allowing the system to learn from actual user behavior and improve recommendation accuracy over time without requiring additional data collection infrastructure.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system transitions from static optimization functions to dynamic calculations that adapt to changing data conditions. The contextual bandits continuously update their policies based on new feedback, and the propensity scores are recalculated based on current user behavior patterns, enabling the system to handle varying data quantities and maintain optimal performance regardless of data availability.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11797891B1Contextual bandits-based ecosystem recommender system for synchronized personalization
Publication Date: 2023.10.24 INTUIT INC
  • US11797891B1 patent drawing
  • US11797891B1 patent drawing
  • US11797891B1 patent drawing

AI summary

The instant systems and methods are directed to a contextual bandits machine learning model configured to enable granular synchronized ecosystem personalization and optimization. The system and methods determine an objective and feed the objective and one more lifecycle model propensity scores as inputs to the contextual bandits machine learning model. The contextual bandits machine learning model then generates one or more potential weighted model rewards, wherein each potential weighted model reward includes at least a desired user action, a weight, a channel, and an expected change to the objective, and selects a weighted model reward that optimizes the objective. An action recommendation is subsequently transmitted to a user device based on the weighted model reward, wherein the action recommendation is presented in a selected channel associated with the weighted model reward. Feedback associated with the action recommendation is collected and used in training and fine-tuning of the model.