Hybrid Recommendation Models for Long-Term Satisfaction in Large Action Spaces

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing web-scale recommendation systems focus on immediate user feedback, failing to enhance long-term user satisfaction by terminating user sessions once immediate needs are satisfied, and reinforcement learning at scale faces challenges due to large action spaces.

Innovation Solution

A distributed training system using supervised learning to warm-start an agent model, combined with a data generator and reinforcement learning trainer, employs a hybrid exploration strategy and transformer-based deep reinforcement learning to select items that increase user engagement and session duration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If reinforcement learning is applied to maximize long-term goals, then long-term user satisfaction is improved, but the dimensionality of RL does not scale due to extremely large action space

Engineering Contradiction:
Improvelong-term user satisfactionVSAvoiddimensionality of RL
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the recommendation system into multiple components: a supervised learning model that generates candidate items, and a reinforcement learning model that selects from these candidates. This segmentation reduces the action space for the RL model from billions of items to a manageable subset, resolving the scalability issue while maintaining long-term satisfaction optimization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The supervised learning model acts as an intermediary between the user and the reinforcement learning model. It pre-processes the vast item space and provides a curated candidate list to the RL model, enabling the RL model to focus on selection rather than search, thus reducing dimensionality while preserving long-term goal optimization.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If supervised learning methods are used to determine the best immediate result, then immediate user feedback is optimized, but long-term user satisfaction is not captured

Engineering Contradiction:
Improveimmediate user feedbackVSAvoidlong-term user satisfaction
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent merges supervised learning and reinforcement learning into a hybrid recommendation system. The supervised learning component handles immediate user feedback optimization, while the reinforcement learning component handles long-term satisfaction. Both components work together in a unified framework, allowing the system to simultaneously optimize for both immediate and long-term goals.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The hybrid system serves multiple functions: it optimizes for immediate user feedback through supervised learning, optimizes for long-term satisfaction through reinforcement learning, and generates candidate items through the supervised model. This multi-functionality allows a single system to address both short-term and long-term objectives without requiring separate systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250259101A1System and methods to recommend items that increase long-term user satisfaction
Publication Date: 2025.08.14 PINTEREST INC
  • US20250259101A1 patent drawing
  • US20250259101A1 patent drawing
  • US20250259101A1 patent drawing

AI summary

Described are systems and methods that resolve the shortcomings of existing recommendation systems that focus on immediate user reward rather than long-term user satisfaction. The disclosed implementations increase user session duration and user session depth during a session by selecting items that incent the user to remain engaged and participating in the session, rather than optimizing for immediate user metrics.