Hybrid Recommendation Models for Long-Term Satisfaction in Large Action Spaces
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing web-scale recommendation systems focus on immediate user feedback, failing to enhance long-term user satisfaction by terminating user sessions once immediate needs are satisfied, and reinforcement learning at scale faces challenges due to large action spaces.
Innovation Solution
A distributed training system using supervised learning to warm-start an agent model, combined with a data generator and reinforcement learning trainer, employs a hybrid exploration strategy and transformer-based deep reinforcement learning to select items that increase user engagement and session duration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If reinforcement learning is applied to maximize long-term goals, then long-term user satisfaction is improved, but the dimensionality of RL does not scale due to extremely large action space
Solution Approach 1:
The patent segments the recommendation system into multiple components: a supervised learning model that generates candidate items, and a reinforcement learning model that selects from these candidates. This segmentation reduces the action space for the RL model from billions of items to a manageable subset, resolving the scalability issue while maintaining long-term satisfaction optimization.
Solution Approach 2:
The supervised learning model acts as an intermediary between the user and the reinforcement learning model. It pre-processes the vast item space and provides a curated candidate list to the RL model, enabling the RL model to focus on selection rather than search, thus reducing dimensionality while preserving long-term goal optimization.
2Productivity
If supervised learning methods are used to determine the best immediate result, then immediate user feedback is optimized, but long-term user satisfaction is not captured
Solution Approach 1:
The patent merges supervised learning and reinforcement learning into a hybrid recommendation system. The supervised learning component handles immediate user feedback optimization, while the reinforcement learning component handles long-term satisfaction. Both components work together in a unified framework, allowing the system to simultaneously optimize for both immediate and long-term goals.
Solution Approach 2:
The hybrid system serves multiple functions: it optimizes for immediate user feedback through supervised learning, optimizes for long-term satisfaction through reinforcement learning, and generates candidate items through the supervised model. This multi-functionality allows a single system to address both short-term and long-term objectives without requiring separate systems.
Data Source
AI summary
Described are systems and methods that resolve the shortcomings of existing recommendation systems that focus on immediate user reward rather than long-term user satisfaction. The disclosed implementations increase user session duration and user session depth during a session by selecting items that incent the user to remain engaged and participating in the session, rather than optimizing for immediate user metrics.


