Digital Avatar Recommendation via Reinforcement Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current recommendation systems are limited by their inability to actively engage users and understand preferences effectively, as they rely on passive content display and cannot accommodate a large quantity of content due to display limitations, leading to suboptimal user experience.
Innovation Solution
A digital avatar recommendation system using reinforcement learning to actively interact with users, mapping state data to target actions, and updating policies based on user feedback to provide personalized content recommendations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a batch of candidate content items are prepared and ranked based on user features, then recommendation accuracy is improved, but the quantity of content items that can be displayed is limited due to display location constraints
Solution Approach 1:
The patent transforms the static recommendation process into a dynamic interactive conversation. The recommendation system continuously adapts based on user responses, allowing the recommendation content to evolve dynamically rather than being fixed in advance. This enables the system to provide personalized recommendations through multi-turn dialogue, effectively increasing the usable recommendation capacity beyond static display limits.
Solution Approach 2:
The patent introduces a temporal dimension to the recommendation process by using multi-turn conversations. Instead of displaying all recommendations simultaneously in a single dimension (display space), the system distributes recommendations across multiple time steps and interaction turns, effectively adding a time dimension that expands the total recommendation capacity beyond spatial display constraints.
2Device complexity
If recommendation content is passively displayed and pushed to the user, then system complexity is reduced, but user engagement and preference understanding are insufficient
Solution Approach 1:
The patent inverts the traditional recommendation paradigm by shifting the initiative from the system to the user. Instead of the system actively pushing recommendations to users, the user initiates the interaction by asking questions or expressing needs, and the system responds accordingly. This inversion significantly boosts user engagement while maintaining manageable system complexity through template-based response generation.
Solution Approach 2:
The patent enables users to actively serve themselves in the recommendation process by allowing them to initiate conversations, ask questions, and guide the recommendation flow according to their own needs and preferences. This self-service approach empowers users to take control of their recommendation experience, significantly enhancing engagement without requiring complex system orchestration.
3Speed
If only a small quantity of content items are displayed on the page, then display performance is maintained, but most content items have no chance to be displayed and user preference understanding is limited
Solution Approach 1:
The patent implements continuous recommendation delivery through multi-turn conversations. Instead of delivering all recommendations at once or limiting to a small fixed set, the system continuously provides new recommendations based on ongoing user interactions. This continuous action ensures that users are exposed to a much larger variety of content items over time, reducing information loss about user preferences while maintaining excellent display performance at each moment.
Data Source
AI summary
Implementations of the present specification provide a digital avatar recommendation method and recommendation system. The digital avatar recommendation system includes a computer-simulated digital avatar, and the corresponding recommendation method includes: obtaining current state data, where the state data includes user information of a target user, scenario information of a current scenario, and history information of an interaction between the target user and the digital avatar; mapping, by an agent in the digital avatar, the state data to a target action in a candidate action set based on a current policy obtained through reinforcement learning, where a candidate action in the candidate action set corresponds to a to-be-recommended content category, and the target action corresponds to a target content category; and performing, by the digital avatar, target interaction with the target user, where the target interaction is used to recommend the target content category. As such, individualized recommendation is provided for the target user by using the digital avatar.


