Simulated-User Reinforcement Learning for Personalized Recommendations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The lack of data poses a significant challenge in reinforcement learning, hindering effective training and application of recommendation systems.

Innovation Solution

A method involving the generation of simulated users based on user data, utilizing pretrained generative artificial intelligence to create synthetic data for training, and updating states and rewards to optimize recommendation sessions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If reinforcement learning is applied to recommendation systems, then personalized recommendations can be provided, but data scarcity hinders effective training

Engineering Contradiction:
Improvepersonalized recommendation capabilityVSAvoidtraining data quantity
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent creates simulated users that copy and simulate the behavior patterns, preferences, and interactions of real users. These synthetic user profiles serve as virtual training data, allowing the reinforcement learning model to train without requiring extensive real user data. The simulated users replicate user states, actions, and rewards, providing a data-rich environment for training while preserving privacy and reducing dependency on actual user data collection.

Inventive Principle:
Principle #26Copying

2Quantity of substance

If simulated users are generated based on user data, then training data quantity increases, but system complexity increases

Engineering Contradiction:
Improvetraining data quantityVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The simulated user generation system serves multiple functions: it generates training data for reinforcement learning, preserves user privacy by not requiring direct access to real user data during training, enables scalable data generation, and can adapt to different user behaviors. This multi-functional approach justifies the added complexity by providing comprehensive benefits across data quantity, privacy, and training flexibility dimensions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of time

If reinforcement learning models are trained with limited data, then training time is reduced, but recommendation accuracy deteriorates

Engineering Contradiction:
Improvetraining timeVSAvoidrecommendation accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system performs preliminary action by pre-generating simulated user data and training the reinforcement learning model in advance with this synthetic data. This pre-training phase allows the model to learn from extensive simulated interactions before deployment, ensuring high recommendation accuracy is achieved before actual use. The preliminary training with simulated data eliminates the need for lengthy real-data training while maintaining model performance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250342403A1Method and server for providing personalized recommendation on basis of reinforcement learning
Publication Date: 2025.11.06 SAMSUNG ELECTRONICS CO LTD
  • US20250342403A1 patent drawing
  • US20250342403A1 patent drawing
  • US20250342403A1 patent drawing

AI summary

A method of providing a recommendation based on reinforcement learning includes: obtaining user data; generating a simulated user corresponding to an actual user based on the user data; determining an action based on a state of the simulated user, where the action corresponds to a recommendation element; updating the state of the simulated user; based on the updating, identifying the state of the simulated user and a reward output from the simulated user based on the recommendation element; generating a recommendation session by determining a plurality of the recommendation elements based on the state of the simulated user and the reward; and outputting the recommendation session.