Constraint-Sampled Reinforcement Learning for Diverse, Novel Recommendations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional recommendation systems prioritize accuracy over diversity and novelty, leading to decreased user satisfaction over time due to boredom from lack of exposure to varied media items.

Innovation Solution

A recommendation network trained using reinforcement learning techniques with a constraint sampling algorithm that optimizes for lifetime values, incorporating factors like accuracy, diversity, and novelty to enhance user satisfaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional recommendation systems focus on accurate predictions of user preferences, then recommendation accuracy is improved, but diversity and novelty of recommended items deteriorate

Engineering Contradiction:
Improverecommendation accuracyVSAvoiddiversity and novelty
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the optimization parameter from short-term accuracy to long-term user satisfaction, which incorporates diversity and novelty as additional factors. The reinforcement learning model optimizes for cumulative rewards over multiple interactions rather than immediate prediction accuracy, allowing the system to balance accuracy with diversity and novelty in its recommendations.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback mechanisms where user interactions with recommended items are continuously observed and used to update the reinforcement learning model. This feedback loop allows the system to learn from user responses and adjust its recommendation strategy to maintain both accuracy and diversity over time, preventing user boredom while sustaining engagement.

Inventive Principle:
Principle #23Feedback

2Reliability

If recommendation systems provide highly accurate predictions based on past user behavior, then short-term user satisfaction is improved, but long-term user engagement deteriorates due to boredom

Engineering Contradiction:
Improveshort-term user satisfactionVSAvoidlong-term user engagement
Core Design Contradiction:
ReliabilityVSDuration of action of stationary object

Solution Approach 1:

The patent makes the recommendation system dynamic by using reinforcement learning that continuously adapts to user responses. Rather than providing static accurate predictions, the system dynamically adjusts its recommendations based on observed user interactions, allowing it to balance short-term satisfaction with long-term engagement by introducing novel items when users show signs of boredom.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent takes preliminary action by proactively introducing diverse and novel items before users become bored. The reinforcement learning model anticipates user needs by exploring the item space and preparing recommendations that maintain engagement, rather than waiting for user dissatisfaction to manifest.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If recommendation systems explore diverse items to prevent boredom, then diversity is improved, but recommendation accuracy deteriorates

Engineering Contradiction:
ImprovediversityVSAvoidrecommendation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies partial exploration by introducing diverse items at controlled rates rather than uniformly. The reinforcement learning model balances exploitation of known user preferences with exploration of novel items, applying diversity partially to maintain accuracy while preventing boredom. This selective exploration ensures that diversity is introduced only when beneficial to long-term engagement.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12380359B2Constraint sampling reinforcement learning for recommendation systems
Publication Date: 2025.08.05 ADOBE INC
  • US12380359B2 patent drawing
  • US12380359B2 patent drawing
  • US12380359B2 patent drawing

AI summary

Systems and methods for sequential recommendation receive a user interaction history including interactions of a user with a plurality of items, select a constraint from a plurality of candidate constraints based on lifetime values observed for the candidate constraints, wherein the lifetime values are based on items predicted for other users using a recommendation network subject to the candidate constraints, and predict a next item for the user based on the user interaction history using the recommendation network subject to the selected constraint.