Reinforcement Learning Model for Long-Term User Engagement Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning agents face challenges in predicting the long-term impact of content item presentation on user engagement, leading to potential negative consequences that can deter users from future content interactions.

Innovation Solution

A system utilizing a machine learning model trained through reinforcement learning to determine whether to present a content item based on predicted long-term engagement scores, which assesses the potential impact of presenting a content item in a specific context, thereby deciding whether to show the content item to maintain user engagement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If reinforcement learning agents present content items to maximize short-term engagement, then immediate user interaction increases, but long-term user engagement deteriorates due to negative consequences

Engineering Contradiction:
Improveshort-term user engagementVSAvoidlong-term user engagement
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary assessment of content items by evaluating their predicted long-term impact on user engagement before presentation. The reinforcement learning model is trained to anticipate future engagement consequences, allowing the system to select content items that are likely to maintain or improve long-term engagement rather than merely maximizing immediate interactions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback mechanism where the reinforcement learning model learns from both short-term engagement outcomes and long-term engagement consequences. By incorporating long-term engagement metrics into the reward function, the model receives feedback that guides it to balance immediate user interaction with sustained future engagement, progressively improving its content selection strategy.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If the system presents all candidate content items to users, then content variety and user choice increase, but user experience deteriorates due to potential negative long-term consequences

Engineering Contradiction:
Improvecontent varietyVSAvoidnegative long-term consequences on user engagement
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The system applies different selection criteria to different content items based on their individual predicted long-term impact. Rather than uniformly presenting all content items or applying a single filter, the reinforcement learning model evaluates each content item's specific characteristics and potential consequences, making localized decisions about which items to present based on their unique engagement profiles.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts content presentation parameters based on predicted long-term impact. The reinforcement learning model modifies selection probabilities, presentation timing, and content item prioritization according to learned patterns about which content characteristics lead to positive or negative long-term engagement outcomes, thereby optimizing the balance between variety and user experience.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10839310B2Selecting content items using reinforcement learning
Publication Date: 2020.11.17 GOOGLE LLC
  • US10839310B2 patent drawing
  • US10839310B2 patent drawing
  • US10839310B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for using a machine learning model that has been trained through reinforcement learning to select a content item. One of the methods includes receiving first data characterizing a first context in which a first content item may be presented to a first user in a presentation environment; and providing the first data as input to a long-term engagement machine learning model, the model having been trained through reinforcement learning to: receive a plurality of inputs, and process each of the plurality of inputs to generate a respective engagement score for each input that represents a predicted, time-adjusted total number of selections by the respective user of future content items presented to the respective user in the presentation environment if the respective content item is presented in the respective context.