Deep Reinforcement Learning Agent for Marketing Strategy Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In complex service scenarios like recommendation marketing, existing machine learning models struggle to efficiently design and select applicable models to measure service execution results across multiple stages and factors, leading to suboptimal service execution effects.

Innovation Solution

A deep reinforcement learning system is employed, where an agent determines marketing behaviors based on marketing strategies and environment status, obtaining multiple execution results and calculating a reward score to update its strategy, thereby simultaneously learning multiple targets on the marketing effect chain.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple machine learning models are applied to different stages of marketing service, then measurement precision of service execution results is improved, but device complexity increases

Engineering Contradiction:
Improvemeasurement precision of service execution resultsVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple machine learning models (classification model, regression model, and deep reinforcement learning model) into a unified marketing service system. These models work together to handle different aspects of marketing execution measurement - the classification model identifies user response types, the regression model quantifies execution results, and the deep reinforcement learning model optimizes marketing strategies based on cumulative rewards from multiple stages, thereby improving measurement precision while managing system complexity through integrated architecture

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a multi-functional measurement system where a single integrated framework performs multiple functions across different marketing stages. The system can simultaneously evaluate classification accuracy, regression precision, and reinforcement learning optimization across selection, sending, and feedback stages, making the measurement apparatus universally applicable to various marketing scenarios without requiring separate dedicated systems for each function

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If deep reinforcement learning system learns multiple targets simultaneously, then service execution effects are improved, but difficulty of detecting and measuring increases

Engineering Contradiction:
Improveservice execution effectsVSAvoiddifficulty of detecting and measuring
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the complex multi-target learning process into distinct measurable components. The deep reinforcement learning system evaluates multiple targets (classification accuracy, regression precision, and optimization metrics) separately at each marketing stage, assigning specific reward values to each target achievement. This segmentation allows the system to track and measure each target independently while still optimizing for overall service execution effects across all targets simultaneously

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11188928B2Marketing method and apparatus based on deep reinforcement learning
Publication Date: 2021.11.30 ADVANCED NEW TECHNOLOGIES CO LTD
  • US11188928B2 patent drawing
  • US11188928B2 patent drawing
  • US11188928B2 patent drawing

AI summary

Embodiments of the present specification provide marketing methods based on a deep reinforcement learning system. One method includes the following: obtaining, from an execution environment of a deep reinforcement learning system, a plurality of execution results generated by a user in response to marketing activities, wherein the plurality of execution results correspond to a plurality of targeted effects on a marketing effect chain; determining a reward score of reinforcement learning based on the plurality of execution results; and returning the reward score to a smart agent of the deep reinforcement learning system, for the smart agent to update a marketing strategy, wherein the smart agent is configured to determine the marketing activities based on the marketing strategy and status of the execution environment.