Reinforcement Learning Bidding Model for Ad Exchange Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing real-time bidding systems in programmatic advertising rely on rigid and reactionary policies, requiring human configuration and monitoring, and lack adaptive strategies to optimize ad placement in real-time.

Innovation Solution

A multi-platform integration system employing reinforcement learning to automate bidding strategies, adjusting based on factors like webpage content, consumer interactions, and time of day, eliminating the need for human oversight and enabling automatic discovery of sophisticated bidding policies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If rigid automatic bidding policies are used, then bidding process is controlled and predictable, but adaptability to changing conditions is poor

Engineering Contradiction:
Improvebidding controlVSAvoidpolicy adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by transitioning from static rigid bidding policies to dynamic reinforcement learning policies that continuously adapt to changing market conditions. The RL agent learns optimal bidding strategies through continuous interaction with the ad exchange environment, allowing the system to dynamically adjust bids based on real-time feedback while maintaining controlled experimentation through the simulation framework.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements self-service by enabling automatic discovery of bidding policies through reinforcement learning without requiring human configuration or monitoring. The RL agent autonomously learns and optimizes bidding strategies by interacting with the simulated ad exchange environment, eliminating the need for manual policy management while maintaining reliable control through the structured simulation framework.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If reinforcement learning is used to automate bidding, then adaptability and automatic discovery improve, but system complexity increases

Engineering Contradiction:
Improvebidding adaptabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a simulation environment as an intermediary between the reinforcement learning agent and the real ad exchange. This intermediary layer allows the RL agent to learn complex bidding policies in a controlled virtual environment before deployment, reducing the complexity burden on the production system while maintaining high adaptability through continuous learning from simulated market conditions.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary action by training the reinforcement learning agent in a simulated ad exchange environment before deploying to production. This pre-training phase allows the complex RL policies to be developed and optimized offline, reducing the complexity impact on the live bidding system while maintaining adaptability through subsequent fine-tuning and continuous learning.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If existing rigid bidding policies are used, then implementation is simple, but real-time optimization capability is limited

Engineering Contradiction:
Improveimplementation simplicityVSAvoidbid optimization efficiency
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent replaces mechanical rule-based bidding systems with intelligence-driven reinforcement learning policies. Instead of relying on pre-programmed bidding rules that require simple implementation but lack optimization capability, the system uses RL agents that automatically learn optimal bidding strategies, significantly improving bid optimization efficiency while managing implementation complexity through the simulation framework.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Reliability

If manual configuration and monitoring of bidding policies is required, then control is maintained, but operational efficiency decreases

Engineering Contradiction:
Improvepolicy controlVSAvoidoperational efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system implements self-service by enabling automatic discovery and optimization of bidding policies through reinforcement learning without requiring human configuration or monitoring. The RL agent autonomously learns optimal strategies by interacting with the simulated ad exchange environment, eliminating manual operational tasks while maintaining reliable control through the structured simulation framework and automated deployment processes.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11941668B2Ad exchange bid optimization with reinforcement learning
Publication Date: 2024.03.26 ZETA GLOBAL CORP
  • US11941668B2 patent drawing
  • US11941668B2 patent drawing
  • US11941668B2 patent drawing

AI summary

A system for training a bidding model comprising: a plurality of tactics stored on at least one database; a plurality of hyperparameters; in response to an available inventory from a publisher relayed through a real time bid server, computing a bid on the available inventory; sending the bid to the real time bid server; receiving an auction result in response to the bid; calculating a plurality of rewards based on the auction result and the tactics; calculate a plurality of q values based on the rewards; calculate a plurality of losses; backpropogating the losses through the bidding model.