Advertisement bidding method, system and device based on unified constraint bidding and medium
Through a unified constraint bidding model based on reinforcement learning, advertising bidding decisions are optimized, the accuracy problem of the RTB advertising bidding algorithm is solved, efficient advertising delivery management is achieved in budget-constrained scenarios, and advertising click revenue and resource allocation efficiency are improved.
Patent Information
- Application Number
- CN202510711090.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-16
AI Technical Summary
The existing RTB advertising bidding algorithm lacks precision, resulting in low budget utilization of advertisers and inability to accurately target target user groups, resulting in ineffective waste of funds and reduced conversion rates and return on investment.
A unified constraint bidding model based on reinforcement learning is adopted. By obtaining historical advertising bidding data, trajectory data is generated, and the evaluation sub-model and bidding sub-model are used for training to optimize advertising bidding decisions and make dynamic adjustments based on real-time budget consumption and market feedback.
Under the premise of ensuring budget control, maximize the revenue from ad clicks, improve the effectiveness of advertising delivery, and realize intelligent and automated advertising resource allocation.
Smart Images

Figure CN120655359A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of advertising delivery technology, and in particular relates to an advertising bidding method, system, device and medium based on unified constraint bidding. Background Art
[0002] In the programmatic advertising industry, demand-side platforms (DSPs) serve as a core tool for advertisers to deliver ads. They use real-time bidding (RTB) technology to precisely display ads to target users and collect commissions from advertisers based on ad performance (such as click-through rate or conversion rate). As competition in the advertising market intensifies, advertisers are increasingly demanding more refined budget allocation. DSPs, on the other hand, need to optimize bidding algorithms to maximize ad effectiveness within budget constraints, thereby improving revenue structure and resource allocation efficiency.
[0003] In programmatic advertising transactions, real-time bidding (RTB) systems enable dynamic allocation of advertising resources through millisecond-level bidding decisions. However, the core technologies of existing RTB ad bidding algorithms still have significant flaws, resulting in low budget utilization for advertisers. Traditional algorithms primarily rely on static prediction models based on historical click-through rates (CTR) or conversion rates (CVR) during the bidding process, failing to fully integrate real-time user behavior trajectories (such as page dwell time, cross-device interactions, and immediate search intent). This results in insufficiently accurate matching of bidding decisions with users' immediate needs.
[0004] Due to the poor effectiveness of existing advertising bidding algorithms, a large number of advertisements cannot accurately hit the target user groups. The budgets invested by advertisers are used to reach non-potential customers, resulting in ineffective waste of funds, significantly reducing the conversion rate and return on investment of advertising, and hindering the healthy and sustainable development of the advertising industry.
[0005] Therefore, there is an urgent need to develop a more efficient and accurate RTB advertising bidding algorithm to improve the effectiveness of advertising delivery, reduce advertiser costs, and achieve optimal allocation of advertising resources. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide an advertising bidding method, system, device and medium based on unified constraint bidding to solve the problem of insufficient accuracy of advertising bidding algorithms in related technologies.
[0007] In order to solve the above technical problems, this application provides the following technical solutions:
[0008] In a first aspect, the present application provides an advertising bidding method based on a unified constraint bid, comprising:
[0009] Acquire historical advertising bidding data, and generate first trajectory data based on the historical advertising bidding data;
[0010] Inputting the first trajectory data into a preset reinforcement learning-based training and evaluation model to generate second trajectory data;
[0011] Inputting the second trajectory data into a preset advertising bidding model based on unified constraint bidding for training, thereby obtaining a trained advertising bidding model;
[0012] The advertisements are bid in real time according to the trained advertisement bidding model.
[0013] Furthermore, the first trajectory data includes: state data, action data, reward data, termination mark data and next state data.
[0014] Furthermore, the preset advertising bidding model based on unified constraint bidding includes: an evaluation sub-model and a bidding sub-model; wherein,
[0015] The evaluation sub-model is used to evaluate the quality of the state and action in the second trajectory data to obtain a data quality evaluation result;
[0016] The bidding sub-model is used to optimize the advertising bid based on the results of the data quality assessment. Furthermore, the evaluation sub-model includes three fully connected layers, and the specific calculation formula is as follows:
[0017] h c1 =ReLU(W c1 s+b c1 )
[0018] ReLU(x)=max(0,x)
[0019] h c2 =ReLU(W c2 [h c1 ,a]+b c2 )
[0020] h c3 =ReLU(W c3 h c2 +b c3 )
[0021] Q(s,a)=sigmoid(W c4 h c3 +b c4 )
[0022]
[0023] Among them, W c1、W c2 、W c3 and W c4 is the weight of each connection layer; b c1 、b c2 、b c3 and b c4 is the bias of each connection layer; h c1 、h c2 and h c3 is the output of each connection layer; ReLU(x) is the activation function; Q(s,a) represents the expected cumulative reward of taking action a in state s; sigmoid is the activation function.
[0024] Furthermore, the bidding sub-model includes two hidden layers, and the specific calculation formula is as follows:
[0025] h a1 =ReLU(W a1 ·s+b a1 )
[0026] h a2 =ReLU(W a2 ·h a1 +b a2 )
[0027] a=tanh(W a3 ·h a2 +b a3 )
[0028]
[0029] Among them, W a1 、W a2 and W a4 is the weight of each layer; b a1 、b a2 and b a4 is the bias of each layer; a is the output action; tanh(x) is the activation function.
[0030] Furthermore, the evaluation sub-model adopts the following loss function:
[0031]
[0032] Among them, L Q is the loss function; N is the total number of samples; Q i is the predicted Q value of the state-action pair of the i-th sample of the Critic network; y i is the target Q value; r i is the immediate reward value obtained by the i-th sample from the environment after performing the action; v iis the total clicks that the model can win in the future time steps under the current bid coefficient; R optimal is the theoretical optimal reward under the current state; In all possible action sequences A={a1,a2,…,a T}, find the action sequence that maximizes the objective function; t represents the state at time step t, a t represents the action in t time steps, T is the total number of time steps; N t is the number of advertisements at time step t; B budget is the total budget, which represents the maximum expenditure allowed in the current time period.
[0033] Furthermore, the bidding sub-model adopts the following loss function:
[0034]
[0035] in, is the loss function; N is the number of samples; s i is the i-th state sample; μ(s i ) is the bidding sub-model in s i The output action; Q(s i ,μ(s i )) is the Q value output by the evaluation sub-model.
[0036] In a second aspect, the present application further provides an advertising bidding system based on a unified constraint bid, comprising:
[0037] an acquisition module, configured to acquire historical advertisement bidding data and generate first trajectory data based on the historical advertisement bidding data;
[0038] a data processing module, configured to input the first trajectory data into a preset reinforcement learning-based training and evaluation model to generate second trajectory data;
[0039] A training module, configured to input the second trajectory data into a preset advertising bidding model based on unified constraint bidding for training, thereby obtaining a trained advertising bidding model;
[0040] The advertisement bidding module is used to conduct real-time bidding on advertisements based on the trained advertisement bidding model.
[0041] In a third aspect, the present application also provides a computer electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above-mentioned advertising bidding methods based on unified constraint bidding are implemented.
[0042] In a fourth aspect, the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any one of the above-mentioned advertising bidding methods based on unified constraint bidding.
[0043] The present application provides an advertising bidding method, system, device, and medium based on unified constraint bidding, which have the following beneficial effects:
[0044] This application uses the USCB (Uniform Constrained Bidding) model to model the advertising bidding environment, automatically sensing and adapting to the budget constraints and market environment changes of different advertisers. During the training phase, the system collects and normalizes historical bidding data, builds an experience replay pool, and uses reinforcement learning algorithms to continuously iterate and optimize bidding strategies. During the actual delivery process, the model can dynamically adjust bidding decisions based on real-time budget consumption, historical bidding results, and market feedback, thereby maximizing ad click revenue while ensuring that the budget is under control. In addition, the system also integrates the estimation and dynamic adjustment mechanism of macro parameters such as daily budget and optimal rate of return to further enhance the adaptability and robustness of the strategy. The overall solution effectively solves the revenue bottleneck of traditional bidding methods in budget-constrained scenarios, and realizes intelligent and automated advertising delivery management. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0046] Figure 1 This is a flowchart of an advertising bidding method based on unified constraint bidding in an embodiment of the present application;
[0047] Figure 2 Schematic diagram of the training process of the reinforcement learning-based training and evaluation model in the embodiment of the present application;
[0048] Figure 3 This is a structural diagram of an advertising bidding system based on unified constraint bidding in an embodiment of the present application;
[0049] Figure 4 It is a structural diagram of a computer electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0051] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be an intermediate element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be an intermediate element. Conversely, when an element is referred to as being "directly on" another element, there is no intermediate element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only.
[0052] In this application, unless otherwise expressly specified or limited, terms such as "mounted," "connected," "connect," and "fixed" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integration; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components or interactions between two components. Those skilled in the art will understand the specific meanings of these terms in this application based on specific circumstances.
[0053] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0054] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the", and "the" used in one or more embodiments of the present application are also intended to include plural forms unless the context clearly indicates otherwise.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used in the template description herein are for the purpose of describing specific embodiments only and are not intended to limit this application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0056] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when..." or "when...".
[0057] Currently, existing advertising bidding methods have the following problems:
[0058] Traditional algorithm based on linear bidding:
[0059] Because linear bidding relies on a fixed scaling factor, α, it lacks flexibility when market prices fluctuate significantly. For example, when increased competition drives up costs, budgets may be exhausted prematurely. Conversely, when competition weakens, insufficient bidding may lead to missed opportunities for high-value traffic and low click-through rates.
[0060] Bidding method based on logistic regression:
[0061] Logistic regression relies heavily on feature engineering, requiring manual design of high-quality features. Furthermore, model parameters are trained based on historical data, making it difficult to adapt to changing market conditions. If market distribution shifts (e.g., due to changes in user behavior), the accuracy of the predicted pCTR decreases, and bidding performance deteriorates.
[0062] Bidding strategy based on greedy algorithm:
[0063] Greedy algorithms focus on short-term gains and ignore long-term budget allocation. For example, overspending on budget in the early stages of a bidding process can lead to missing out on high-quality traffic due to insufficient resources later on, resulting in lower-than-expected overall returns.
[0064] Reinforcement Learning Methods:
[0065] Traditional reinforcement learning relies on extensive trial-and-error to estimate state-action values (such as Q-values), resulting in high computational costs and slow convergence. In the high-frequency, low-latency (millisecond-level) RTB scenario, frequent environmental interactions are difficult to implement, making the algorithm difficult to implement practically. Furthermore, traditional reinforcement learning models incur high costs when interacting with real-world bidding environments.
[0066] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes in certain embodiments will not be repeated. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0067] Please refer to Figure 1 The advertising bidding method based on unified constraint bidding provided in the embodiment of the present application includes at least the following steps:
[0068] S10: Acquire historical advertisement bidding data, and generate first trajectory data based on the historical advertisement bidding data.
[0069] Specifically, in this application, in this example, first, the bidding data of historical advertisements can be obtained through the advertisement bidding platform.
[0070] In one embodiment, after obtaining the historical advertising bidding data, the historical advertising bidding data may be preprocessed, wherein the preprocessing includes:
[0071] Noise removal: Various noises may exist in the data, such as sensor measurement errors, environmental interference, etc. These noises will affect the learning effect of the algorithm and need to be removed through filtering, smoothing and other methods.
[0072] Handling missing values: Missing values in the data need to be handled according to the specific situation. Common methods include deleting samples with missing values, filling missing values with statistics such as mean, median, or mode, or using more complex interpolation methods.
[0073] Remove outliers: Outliers may be caused by data collection errors or special circumstances, and they may have a significant impact on model training. You can identify and remove outliers by setting thresholds, using boxplots, and other methods.
[0074] Secondly, after pre-processing the historical advertising bidding data, it is necessary to convert the historical advertising bidding data into first track data.
[0075] In one embodiment, the first trajectory data includes: state data, action data, reward data, termination mark data, and next state data.
[0076] Specifically, the format of the generated first trajectory data is as follows:
[0077] Taking daily historical advertising data as input, a set of (S t ,A t ,R t ,S t+1 ,done) format for reinforcement learning trajectory, which is used for subsequent strategy training and evaluation. The following is the definition of trajectory data:
[0078] 1. State data definition:
[0079] Each state S tIt is a fixed-length feature vector that fully describes the environment, resource constraints, and historical behavior trends at the current time step. The state vector consists of the following 16 dimensions:
[0080]
[0081]
[0082] 2. The action data action is defined as follows:
[0083] Action A t Represents the overall bidding tendency of the bidding agent at the current time step, which is calculated based on the bids and estimated values of all impression opportunities at the current time step:
[0084]
[0085] Among them, bidd i A unique identifier representing a single bid, pctr i Indicates the estimated ctr for this impression opportunity.
[0086] 3. The reward data Reward is defined as follows:
[0087] This study defines rewards as the total number of clicks on impressions won by the bidding agent:
[0088]
[0089] Click indicates whether a click occurred, 1 for click occurred, and 0 for click did not occur.
[0090] 4. Termination mark data done
[0091] At each time step, it is calculated whether the current trajectory has reached the terminal state. The calculation rules are as follows:
[0092]
[0093] Among them, step represents the time step of the bidding, and remainingBudget represents the remaining budget after the end of this step.
[0094] It should be noted that once the mark is done = 1, it indicates that the current trajectory is the end state and no S will be generated later. t+1 (Next state data)
[0095] 5. Next state data S t+1 :
[0096] In the non-terminal state, the system will set the state S of the next time step to t+1is constructed as the successor state of the current trajectory; and for the terminal time step (done=1), its S t+1 Set to None to mark the end of the track.
[0097] S20: Input the first trajectory data into a preset training and evaluation model based on reinforcement learning to generate second trajectory data.
[0098] For the sake of subsequent description, the preset reinforcement learning-based training and evaluation model will be referred to as the USCBTrainer model.
[0099] In one embodiment of the application, the preset reinforcement learning-based training and evaluation model includes: a data loading and preprocessing submodule, a parameter initialization submodule, a model training submodule, and a model evaluation submodule.
[0100] Please refer to the training process Figure 2 Specifically, the data loading and preprocessing submodule is used to load and organize the data management class of the advertising bidding experiment data, support the separate management of training and testing data, and provide step-by-step traffic samples when simulating bidding. It includes the following functions:
[0101] 1. Organize data by advertiser and training / testing phase.
[0102] 2. Provide input samples in time steps for reinforcement learning and bidding simulation.
[0103] 3. Output key data streams such as click, market_price (market price, i.e., the bid price for the display opportunity), and pctr (estimated click-through rate) for reward calculation.
[0104] The parameter initialization submodule is used to set the initial values of various parameters and define important information such as data storage path, model parameter storage path, budget information storage path, etc. It specifically includes the following functions:
[0105] 1. Set the budget ratio and advertiser ID.
[0106] 2. Load normalization parameters (state_mean, state_std).
[0107] 3. Load the trained optimal USCB (Uniform Constrained Bidding) model.
[0108] It should be noted that the calculation logic of the normalization parameter is as follows: for all state vectors during training, we can calculate the mean and standard deviation of each dimension. Specifically, assuming you have N state samples, each state is a d-dimensional vector:
[0109] Calculate the mean of each dimension (state_mean):
[0110]
[0111] Among them, s ij Represents the value of the jth dimension in the i-th state vector, μ j represents the mean of the j-th dimension.
[0112] Calculate the standard deviation of each dimension (state_std):
[0113]
[0114] Among them, σ j represents the standard deviation of the j-th dimension.
[0115] Normalize the state:
[0116]
[0117] Among them, μ i is the mean of the i-th dimension, σ i is the standard deviation of the i-th dimension.
[0118] Model training submodule (Run module): The main task of the Run module is to construct a USCB model instance and organize training based on the preset directory address and the relevant information passed in the init module. After the training is completed, the evaluate module is called to evaluate the model effect.
[0119] The model evaluation submodule, implemented through the evaluate() function, aims to simulate the actual performance of the current model in a real-world ad bidding scenario. The core steps of the evaluation process include: Based on daily ad traffic data, the bidding agent USCBAgent (whose function is simply to process the output action of the USCB input into a bid) loads the USCB model and executes a simulated bidding process. This simulated bidding process will be described in detail later. When the model budget is exhausted or the daily time limit is reached, the evaluation process terminates and all clicks received by the model are counted, which serves as a benchmark for evaluating model performance.
[0120] S30: Input the second trajectory data into a preset advertising bidding model based on unified constraint bidding for training, to obtain a trained advertising bidding model.
[0121] Specifically, in this embodiment, the second trajectory data is input into the advertising bidding model based on unified constraint bidding through the USCBTrainer model for training.
[0122] For the sake of subsequent description, in the subsequent description, the preset advertising bidding model based on the unified constrained bidding model will be referred to as the USCB model.
[0123] In one embodiment of the present application, the preset advertising bidding model based on unified constraint bidding includes: an evaluation sub-model () and a bidding sub-model; wherein,
[0124] The evaluation sub-model is used to evaluate the quality of the state and action in the second trajectory data to obtain a data quality evaluation result;
[0125] The bidding sub-model is used to optimize the advertising bid according to the result of the data quality assessment.
[0126] To facilitate subsequent descriptions, the evaluation sub-model will be referred to as the Critic network, and the bidding sub-model will be referred to as the Actor network.
[0127] Specifically, during training, the main function of the Critic network is to evaluate the quality of state-action pairs and guide the learning of the Actor network by estimating the Q value (expected cumulative reward) of a given state and action. During inference, the Actor network maps states to actions and determines the agent's actions based on the input state.
[0128] It should be noted that the Critic network is a neural network designed to estimate the Q-value of a state-action pair. During training, the Critic network updates its parameters by minimizing the error between the predicted Q-value and the target value, and provides gradients to the Actor network to optimize the policy. In the USCB model, its structure consists of the following components:
[0129] Input: State s: The state vector representing the environment. Action a: Generated by the Actor network or sampled from data.
[0130] The structure of the Critic network is as follows:
[0131] 1. FC1 (fully connected layer 1): maps the state s to a 10-dimensional hidden layer, using the ReLU activation function:
[0132] h c1 =ReLU(W c1 s+b c1 )
[0133] The formula for the ReLU function is:
[0134] ReLU(x)=max(0,x)
[0135] 2. FC2 (fully connected layer 2): concatenates the 10-dimensional state representation with action a to generate a 10+dim_action (action dimension) dimensional vector, which is then mapped to a 50-dimensional hidden layer using the ReLU activation function:
[0136] h c2 =ReLU(W c2 [h c1 ,a]+b c2 )
[0137] 3. FC3 (fully connected layer 3): maps the 50-dimensional hidden layer to a 10-dimensional hidden layer, using the ReLU activation function:
[0138] b c3 =ReLU(W c3 h c2 +b c3 )
[0139] 4. FC4 (Fully Connected Layer 4): Maps the 10-dimensional hidden layer to a single Q value, and the output is constrained to [0, 1] through the sigmoid activation function:
[0140] Q(s,a)=sigmoid(W c4 h c3 +b c4 )
[0141]
[0142] Output: A single scalar Q(s,a) representing the expected cumulative reward of taking action a in state s, with values in the range [0,1] (due to the sigmoid activation).
[0143] In the above calculation formula, W c1 、W c2 、W c3 and W c4 is the weight of each connection layer; b c1 、b c2 、b c3 and b c4 is the bias of each connection layer; h c1 、h c2 and h c3 is the output of each connection layer; ReLU(x) is the activation function; Q(s,a) represents the expected cumulative reward of taking action a in state s; sigmoid is the activation function.
[0144] In one embodiment of the application, the loss function of the evaluation sub-model (Critic network) is calculated as follows:
[0145]
[0146] Among them, L Q is the loss function; N is the total number of samples; Q i is the predicted Q value of the state-action pair of the i-th sample of the Critic network; y i is the target Q value; r i is the immediate reward value obtained by the i-th sample from the environment after performing the action; v i is the total clicks that the model can win in the future time steps under the current bid coefficient; R optimal is the theoretical optimal reward under the current state; In all possible action sequences A={a1,a2,…,a T}, find the action sequence that maximizes the objective function; t represents the state at time step t, a t represents the action in t time steps, T is the total number of time steps; N t is the number of advertisements at time step t; B budget is the total budget, which represents the maximum expenditure allowed in the current time period.
[0147] Actor Network:
[0148] The main task of the Actor network is to output a bidding action (i.e.,
[0149] The composition structure is as follows:
[0150] 1. First layer: input layer -> hidden layer 1:
[0151] h a1 =ReLU(W a1 ·s+b a1 )
[0152] 2. Second layer: hidden layer 1 -> hidden layer 2
[0153] h a2 =ReLU(W a2 ·h a1 +b a2 )
[0154] 3. Third layer: hidden layer 2->output layer (action)
[0155] a=tanh(W a3 ·h a2 +b a3 )
[0156]
[0157] In the above formula, W a1、W a2 and W a4 is the weight of each layer; b a1 、b a2 and b a4 is the bias of each layer; a is the output action; tanh(x) is the activation function.
[0158] In one embodiment of the present application, the calculation formula of the loss function of the bidding sub-model (Actor network) is as follows:
[0159]
[0160] in, is the loss function; N is the number of samples; s i is the i-th state sample; μ(s i ) is the bidding sub-model in s i The output action; Q(s i ,μ(s i )) is the Q value output by the evaluation sub-model.
[0161] S40: Conduct real-time bidding on advertisements based on the trained advertisement bidding model.
[0162] Specifically, in this embodiment, real-time bidding is performed on advertisements using the trained advertisement bidding model.
[0163] Specifically, after obtaining the trained advertising bidding model, you can bid for ads in real time based on real-time data. The bidding formula is as follows:
[0164] bid i =α0·(1+α)·pCTR i
[0165] Among them, bid i is the bid for the i-th advertisement; α0 is the initial bid coefficient; α is the bid adjustment coefficient output by the USCB model; pCTR i is the predicted click-through rate of the i-th advertisement.
[0166] It should be noted that the calculation process of the initial bid coefficient is as follows:
[0167] First, all ad display opportunities within the target date need to be obtained. Second, the cost-effectiveness of each ad display opportunity is calculated. Next, all display opportunities for that day are sorted from high to low by cost-effectiveness. Based on this sorting, the agent selects display opportunities for purchase in descending order of cost-effectiveness until the budget is exhausted. The inverse of the cost-effectiveness of the last successfully purchased display opportunity on that day is used as the initial bid coefficient. This initial bid coefficient reflects the lowest cost-effectiveness that the agent can accept within the budget boundary conditions. The cost-effectiveness of each ad display opportunity is calculated using the following formula:
[0168] Cost Performance=pctr / market_price
[0169] Among them, Cost Performance is the cost-effectiveness of the ad display opportunity; pctr is the estimated click-through rate of the ad display opportunity; market_price is the market price of the ad display opportunity.
[0170] In one embodiment of the present application, the method further includes:
[0171] S50: Evaluate the advertisement bidding model, and optimize the advertisement bidding model according to the evaluation result.
[0172] It is understandable that in order to make the bidding strategy of the advertising bidding model more accurate, this application also evaluates the trained advertising bidding model, and then optimizes the bidding strategy of the advertising bidding model based on the evaluation results. The evaluation process is as follows:
[0173] In this embodiment, the model evaluation mechanism is implemented through the evaluate() function of the USCBTrainer model, and its purpose is to simulate the actual performance of the current model in a real advertising bidding scenario. The core steps of the evaluation process include: for daily advertising traffic data, the bidding agent DDAgent (whose function is only to process the output action obtained by USCB through input into a bid) loads the USCB model and executes a simulated bidding process. The simulated bidding process will be described in detail later. When the model budget is exhausted or the daily time limit is reached, the evaluation process will terminate, and all clicks obtained by the model will be counted as a benchmark for evaluating model performance.
[0174] It should be noted that the basis for simulated bidding is the display opportunities that have been won in the past. The winning prices of these display opportunities are used as their market prices and provided to the intelligent agent (DDAgent) for bidding. If the intelligent agent's bid is greater than the market price of the display opportunity, it is deemed that the intelligent agent has won the display opportunity.
[0175] DDAgent's bidding follows the following formula:
[0176] bid i =α0·(1+α)·pCTR i
[0177] Among them, bid i is the bid for the i-th advertisement; α0 is the initial bid coefficient; α is the bid adjustment coefficient output by the USCB model; pCTR i is the predicted click-through rate of the i-th advertisement.
[0178] Ultimately, the bidding results for each time step will be saved in a data frame, and then the bidding strategy of the advertising bidding model will be optimized based on these data frames.
[0179] The present application provides an advertising bidding method based on unified constraint bidding, which has the following beneficial effects:
[0180] This application uses the USCB (Uniform Constrained Bidding) model to model the advertising bidding environment, automatically sensing and adapting to the budget constraints and market environment changes of different advertisers. During the training phase, the system collects and normalizes historical bidding data, builds an experience replay pool, and uses reinforcement learning algorithms to continuously iterate and optimize bidding strategies. During the actual delivery process, the model can dynamically adjust bidding decisions based on real-time budget consumption, historical bidding results, and market feedback, thereby maximizing ad click revenue while ensuring that the budget is under control. In addition, the system also integrates the estimation and dynamic adjustment mechanism of macro parameters such as daily budget and optimal rate of return to further enhance the adaptability and robustness of the strategy. The overall solution effectively solves the revenue bottleneck of traditional bidding methods in budget-constrained scenarios, and realizes intelligent and automated advertising delivery management.
[0181] See also Figure 3 The embodiment of the present application further provides an advertising bidding system 200 based on a decision transformer, including:
[0182] An acquisition module 201 is configured to acquire historical advertisement bidding data and generate first trajectory data based on the historical advertisement bidding data;
[0183] A data processing module 202 is configured to input the first trajectory data into a preset reinforcement learning-based training and evaluation model to generate second trajectory data;
[0184] A training module 203 is configured to input the second trajectory data into a preset advertising bidding model based on unified constraint bidding for training, thereby obtaining a trained advertising bidding model;
[0185] The advertisement bidding module 204 is configured to conduct real-time bidding on advertisements based on the trained advertisement bidding model.
[0186] See also Figure 4 An embodiment of the present application also provides a computer electronic device 300, including a memory 303 and a processor 302, wherein the memory 303 stores a computer program, and when the processor executes the computer program, it implements the steps of the advertising bidding method based on unified constraint bidding as described above.
[0187] Specifically, the electronic device 300 includes: a transceiver 301, a bus interface and a processor 302. The processor 302 is used to obtain historical advertising bidding data and generate first trajectory data based on the historical advertising bidding data; input the first trajectory data into a preset training and evaluation model based on reinforcement learning to generate second trajectory data; input the second trajectory data into a preset advertising bidding model based on unified constraint bidding for training to obtain a trained advertising bidding model; and perform real-time bidding on advertisements based on the trained advertising bidding model.
[0188] In the embodiment of the present application, the electronic device 300 further includes: a memory 303. Figure 4 In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits such as one or more processors represented by processor 302 and memory represented by memory 303. The bus architecture may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 301 may be a plurality of components, i.e., a transmitter and a receiver, providing a unit for communicating with various other devices over a transmission medium. The processor 302 is responsible for managing the bus architecture and general processing, and the memory 303 may store data used by the processor 302 when performing operations.
[0189] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any one of the above-mentioned advertising bidding methods based on unified constraint bidding.
[0190] In this embodiment, the computer-readable storage medium may be a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium may include, but is not limited to, a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0191] In all examples shown and described herein, any specific values should be interpreted as merely exemplary and not limiting, and thus other examples of the exemplary embodiments may have different values.
[0192] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0193] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or flowchart, and the combination of boxes in the structure diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0194] In addition, the functional modules or units in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0195] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a terminal device (which can be a smart phone, personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
[0196] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. An advertising bidding method based on unified constraint bidding, characterized in that: include: Acquire historical advertising bidding data, and generate first trajectory data based on the historical advertising bidding data; Inputting the first trajectory data into a preset reinforcement learning-based training and evaluation model to generate second trajectory data; Inputting the second trajectory data into a preset advertising bidding model based on unified constraint bidding for training, thereby obtaining a trained advertising bidding model; The advertisements are bid in real time according to the trained advertisement bidding model.
2. The advertising bidding method according to claim 1, characterized in that: The first trajectory data includes: state data, action data, reward data, termination mark data and next state data.
3. The advertising bidding method according to claim 1, characterized in that: The preset advertising bidding model based on unified constraint bidding includes: an evaluation sub-model and a bidding sub-model; wherein, The evaluation sub-model is used to evaluate the quality of the state and action in the second trajectory data to obtain a data quality evaluation result; The bidding sub-model is used to optimize the advertising bid according to the result of the data quality assessment.
4. The advertising bidding method according to claim 3, characterized in that: The evaluation sub-model includes three fully connected layers, and the specific calculation formula is as follows: h c1 =ReLU(W c1 s+b c1 ) ReLU(x)=max(0,x) h c2 =ReLU(W c2 [h c1 ,a]+b c2 ) h c3 =ReLU(W c3 h c2 +b c3 ) Q(s,a)=sigmoid(W c4 h c3 +b c4 ) Among them, W c1 、W c2 、W c3 and W c4 is the weight of each connection layer; b c1 、b c2 、b c3 and b c4 is the bias of each connection layer; h c1 、h c2 and h c3 is the output of each connection layer; ReLU(x) is the activation function; Q(s,a) represents the expected cumulative reward of taking action a in state s; sigmoid is the activation function.
5. The advertising bidding method according to claim 3, characterized in that: The bidding sub-model includes two hidden layers, and the specific calculation formula is as follows: h a1 =ReLU(W a1 s+b a1 ) h a2 =ReLU(W a2 h a1 +b a2 ) a=tanh(W a3 ·h a2 +b a3 ) Among them, W a1 、W a2 and W a4 is the weight of each layer; b a1 、b a2 and b a4 is the bias of each connection layer; a is the output action; tanh(x) is the activation function.
6. The advertising bidding method according to claim 3, characterized in that: The evaluation sub-model adopts the following loss function: Among them, L Q is the loss function; N is the total number of samples; Q i is the predicted Q value of the state-action pair of the i-th sample of the Critic network; y i is the target Q value; r i is the immediate reward value obtained by the i-th sample from the environment after performing the action; v i is the total clicks that the model can win in the future time steps under the current bid coefficient; R optimal is the theoretical optimal reward under the current state; Indicates that in all possible action sequences a={a1,a2,…,a T }, find the action sequence that maximizes the objective function; t represents the state at time step t, a t represents the action in t time steps, T is the total number of time steps; N t is the number of advertisements at time step t; B budget is the total budget, which represents the maximum expenditure allowed in the current time period.
7. The advertising bidding method according to claim 3, characterized in that: The bidding sub-model adopts the following loss function: in, is the loss function; N is the number of samples; s i is the i-th state sample; μ(s i ) is the bidding sub-model in s i The output action; Q(s i ,μ(s i )) is the Q value output by the evaluation sub-model.
8. An advertising bidding system based on unified constraint bidding, characterized in that: include: an acquisition module, configured to acquire historical advertisement bidding data and generate first trajectory data based on the historical advertisement bidding data; a data processing module, configured to input the first trajectory data into a preset reinforcement learning-based training and evaluation model to generate second trajectory data; A training module, configured to input the second trajectory data into a preset advertising bidding model based on unified constraint bidding for training, thereby obtaining a trained advertising bidding model; The advertisement bidding module is used to conduct real-time bidding on advertisements based on the trained advertisement bidding model.
9. A computer electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the advertising bidding method based on unified constraint bidding according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the advertising bidding method based on unified constraint bidding according to any one of claims 1 to 7 are implemented.