Training method and decision method of decision model
Patent Information
- Application Number
- CN202511892244.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-12-15
AI Technical Summary
[0003]然而,由于实际过程中,量价关系并非平滑曲线,因此根据机器学习模型预测的方式,容易陷入局部最优的陷阱
[0010]根据本说明书实施例的第六方面,提供了一种计算机可读存储介质,其存储有计算机程序或指令,该计算机程序或指令被处理器执行时实现上述方法的步骤。
Smart Images

Figure CN122089406B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of computer technology, and in particular to a training method and a decision-making method for a decision model. Background Technology
[0002] With the continuous development of computer technology, driving traffic to stores and increasing store revenue through external website advertising has become an increasingly popular marketing method for advertisers. The process of advertising on external websites typically involves a prediction-plus-decision approach. In the prediction phase, machine learning models predict the price-volume relationship under the current bidding environment. In the decision-making phase, operations research or control algorithms are used to find target configurations and determine the advertiser's bidding strategy to achieve better advertising performance.
[0003] However, since the relationship between volume and price is not a smooth curve in practice, predictions based on machine learning models are prone to falling into local optima. Furthermore, errors in the volume-price model during the prediction phase are amplified during the decision-making phase, leading to cascading errors. Additionally, the training phase of the prediction and decision-making process is heavily influenced by historical bidding strategies, resulting in significant sample selection bias. Moreover, because users rarely modify their bidding configurations, the training data is very sparse for modeling the volume-price relationship. Therefore, using these techniques makes it difficult to achieve good advertising buying performance. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a training method and a decision-making method for a decision model. One or more embodiments of this specification also relate to a training apparatus for a decision model, a decision-making apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a method for training a decision model is provided, comprising: Acquire training sequence data, wherein the training sequence data consists of multiple time step data, and each time step data includes action parameters and reward parameters; The training sequence data is input into the initial decision model to obtain the predicted sequence data output by the initial decision model. The predicted sequence data includes the predicted action parameters, predicted reward parameters, and predicted surplus parameter information corresponding to each time step data. The predicted reward parameters are predicted based on multiple reward parameters. Based on the training sequence data and the prediction sequence data, the target loss of the initial decision model is determined, wherein the target loss includes action loss, multi-attribution loss and quantile loss, the action loss is determined based on action parameters and predicted action parameters, and the multi-attribution loss and quantile loss are determined based on reward parameters, predicted reward parameters and predicted surplus parameters. Based on the target loss of the initial decision model, the parameters of the initial decision model are adjusted to obtain the final decision model.
[0006] According to a second aspect of the embodiments of this specification, a decision-making method is provided, the method comprising: Determine the decision-making time step and obtain the current state parameters, historical sequence data, and current balance parameter information corresponding to the decision-making time step; Based on the current state parameters, historical sequence data, and current balance parameter information, determine the sequence data corresponding to the time step to be decided; The sequence data is input into the decision model to obtain the predicted action parameters of the next time step associated with the time step to be decided, which are output by the decision model. The decision model is obtained based on the training method of any decision model involved in the first aspect of the embodiments of this specification.
[0007] According to a third aspect of the embodiments of this specification, a training apparatus for a decision model is provided, comprising: The acquisition unit is configured to acquire training sequence data, wherein the training sequence data consists of multiple time step data, and each time step data includes action parameters and reward parameters; The prediction unit is configured to input the training sequence data into an initial decision model to obtain the prediction sequence data output by the initial decision model. The prediction sequence data includes prediction action parameters, prediction reward parameters, and prediction surplus parameter information corresponding to each time step data. The prediction reward parameters are predicted based on multiple reward parameters. The determining unit is configured to determine the target loss of the initial decision model based on the training sequence data and the prediction sequence data, wherein the target loss includes action loss, multi-attribution loss and quantile loss, the action loss is determined based on action parameters and predicted action parameters, and the multi-attribution loss and quantile loss are determined based on reward parameters, predicted reward parameters and predicted surplus parameter information; The adjustment unit is configured to adjust the parameters of the initial decision model based on the target loss of the initial decision model to obtain a decision model.
[0008] According to a fourth aspect of the embodiments of this specification, a decision-making device is provided, comprising: The first determining unit is configured to determine the decision-making time step and obtain the current state parameters, historical sequence data and current balance parameter information corresponding to the decision-making time step. The second determining unit is configured to determine the sequence data corresponding to the time step to be decided based on the current state parameters, historical sequence data, and current balance parameter information. The prediction unit is configured to input the sequence data into the decision model to obtain the predicted action parameters of the next time step associated with the time step to be decided, output by the decision model.
[0009] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions, which, when executed by the processor, implement the steps of the above method.
[0010] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program or instructions which, when executed by a processor, implement the steps of the above-described method.
[0011] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0012] According to one embodiment of this specification, the model directly outputs predicted action parameters and performs end-to-end optimization using a target loss that includes action losses. This avoids the phenomenon of errors in the prediction stage being amplified in subsequent optimization stages, improving the accuracy and consistency of decision-making. Based on multi-attribution loss, the model can predict the predicted return parameters based on multiple return parameters. This avoids the bias caused by using a single-attribution loss, enhances the model's ability to learn robust patterns from sparse and biased data, and improves generalization and robustness. A quantile loss is introduced and applied to the predicted surplus parameter information. This encourages the model to explore strategies that bring higher long-term returns in complex, non-convex advertising revenue environments, helping to discover globally superior bidding strategies. Attached Figure Description
[0013] Figure 1 A flowchart is shown of a training method for a decision model according to an embodiment of this specification; Figure 2 A flowchart illustrating the processing steps of a decision-making method provided in one embodiment of this specification is shown. Figure 3A schematic diagram of the structure of a training device for a decision model provided in one embodiment of this specification is shown; Figure 4 This specification shows a schematic diagram of the structure of a decision-making device according to one embodiment; Figure 5 A structural block diagram of a computing device according to an embodiment of this application is shown. Detailed Implementation
[0014] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0017] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0018] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model or basic model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include large-scale language models (LLMs) and multi-modal pre-training models.
[0019] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0020] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0021] Billing models include the cost-per-performance (Cost Per X) model. Cost Per X can be understood as a series of billing models where you pay based on X results. Here, "X" is a variable representing different levels and types of target results sought by the advertiser, such as cost-per-thousand impressions (CPM) and cost-per-action (CPA).
[0022] Click-through rate (CTR) can be understood as the percentage of times an ad is displayed (exposure) and a user clicks on it.
[0023] External advertising can be understood as advertising that is placed outside the platform. Specifically, it can be understood as advertising that merchants purchase on external platforms within a platform or ecosystem in order to acquire traffic and users from other external platforms.
[0024] The dynamic pricing problem can be understood as how advertisers can automatically and intelligently adjust the price of goods or services based on real-time changes in supply and demand, market environment, competitor behavior, and customer characteristics in order to maximize goals such as revenue, profit, or market share.
[0025] With the continuous development of computer technology, and because bidding constraints are closely related to the quality and efficiency of advertising buying, advertisers modify their bidding strategies to achieve better advertising performance. Therefore, how to optimize advertising buying bidding strategies has become an urgent problem to be solved.
[0026] In related technologies, advertising bidding is typically determined by a combination of prediction and decision-making. However, this approach is prone to local optima because the price-volume curve is not smooth in real-world scenarios. Furthermore, since the decision-making process depends on the prediction results, cascading errors can easily occur, affecting the quality of the decisions. Moreover, the training data is influenced by historical bidding strategies, leading to sample selection bias. Additionally, the infrequent modification of bidding configurations results in sparse training data for modeling the price-volume relationship.
[0027] Therefore, this specification provides a training method and a decision-making method for a decision model. This specification also relates to a training device for a decision model, a decision-making device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0028] See Figure 1 , Figure 1 A flowchart is shown of a training method for a decision model according to an embodiment of this specification, specifically including the following steps 102-108.
[0029] Step 102: Obtain training sequence data, wherein the training sequence data consists of multiple time step data, and each time step data includes action parameters and reward parameters.
[0030] The decision model can be understood as an artificial intelligence model that learns from data and outputs decisions based on input state or sequence information. In this specification, the decision model can be understood as a generative model based on the Transformer architecture used for optimizing ad bidding configuration.
[0031] Training sequence data can be understood as an ordered set of data used to train the model. In this specification, training sequence data can be understood as a complete historical advertising delivery trajectory, divided by time steps (e.g., hourly), with each slice containing data for that time period.
[0032] Time step data can be understood as a data unit within a training sequence. It includes action parameters, reward parameters, and state parameters. Action parameters can be understood as the bid for the current time step, such as the CPX value. Reward parameters can be understood as the results generated at the current time step (such as clicks, cost per conversion, GMV, etc.). State parameters can be understood as the current environment, such as the traffic quality at the current time step.
[0033] In one specific embodiment provided in this specification, training sequence data is obtained. Advertising campaign logs are extracted from a historical database and organized into a structured sequence dataset in chronological order.
[0034] Specifically, taking the construction of advertising placement training sequence data as an example, historical advertising placement data is collected, and based on this data, the state characteristics, action characteristics, and reward characteristics corresponding to different time steps are determined. For ease of understanding, this specification uses a)-c) to explain these characteristics.
[0035] a) State features, used to reflect the current environment of ad delivery. The constructed state features include: Budget-related factors: remaining budget, budget consumption rate, and budget consumption percentage, etc.
[0036] Time-related information: current time period, day of the week, deployment progress (deployment time or total planned time), etc.
[0037] Environmental factors include: intensity of competition (reflected by indicators such as win rate or number of bids), and traffic quality.
[0038] b) Action characteristics, used to reflect bid configurations, such as target CPA bid value. It's important to understand that in historical data, actions represent the actual bid value set by the advertiser.
[0039] c) Return characteristics and RTG (Return-to-Go): The reward can be understood as the immediate return obtained at each time step, such as the GMV (Gross Merchandise Volume) generated at that time step. It is important to understand that due to data latency, the reward data may not be known until several hours after the action occurs.
[0040] RTG can be understood as the cumulative return from the current time to the end of the sequence. It's important to understand that in this specification, RTG is characterized based on the balance parameter information.
[0041] Given different features, training sequence data is constructed in chronological order based on the state parameters, reward parameters, and action parameters corresponding to different time steps.
[0042] Step 104: Input the training sequence data into the initial decision model to obtain the predicted sequence data output by the initial decision model.
[0043] The predicted sequence data includes predicted action parameters, predicted reward parameters, and predicted surplus parameter information corresponding to each time step data. The predicted reward parameters are predicted based on multiple reward parameters.
[0044] The initial decision model can be understood as a model architecture that is untrained or incompletely trained, with parameters in a random or pre-initialized state.
[0045] Predicted sequence data can be understood as the corresponding predicted values output by the model for the input training sequence, including predicted action parameters, predicted reward parameters, and predicted surplus parameter information. For ease of understanding, this specification uses a)-c) as examples: a) Predicted action parameters can be understood as the action (bid) that the initial decision model believes should be taken at this time step.
[0046] (b) The predicted return parameter can be understood as the return predicted by the initial decision model at this time step. It should be understood that the predicted return parameter involved in this specification is derived based on the combination of multiple attribution strategies.
[0047] c) The predicted residual parameter information can be understood as the cumulative remaining reward (Return to Go, RTG) predicted by the model from the current time step to the end of the task.
[0048] In one specific implementation provided in this specification, the model performs forward propagation. A training sequence is input into the initial decision model. After internal processing, the model outputs the predicted action, predicted reward, and predicted RTG for the entire sequence.
[0049] Based on the above, the training sequence data is input into the initial decision model to obtain the predicted sequence data output by the initial decision model, including: The training sequence data is input into the initial decision model to obtain at least one attribution reward parameter corresponding to each time step determined by the initial decision model; Based on at least one attribution reward parameter corresponding to each time step, predict the predicted reward parameter corresponding to each time step. Multiple predicted return parameters are fused to obtain fused features, and action prediction parameters and predicted surplus parameters are generated based on the fused features.
[0050] In this context, at least one attribution reward parameter corresponding to each time step can be understood as the expected reward value independently predicted by the model in the intermediate layer for each input time step, corresponding to at least one attribution strategy perspective.
[0051] The predicted return parameters for each time step can be understood as the single, comprehensive predicted return value for that time step, which is further generated by the model after obtaining the preliminary predictions from the above multiple attribution perspectives.
[0052] Fusing multiple predicted return parameters can be understood as performing feature aggregation, which involves interacting and integrating the predicted return parameters of all (or within a window) time steps in the sequence according to the model's sequence processing capabilities (such as the Self-Attention mechanism of Transformer).
[0053] Fusion features can be understood as a high-dimensional feature vector or vector sequence that has undergone in-depth processing and aggregates information from multiple time steps and multiple perspectives in the sequence.
[0054] Action prediction parameters can be understood as the decision recommendations for each time step that the model ultimately decodes based on the fused features, i.e., the optimal output value (CPX).
[0055] The predicted balance parameter information can be understood as the model's prediction of the future cumulative return (RTG) starting from the current time step, based on the fused feature decoding.
[0056] In one specific embodiment provided in this specification, the model's front-end network (e.g., a shared feature extraction layer followed by multiple parallel output heads) directly generates a set of preliminary predictions in parallel for each time step of the input sequence. A portion (or all) of these predictions is used to estimate the value of that time step under different attribution strategies, i.e., at least one attribution reward parameter. For each time step, the attribution predictions (at least one attribution reward parameter) from its multiple perspectives are summarized or transformed to form a unified, scalar, or low-dimensional vector of predicted reward parameters.
[0057] Step 106: Based on the training sequence data and the predicted sequence data, determine the target loss of the initial decision model.
[0058] The target loss includes action loss, multi-attribution loss, and quantile loss. The action loss is determined based on action parameters and predicted action parameters, while the multi-attribution loss and quantile loss are determined based on reward parameters, predicted reward parameters, and predicted surplus parameters.
[0059] The target loss can be understood as a comprehensive error function used to measure the difference between the model's predicted values and the true values. Specifically, the target loss includes action loss, multi-attribution loss, and quantile loss. For ease of understanding, this specification uses a)-c) to explain the components of the target loss.
[0060] a) Action loss can be understood as the deviation between the action prediction output by the model and the historical actual action. That is, the deviation between the action prediction parameters and the action parameters.
[0061] b) Multi-attribution loss can be understood as the deviation between the model's predicted returns from multiple different attribution perspectives and the true return labels redistributed according to the corresponding attribution rules.
[0062] c) Quantile loss can be understood as an upper limit for the potential future return (RTG) in a given state, rather than just the average, thereby incentivizing the model to explore strategies with higher returns.
[0063] In one specific embodiment provided in this specification, a composite target loss is calculated. The predicted sequence data output by the model is compared with the input training sequence data. Action loss, multi-attribution loss, and quantile loss are calculated. The three losses are then weighted and summed to obtain the final target loss value.
[0064] For ease of understanding, the method for determining the target loss of the initial decision model is explained in the following manner in this specification.
[0065] In one specific embodiment provided in this specification, the multi-attribution loss is determined using methods S1062-S1068: S1062. Determine the time step to be processed and at least one associated time step corresponding to the time step to be processed.
[0066] In this context, the pending time step can be understood as the specific point in time (i.e., a time step t in the sequence) where the loss is currently being evaluated and calculated. For example, the pending time step could be the time step in which the target event (e.g., conversion) occurs. It's important to understand that since the final value of ad reach can only be determined at the time step in which the target event occurs, attribution needs to be performed based on the pending time step.
[0067] In one specific implementation provided in this specification, the training sequence data is traversed to identify all time points where a "conversion" event occurred, and these are marked as time steps to be processed. For each time step to be processed, the process is backtracked on the timeline to find all associated time steps that generated valid interactions (such as clicks) with this advertising campaign before the current conversion.
[0068] Specifically, determining the time step to be processed and at least one associated time step corresponding to the time step to be processed includes: The time step corresponding to the target behavior information is determined as the time step to be processed. Determine the associated behavior information corresponding to the target behavior information, and determine at least one associated time step based on the associated behavior information.
[0069] Among them, target behavior information can be understood as the information carried by the core result event after conversion, which is ultimately defined during the process of advertising and user interaction.
[0070] Related behavioral information can be understood as information that occurs before the occurrence of the target behavioral information and has a pre-defined sequential relationship with the occurrence of the target behavioral information.
[0071] In one specific embodiment provided in this specification, the entire training sequence data is scanned to identify all time points that record target behavior information (e.g., conversion), and these time points are determined as pending time steps requiring subsequent attribution calculations. Further, for each identified target behavior information, associated behavior information is determined according to preset business rules. The occurrence time point of each associated behavior information is extracted, and the time step to which that time point belongs in the training sequence is found. Based on the set of all time steps, at least one associated time step corresponding to the pending time step is determined.
[0072] According to a specific implementation method provided in this specification, related behavioral information is determined based on target behavioral information, and then irrelevant behavioral information is filtered out, thereby providing accurate supervision signals for the model, avoiding noise interference during the training process, and accelerating the model convergence speed.
[0073] S1064. Determine the attribution reward label corresponding to the time step to be processed based on the time step to be processed and the reward parameters corresponding to each associated time step.
[0074] In this context, a related time step can be understood as one or more earlier time steps that have a causal or sequential relationship with the final conversion event represented by the time step to be processed. For example, a related time step can be understood as all the time points during which the same user has been exposed to different advertising plans under this advertising campaign before the final conversion occurs, i.e., between the time steps to be processed. The conversion corresponding to the time step to be processed is triggered by the triggering of all the time points corresponding to the related time steps.
[0075] Attribution reward tags can be understood as new reward values for a given time step. These values are calculated based on pre-defined attribution rules, recalculating and redistributing a portion of the total conversion value (i.e., reward parameters) to the current time step and any related time steps. It's important to understand that attribution reward tags are artificially constructed values based on the attribution strategy and differ from the rewards recorded in the original logs for that time step. A single time step may correspond to multiple attribution reward tags. For example, one tag might be obtained using the last-click attribution strategy, another using a linear attribution strategy, and so on. It's also important to understand that the original logs may only record reward parameters obtained from direct clicks, not the attributed reward parameters.
[0076] In one specific embodiment provided in this specification, the attribution reward parameters obtained from the final transformation of the time step to be processed are acquired. The total attribution reward parameters are allocated to the relevant time steps according to one or more predefined attribution strategies. For the current time step to be processed, the reward parameters allocated to the time step according to each attribution strategy are extracted. These reward parameters are then determined as the attribution reward label for that time step under the current attribution strategy.
[0077] For example, suppose a user clicks on ad A in the 3rd hour, clicks on ad B in the 5th hour, and converts in the 8th hour. Using a linear attribution strategy, the 100 yuan is evenly distributed across the 3rd, 5th, and 8th hours. Therefore, for the pending time step (the 8th hour), the attribution reward label would be 33.33 yuan.
[0078] Specifically, based on the pending time step and the reward parameters corresponding to each associated time step, the attribution reward label corresponding to the pending time step is determined, including: The conversion reward parameters are determined based on the time step to be processed and the reward parameters corresponding to each associated time step. Based on multiple attribution strategies and the conversion return parameters, the attribution return parameters corresponding to each attribution strategy for the time step to be processed are determined. Attribution reward labels are generated for the time steps to be processed based on each attribution reward parameter.
[0079] The conversion return parameter can be understood as a quantitative indicator representing the total value created by this conversion, extracted from the original return parameters corresponding to the time step to be processed.
[0080] Multiple attribution strategies can be understood as a predefined set of mathematical rules or business logic used to assign conversion reward parameters to relevant time steps on the conversion path. For example, according to any attribution strategy, conversion reward parameters extracted based on the time step to be processed can be assigned to one or more related time steps in the conversion path.
[0081] Attribution reward parameters can be understood as the value allocated from the transformation reward parameters to a specific time step to be processed, based on the rules of a specific attribution strategy.
[0082] Attribution reward labels can be understood as ground truth labels for the reward at a given time step, used to supervise model training. It's important to understand that the attribution reward labels discussed in this specification are determined by multiple attribution reward parameters calculated for that time step under various attribution strategies. That is, each time step has corresponding attribution reward parameters for different attribution strategies, and the attribution reward labels for that time step under different attribution strategies are determined based on these parameters.
[0083] In one specific embodiment provided in this specification, the total value corresponding to the target behavior is parsed or calculated from the reward parameters based on the time step to be processed. For each preset attribution strategy, an allocation calculation is performed independently to extract the value allocated to the time step to be processed, thereby determining the attribution reward parameters. This results in multiple attribution reward parameters. Based on these multiple attribution reward parameters, the attribution reward label corresponding to each attribution reward parameter is obtained.
[0084] According to a specific implementation method provided in this specification, a multidimensional supervision signal space is constructed. During model training, the model simultaneously fits the allocation of reward parameters corresponding to multiple attribution strategies, thereby preventing the model from over-relying on any potentially biased attribution logic.
[0085] Specifically, for ease of understanding, this manual explains the attribution parameters corresponding to the time steps to be processed, based on different attribution strategies and conversion return parameters, in the following manner.
[0086] In one specific embodiment provided in this specification, the attribution reward parameter for the time step to be processed at each attribution strategy is determined based on multiple attribution strategies and the conversion reward parameter, including: Determine an initial attribution strategy, wherein the initial attribution strategy is any one of a plurality of attribution strategies; Based on the initial attribution strategy and the conversion reward parameters, the attribution reward parameters corresponding to the initial attribution strategy are determined for the time step to be processed.
[0087] The initial attribution strategy can be understood as the attribution strategy that is currently selected and calculated when iteratively processing multiple sets of attribution strategies.
[0088] In one specific implementation provided in this specification, under the specific mathematical rules defined by the initial attribution strategy, the conversion reward parameter (total value) is assigned to all relevant time steps on the conversion path, and the values assigned to the time steps to be processed are extracted.
[0089] Based on the above, it should be understood that the time step to be processed and at least one associated time step can be understood as sequential data arranged according to the transformation behavior and in chronological order. The time point corresponding to each associated time step is earlier than the time point corresponding to the time step to be processed. Furthermore, for ease of explanation, in this specification, the earliest occurring time step among the associated time steps is defined as the initial associated time step; that is, the time point corresponding to the initial associated time step is earlier than the other associated time steps besides the initial associated time step.
[0090] Furthermore, the multiple attribution strategies include: last click attribution strategy, first click attribution strategy, linear attribution strategy, and time decay attribution strategy; the time point of each associated time step is earlier than the time step to be processed.
[0091] The last-click attribution strategy can be understood as a strategy that attributes all the value generated by a conversion (conversion return parameter) entirely to the time step of the last click on the ad before the conversion occurred. For example, attributing all the value generated by the conversion to the pending time step. In this case, the attribution return parameter corresponding to the associated time step is zero.
[0092] First-click attribution strategy can be understood as a strategy that attributes all the value generated by a conversion to the time step of the first ad click on that conversion path. For example, it attributes all the value generated by the conversion to the initial associated time step. Based on this, the attribution reward parameter for time steps other than the initial associated time step is zero.
[0093] A linear attribution strategy can be understood as a strategy that distributes the total value generated by a transformation equally among all relevant time steps (including all associated time steps and pending time steps) along the transformation path.
[0094] The time decay attribution strategy can be understood as a strategy that allocates total value to each relevant time step based on the proximity of the reach time point to the conversion time point, using exponential decay or other decreasing functions as weights. The closer the reach time point is to the conversion, the higher the weight it receives.
[0095] The weighting information can be understood as a set of coefficients used to calculate the proportion of value that should be allocated to each time step under a time decay attribution strategy. This set of coefficients is a function of time, and generally satisfies the following conditions: the closer to the transformation time step, the greater the weight; and the sum of all weights is 1.
[0096] Based on the above, and using the initial attribution strategy and the conversion reward parameters, the attribution reward parameters corresponding to the initial attribution strategy for the time step to be processed are determined, including: When the initial attribution strategy is the last click attribution strategy, the time step to be processed is determined to be the conversion time step, and the conversion reward parameter is assigned to the conversion time step. When the initial attribution strategy is the first click attribution strategy, the attribution reward parameter for the time step to be processed is determined to be zero. When the initial attribution strategy is the linear attribution strategy, the conversion return parameter is evenly distributed to the time step to be processed and each associated time step; When the initial attribution strategy is the time decay attribution strategy, the allocation weight information corresponding to the time step to be processed and each associated time step is determined according to the time order, and the conversion reward parameter is allocated to the time step to be processed and each associated time step according to the allocation weight information.
[0097] In the above content, it is important to understand that when the time step to be processed is the time step in which the transformation occurs, the time step to be processed can be understood as the transformation time step.
[0098] To facilitate understanding, based on the above content, the attribution reward parameters corresponding to the initial attribution strategy at the time step to be processed are determined, and the following examples are provided in a)-d).
[0099] In one example, suppose a user converts at time step 3, and the conversion reward parameter is 10. Tracing back to the associated time steps corresponding to time step 3 where the conversion occurred, we obtain associated time step 1 and associated time step 2. Based on this, according to the chronological order, the path is associated time step 1, associated time step 2, and the pending time step 3.
[0100] a) According to the last click attribution strategy, the conversion reward parameter is allocated between the pending time step and each associated time step as follows: the conversion reward parameter is allocated to the pending time step, resulting in an attribution reward parameter of 10 for the pending time step. In this case, the attribution reward parameter for each associated time step is 0.
[0101] b) According to the initial click attribution strategy, the conversion reward parameter is allocated between the pending time step and each associated time step as follows: the conversion reward parameter is allocated to associated time step 1, resulting in an attribution reward parameter of 10 for associated time step 1. In this case, the attribution reward parameters for associated time step 2 and pending time step 3 are 0.
[0102] c) According to the linear attribution strategy, the distribution of the conversion return parameter among the pending time step and each associated time step is as follows: The number of associated time steps and pending time steps is determined to be 3. Based on this, the linear attribution weights of the linear attribution strategy are set as follows: And based on the linear attribution weights, the attribution reward parameters corresponding to each time step are determined.
[0103] Where pi represents the linear attribution weight corresponding to each time step, and N represents the number of time steps to be processed and associated time steps that have undergone transformation when transformation occurs.
[0104] For example, in the scenario above, given that the number of time steps is 3, the attribution weight for linear attribution is determined to be 0.333. Based on this, the attribution reward parameter corresponding to the time step to be processed is 3.33, the attribution reward parameter corresponding to associated time step 1 is 3.33, and the attribution reward parameter corresponding to associated time step 2 is 3.33.
[0105] d) According to the time decay attribution strategy, the conversion return parameter is allocated between the time step to be processed and each associated time step as follows: the weight of each time step is determined according to the time order, where the weight of the time step to be processed 3 is greater than the weight of the associated time step 2. The weight of the associated time step 2 is greater than the weight of the associated time step 1. For example, the weight allocation for the time decay attribution strategy is set as follows:
[0106] Where pi represents the assigned weight for each time step, and N represents the number of time steps to be processed and associated time steps that have undergone transformation. This is a hyperparameter, set to 0.5.
[0107] Based on the above, the allocation weights corresponding to time decay attribution are obtained.
[0108] Based on the above, we can obtain the corresponding attribution reward labels under different attribution strategies.
[0109] S1066. Calculate the multi-factor loss corresponding to the time step to be processed based on the attribution reward label and predicted reward parameter corresponding to the time step to be processed.
[0110] Multi-attribution loss can be understood as calculating the loss component of the difference between the model's predicted return parameters and one or more attribution return labels for a single time step. It can be understood as single-time-step multi-attribution loss.
[0111] In one specific embodiment provided in this specification, the predicted return parameters of the model output for the current time step are obtained. The model's predicted return parameters are compared with the attribution return labels. The error between the two (e.g., mean squared error) is calculated.
[0112] Furthermore, for multiple attribution reward labels, the loss between the predicted reward parameters obtained by the model and each attribution reward label is calculated, thereby obtaining the multi-attribution factor loss corresponding to that time step.
[0113] S1068. Determine the multi-attribution loss based on the multi-attribution loss corresponding to each time step to be processed.
[0114] Among them, multi-attribution loss can be understood as the final loss value obtained by aggregating (such as summing or averaging) the multi-attribution losses calculated for all time steps to be processed over the entire training sequence.
[0115] In one specific embodiment provided in this specification, the multi-attribution loss for all time steps marked as unprocessed time steps (i.e., all time points where transformation occurs) in a training sequence data is summarized. The summarization method can be simple summation, arithmetic mean, or a weighted average based on transformation value. The summarized result is the multi-attribution loss for the reward parameter of this training sample.
[0116] For example, based on the above, the attribution reward labels for the time step to be processed under different attribution strategies are obtained. These attribution reward labels are then compared with the predicted reward parameters to obtain the loss functions corresponding to different attribution strategies. These loss functions are then weighted and fused to obtain the multi-attribution loss determined based on different attribution strategies.
[0117] Furthermore, for example, with Identify multi-attribution losses. Among them, These are the weights of the loss functions for different attribution strategies, summing to 1. The predicted return parameters obtained from the characterization model are... The attribution reward labels are represented according to each attribution strategy.
[0118] According to a specific implementation provided in this specification, the calculation of multi-attribution loss relies on one or more logically valid attribution rules defined by the user. Compared with directly fitting noisy and biased raw data, this reduces the risk of overfitting the model to random and biased patterns in historical data, thereby making the subsequently predicted action parameters more accurate.
[0119] In one specific embodiment provided in this specification, the quantile loss is determined in the following manner: Determine the target time step and the corresponding state parameters, wherein the target time step is any one of the time steps in the training sequence data; Based on the state parameters corresponding to the target time step, determine the upper limit of the remaining parameters in the training sequence data, and obtain the remaining parameters corresponding to the target time step; The surplus parameters are input into the initial decision model to obtain the upper limit of the predicted surplus parameters output by the initial decision model; Based on the upper limit of the surplus parameter and the upper limit of the predicted surplus parameter, the quantile loss is determined.
[0120] The target time step can be understood as a reference time point selected from the training dataset for loss calculation.
[0121] The surplus parameter can be understood as the cumulative return that can be obtained from the current time step to the end of the sequence.
[0122] The upper limit of the balance parameter can be understood as the higher level of return observed in historical data under a given state.
[0123] The upper limit of the predicted surplus parameter can be understood as the output of the decision model's estimate of the upper limit of the surplus parameter.
[0124] Quantile loss can be understood as a loss function that measures the difference between the predicted value and the actual high return level.
[0125] Specifically, in one embodiment provided in this specification, in order to enable the initial decision model to explore potentially higher residual parameters, this specification estimates the upper limit of the residual parameters under specific state parameters based on a loss function of quantile regression. For example: .
[0126] Here, 'a' is a hyperparameter. When 'a' is close to 0, the initial decision model tends to maximize the surplus parameter, thus obtaining an upper limit value for the surplus parameter. Furthermore, to prevent the upper limit value of the surplus parameter from deviating from the surplus parameter, a hyperparameter is introduced... This ensures the accuracy of the initial decision model's prediction of the remaining parameters.
[0127] It should be understood that the aforementioned determination of the upper limit of the residual parameter in the training sequence data based on the state parameters corresponding to the target time step can be understood as follows: based on the state parameters corresponding to the target time step, determine the state parameters of other time steps in the training sequence data that are similar to the state parameters corresponding to the target time step; and then determine the residual parameters corresponding to other time steps based on the state parameters of these other time steps. The upper limit of the residual parameter is determined based on the residual parameters corresponding to the target time step and other time steps, serving as the prediction basis for the quantile regression loss function in predicting the upper limit of the residual parameter. Based on this, using the initial decision model, the predicted residual parameter is predicted based on the residual parameter corresponding to the target time step, thereby determining the quantile regression loss under the state parameters corresponding to the target time step, and thus guiding the model to explore higher residual parameters.
[0128] Step 108: Based on the target loss of the initial decision model, adjust the parameters of the initial decision model to obtain the decision model.
[0129] In one specific implementation provided in this specification, the backpropagation algorithm is used to update all trainable parameters in the initial decision model based on the gradient of the target loss. Through iterative data processing, a well-trained decision model capable of making better bidding decisions is finally obtained.
[0130] Specifically, based on the target loss of the initial decision model, the parameters of the initial decision model are adjusted to obtain the decision model, including: Based on the target loss of the initial decision model, the parameters of the initial decision model are adjusted to obtain the target initial decision model.
[0131] Based on the target initial decision model and the training sequence data, new sequence data are determined.
[0132] Furthermore, based on the initial decision model and the training sequence data, new sequence data is determined, including: Based on the perturbation parameters and the target initial decision model, the residual parameters in the training sequence data are adjusted to obtain the simulation sequence data; The simulated sequence data and the training sequence data are linearly interpolated to obtain the new sequence data.
[0133] Specifically, based on the perturbation parameters, the residual parameters in the training sequence data are perturbed, and simulation sequence data is generated using the target initial decision model based on the perturbed residual parameters. The simulation sequence data and the training sequence data are then mixed together using a mixup method to obtain new sequence data.
[0134] For example, new sequence data can be determined as follows:
[0135] in, For difference parameters, Characterize training sequence data, Characterizing the simulated sequence data, Characterize newly added sequence data.
[0136] It is important to understand that, since the number of times the bidding frequency can be modified is limited, and each modification to the bidding frequency can easily lead to changes in behavioral parameters, the number of times the perturbation parameter can be modified is set to a fixed value during the creation of simulation sequence data.
[0137] Based on the newly added sequence data, determine the new loss corresponding to the target initial decision model.
[0138] Based on the new loss and the target loss, the parameters of the initial decision model for the target are adjusted to obtain the decision model.
[0139] In a specific embodiment provided in this specification, based on the above, the target loss is obtained by weighted summation of action loss, multi-attribution loss, and quantile loss: Based on the target loss, the initial decision model for the target is obtained.
[0140] in, The action loss is represented by Lr, which guarantees the multi-attribution loss determined by different attribution strategies, and Le represents the quantile loss determined based on the state parameters.
[0141] Building upon the above, new sequence data is constructed using the initial target decision model and training sequence data. It's important to understand that the distribution of the new sequence data is the same as that of the training sequence data. Based on the new sequence data, the initial target decision model is trained to obtain the corresponding new loss. The parameters of the initial target decision model are then adjusted based on the new loss and the target loss to obtain the final decision model.
[0142] For example, the decision model can be obtained in the following way:
[0143] Where, when time step data i comes from hour, The value is 1 when time step data i comes from hour, The value is 0.1, and N represents the number of time steps in the training sequence data.
[0144] It should be understood that the method of training the initial decision model of the target based on the newly added sequence data is the same as the method of training the initial decision model of the target based on the training data, and will not be repeated in this manual.
[0145] According to a specific implementation provided in this specification, the model directly outputs predicted action parameters and performs end-to-end optimization using a target loss that includes action losses. This avoids the phenomenon of errors in the prediction stage being amplified in subsequent optimization stages, improving the accuracy and consistency of decision-making. Based on multi-attribution loss, the model can predict the predicted return parameters based on multiple return parameters. This avoids the bias caused by using a single-attribution loss, enhances the model's ability to learn robust patterns from sparse and biased data, and improves generalization and robustness. A quantile loss is introduced and applied to the predicted residual parameter information (RTG). This encourages the model to explore strategies that bring higher long-term returns (upper bound of RTG) in complex, non-convex advertising revenue environments, helping to discover globally better bidding strategies.
[0146] The following is in conjunction with the appendix Figure 2 Taking the application of the decision-making model training method provided in this manual to decision-making based on action parameters as an example, the decision-making method will be further explained. Among other things, Figure 2 A flowchart illustrating the processing steps of a decision-making method provided in one embodiment of this specification is shown, specifically including the following steps.
[0147] Step 202: Determine the decision-making time step and obtain the current state parameters, historical sequence data and current balance parameter information corresponding to the decision-making time step.
[0148] Step 204: Based on the current state parameters, historical sequence data, and current balance parameter information, determine the sequence data corresponding to the time step to be decided.
[0149] Step 206: Input the sequence data into the decision model to obtain the predicted action parameters of the next time step associated with the time step to be decided, output by the decision model.
[0150] In one specific implementation provided in this specification, the trained decision-making model is deployed online to generate bidding actions in real time based on the current state and the objective.
[0151] The decision model receives the state parameters corresponding to the current time step, constructs a sequence of data based on historical state parameters, action parameters, and RTG parameters, and outputs the bidding action for the next time step. For example, it combines historical sequences with the state parameters corresponding to the current time step, the action parameters corresponding to the previous time step, and the current RTG to form a sequence. This sequence is then input into the trained decision model to obtain predicted action parameters. These predicted action parameters are then sent to the media platform for ad delivery.
[0152] After the advertisement is delivered, the new state parameters, action parameters, and updated RTG parameters are added to the historical sequence to prepare for the next inference.
[0153] It is important to understand that, due to the delay in media data, after obtaining the delayed effect data, it is necessary to update the rewards and RTG in the sequence and recalculate the multi-attribution rewards.
[0154] To facilitate understanding of the training methods and decision-making methods of the decision-making models mentioned above in this manual, this manual uses advertising data as training sequence data for explanation.
[0155] In one specific embodiment provided in this specification, historical advertising data is collected, sequence data consisting of state features, action features, and reward features is constructed, and multi-attribution rewards are calculated.
[0156] Specifically, this involves acquiring advertising data from external websites; and constructing status characteristics, action characteristics, and return characteristics based on this advertising data. This includes a)-c): a) The state reflects the current environment of ad delivery, and the constructed state features include: Budget-related factors: remaining budget, budget consumption rate, and budget consumption percentage, etc.
[0157] Time-related information: current time period, day of the week, deployment progress (ratio of deployed time to total planned time), etc.
[0158] Environmental factors include: intensity of competition (reflected by indicators such as win rate or number of bids), and traffic quality.
[0159] b) Action characteristics reflect bid configuration, such as target CPA bid value. It's important to understand that in historical data, actions represent the actual bid value set.
[0160] c) Return-to-Go (RTG): The reward can be understood as the immediate return obtained at each time step, such as the GMV (Gross Merchandise Volume) generated at that time step. It is important to understand that due to data latency, the reward data may not be known until several hours after the action occurs.
[0161] RTG can be understood as the cumulative reward from the current time to the end of the sequence. For each time step, there is the following relationship between RTG and the cumulative reward: Where T is the time step at which the sequence ends, and RTG represents the expected total return from the current moment to the end.
[0162] Sequence Construction: States, actions, and RTGs are organized into a sequence in chronological order. Each sequence sample takes the following form:
[0163] Each sequence sample contains multiple time steps, and each time step includes a triplet of state, action, and reward. The sequence length is dynamically adjusted according to the business scenario and typically covers a complete campaign cycle or user conversion path.
[0164] Furthermore, multi-attribution return calculation can be understood as follows: for each user who converts, all the time points of ad touchpoints before the conversion are traced back, and the conversion value is redistributed using four attribution strategies (last click, first click, linear, and time decay).
[0165] It should be understood that the four attribution strategies are detailed in the embodiments described above in this specification, and will not be repeated here.
[0166] Based on this, four reward parameters corresponding to the attribution strategies are generated for each time step (each reach), and these parameters are used as attribution reward labels for multi-attribution learning.
[0167] Construct a DT model. This DT model includes an embedding layer corresponding to state features, action features, and reward features, a transformer encoder, and multiple output heads.
[0168] The output heads include an action prediction head, a multi-attribution reward prediction head, and a quantile regression head.
[0169] In the embedding layer, the state, action, and RTG are mapped to hidden dimensions, respectively. Since the state may be a multi-dimensional vector, a linear layer is used for embedding; actions and RTG are scalars and are also embedded through linear layers. In the transformer encoder, the embedded sequence is input. Above the transformer output, four independent linear layers are connected to predict the reward values corresponding to the four attribution strategies. This results in four predicted reward parameters output by the multi-attribution reward prediction head.
[0170] Furthermore, the quantile regression head is used to predict the upper limit of the RTG (i.e., to predict the remaining parameter information). The action prediction head is used to predict the next action (bid).
[0171] In this specification, the quantile loss is determined based on the upper limit of the predicted RTG obtained from the quantile regression head, and the action loss is determined based on the predicted action parameters obtained from the action prediction head.
[0172] It is important to understand that this specification extracts features from the hidden layers of the four attribution reward prediction heads (i.e., features before the linear layers of the attribution heads), concatenates them, and then fuses them through an MLP (Multilayer Perceptron). The fused features are then added to the backbone features of the Transformer (residual connection) for the final action prediction and quantile regression prediction.
[0173] Furthermore, by combining action prediction loss, multi-attribution reward loss, and quantile regression loss, the target loss is determined.
[0174] It should be understood that the method for determining the target loss involved in this specification is detailed in the embodiments mentioned above, and will not be repeated here.
[0175] Based on the above, a pre-trained base model is used as a data simulator to perturb the RTG in the advertising data. Then, the pre-trained base model is used to generate new action sequences. Sequences in the generated new action sequences with action parameter changes not exceeding 10% and an RTG higher than that in the advertising data are selected as simulation sequence data.
[0176] Linear interpolation is performed on the state, action, and RTG in the sequence data corresponding to the advertising data and the simulation sequence data to obtain the new sequence data.
[0177] The weight of the newly added sequence data is set to 0.1, and the weight of the sequence data corresponding to the advertising data is set to 1, thus obtaining the decision model.
[0178] Corresponding to the above method embodiments, this specification also provides embodiments of a training device for a decision model. Figure 3 A schematic diagram of a training apparatus for a decision model provided in one embodiment of this specification is shown. Figure 3 As shown, the device includes: The acquisition unit 302 is configured to acquire training sequence data, wherein the training sequence data consists of multiple time step data, and each time step data includes action parameters and reward parameters; The prediction unit 304 is configured to input the training sequence data into an initial decision model to obtain the prediction sequence data output by the initial decision model. The prediction sequence data includes prediction action parameters, prediction reward parameters, and prediction surplus parameter information corresponding to each time step data. The prediction reward parameters are predicted based on multiple reward parameters. The determining unit 306 is configured to determine the target loss of the initial decision model based on the training sequence data and the predicted sequence data, wherein the target loss includes action loss, multi-attribution loss and quantile loss, the action loss is determined based on action parameters and predicted action parameters, and the multi-attribution loss and quantile loss are determined based on reward parameters, predicted reward parameters and predicted surplus parameter information. The adjustment unit 308 is configured to adjust the parameters of the initial decision model based on the target loss of the initial decision model to obtain a decision model.
[0179] Furthermore, multi-attribution loss is determined in the following manner: Determine the time step to be processed and at least one associated time step corresponding to the time step to be processed; Based on the time step to be processed and the reward parameters corresponding to each associated time step, determine the attribution reward label corresponding to the time step to be processed; Calculate the multi-factor loss corresponding to the time step to be processed based on the attribution reward label and predicted reward parameter corresponding to the time step to be processed; The multi-attribution loss is determined based on the multi-attribution factor loss corresponding to each time step to be processed.
[0180] Furthermore, the determining unit 306 is further configured as follows: The time step corresponding to the target behavior information is determined as the time step to be processed. Determine the associated behavior information corresponding to the target behavior information, and determine at least one associated time step based on the associated behavior information.
[0181] Furthermore, the determining unit 306 is further configured as follows: The conversion reward parameters are determined based on the time step to be processed and the reward parameters corresponding to each associated time step. Based on multiple attribution strategies and the conversion return parameters, the attribution return parameters corresponding to each attribution strategy for the time step to be processed are determined. Attribution reward labels are generated for the time steps to be processed based on each attribution reward parameter.
[0182] Furthermore, the determining unit 306 is further configured as follows: Determine an initial attribution strategy, wherein the initial attribution strategy is any one of a plurality of attribution strategies; Based on the initial attribution strategy and the conversion reward parameters, the attribution reward parameters corresponding to the initial attribution strategy are determined for the time step to be processed.
[0183] Furthermore, the multiple attribution strategies include: last click attribution strategy, first click attribution strategy, linear attribution strategy, and time decay attribution strategy; the time point of each associated time step is earlier than the time step to be processed. Unit 306 is further configured as follows: When the initial attribution strategy is the last click attribution strategy, the time step to be processed is determined to be the conversion time step, and the conversion reward parameter is assigned to the conversion time step. When the initial attribution strategy is the first click attribution strategy, the attribution reward parameter for the time step to be processed is determined to be zero. When the initial attribution strategy is the linear attribution strategy, the conversion return parameter is evenly distributed to the time step to be processed and each associated time step; When the initial attribution strategy is the time decay attribution strategy, the allocation weight information corresponding to the time step to be processed and each associated time step is determined according to the time order, and the conversion reward parameter is allocated to the time step to be processed and each associated time step according to the allocation weight information.
[0184] Furthermore, prediction unit 304 is further configured as follows: The training sequence data is input into the initial decision model to obtain at least one attribution reward parameter corresponding to each time step determined by the initial decision model; Based on at least one attribution reward parameter corresponding to each time step, predict the predicted reward parameter corresponding to each time step. Multiple predicted return parameters are fused to obtain fused features, and action prediction parameters and predicted surplus parameters are generated based on the fused features.
[0185] Furthermore, the time step data also includes state parameters; the quantile loss is determined through the following steps: Determine the target time step and the corresponding state parameters, wherein the target time step is any one of the time steps in the training sequence data; Based on the state parameters corresponding to the target time step, determine the upper limit of the remaining parameters in the training sequence data, and obtain the remaining parameters corresponding to the target time step; The surplus parameters are input into the initial decision model to obtain the upper limit of the predicted surplus parameters output by the initial decision model; Based on the upper limit of the surplus parameter and the upper limit of the predicted surplus parameter, the quantile loss is determined.
[0186] Furthermore, prediction unit 304 is further configured as follows: Based on the target loss of the initial decision model, the parameters of the initial decision model are adjusted to obtain the target initial decision model; Based on the target initial decision model and the training sequence data, new sequence data is determined; Based on the newly added sequence data, determine the new loss corresponding to the target initial decision model; Based on the new loss and the target loss, the parameters of the initial decision model for the target are adjusted to obtain the decision model.
[0187] Furthermore, prediction unit 304 is further configured as follows: Based on the perturbation parameters and the target initial decision model, the residual parameters in the training sequence data are adjusted to obtain the simulation sequence data; The simulated sequence data and the training sequence data are linearly interpolated to obtain the new sequence data.
[0188] Corresponding to the above method embodiments, this specification also provides embodiments of a decision-making device. Figure 4 A schematic diagram of a decision-making device according to one embodiment of this specification is shown. Figure 4 As shown, the device includes: The first determining unit 402 is configured to determine the decision-making time step and obtain the current state parameters, historical sequence data and current balance parameter information corresponding to the decision-making time step. The second determining unit 404 is configured to determine the sequence data corresponding to the time step to be decided based on the current state parameters, historical sequence data and current balance parameter information; The prediction unit 406 is configured to input the sequence data into the decision model to obtain the predicted action parameters of the next time step associated with the time step to be decided, output by the decision model.
[0189] The above is an illustrative scheme of a training device and a decision-making device for a decision model according to this embodiment. It should be noted that the technical solutions of the training device and the decision-making device for this decision model belong to the same concept as the technical solutions of the training method and the decision-making method of the decision model described above. For details not described in detail in the technical solutions of the training device and the decision-making device for the decision model, please refer to the description of the technical solutions of the training method and the decision-making method of the decision model described above.
[0190] Figure 5 A structural block diagram of a computing device according to an embodiment of this application is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.
[0191] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0192] In one embodiment of this application, the aforementioned components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.
[0193] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 500 can also be a mobile or stationary server.
[0194] The processor 520 is used to execute the following computer program or instructions, which, when executed by the processor, implement the steps of the training method and decision method of the above-mentioned decision model.
[0195] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the technical solutions of the training method and decision method of the decision model described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the training method and decision method of the decision model described above.
[0196] An embodiment of this specification also provides a computer-readable storage medium storing a computer program or instructions that, when executed by a processor, implement the steps of the training method and decision method of the above-described decision model.
[0197] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are described simply because they are substantially similar to the training method and decision method embodiments of the decision model; relevant parts can be referred to the descriptions of the training method and decision method embodiments of the decision model.
[0198] An embodiment of this specification also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the training method and decision method of the above-described decision model.
[0199] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solutions of the above-described training method and decision method of the decision model. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solutions of the above-described training method and decision method of the decision model.
[0200] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0201] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0202] It should be noted that the above description describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0203] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0204] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for training a decision model, comprising: Acquire training sequence data, wherein the training sequence data consists of multiple time step data, each time step data includes action parameters and reward parameters, and the action parameters are determined based on the bid corresponding to each time step; The training sequence data is input into the initial decision model to obtain the predicted sequence data output by the initial decision model. The predicted sequence data includes the predicted action parameters, predicted reward parameters, and predicted surplus parameter information corresponding to each time step data. The predicted reward parameters are predicted based on multiple reward parameters. Based on the training sequence data and the prediction sequence data, the target loss of the initial decision model is determined, wherein the target loss includes action loss, multi-attribution loss and quantile loss, the action loss is determined based on action parameters and predicted action parameters, and the multi-attribution loss and quantile loss are determined based on reward parameters, predicted reward parameters and predicted surplus parameters. Based on the target loss of the initial decision model, the parameters of the initial decision model are adjusted to obtain the final decision model.
2. The method of claim 1, wherein the multi-attribution loss is determined in the following manner: Determine the time step to be processed and at least one associated time step corresponding to the time step to be processed; Based on the time step to be processed and the reward parameters corresponding to each associated time step, determine the attribution reward label corresponding to the time step to be processed; Calculate the multi-factor loss corresponding to the time step to be processed based on the attribution reward label and predicted reward parameter corresponding to the time step to be processed; The multi-attribution loss is determined based on the multi-attribution factor loss corresponding to each time step to be processed.
3. The method of claim 2, wherein determining the time step to be processed and at least one associated time step corresponding to the time step to be processed comprises: The time step corresponding to the target behavior information is determined as the time step to be processed. Determine the associated behavior information corresponding to the target behavior information, and determine at least one associated time step based on the associated behavior information.
4. The method as described in claim 2, wherein determining the attribution reward label corresponding to the time step to be processed based on the reward parameters corresponding to the time step to be processed and each associated time step, includes: The conversion reward parameters are determined based on the time step to be processed and the reward parameters corresponding to each associated time step. Based on multiple attribution strategies and the conversion return parameters, the attribution return parameters corresponding to each attribution strategy for the time step to be processed are determined. Attribution reward labels are generated for the time steps to be processed based on each attribution reward parameter.
5. The method of claim 4, wherein determining the attribution reward parameter corresponding to each attribution strategy for the time step to be processed, based on multiple attribution strategies and the conversion reward parameter, includes: Determine an initial attribution strategy, wherein the initial attribution strategy is any one of a plurality of attribution strategies; Based on the initial attribution strategy and the conversion reward parameters, the attribution reward parameters corresponding to the initial attribution strategy are determined for the time step to be processed.
6. The method of claim 5, wherein the plurality of attribution strategies include: Last-click attribution strategy, first-click attribution strategy, linear attribution strategy, time decay attribution strategy; The time points of each associated time step are earlier than the time step to be processed; Based on the initial attribution strategy and the conversion reward parameters, the attribution reward parameters corresponding to the initial attribution strategy for the time step to be processed are determined, including: When the initial attribution strategy is the last click attribution strategy, the time step to be processed is determined to be the conversion time step, and the conversion reward parameter is assigned to the conversion time step. When the initial attribution strategy is the first click attribution strategy, the attribution reward parameter for the time step to be processed is determined to be zero. When the initial attribution strategy is the linear attribution strategy, the conversion return parameter is evenly distributed to the time step to be processed and each associated time step; When the initial attribution strategy is the time decay attribution strategy, the allocation weight information corresponding to the time step to be processed and each associated time step is determined according to the time order, and the conversion reward parameter is allocated to the time step to be processed and each associated time step according to the allocation weight information.
7. The method of claim 1, wherein the training sequence data is input into an initial decision model to obtain the predicted sequence data output by the initial decision model, comprising: The training sequence data is input into the initial decision model to obtain at least one attribution reward parameter corresponding to each time step determined by the initial decision model; Based on at least one attribution reward parameter corresponding to each time step, predict the predicted reward parameter corresponding to each time step. Multiple predicted return parameters are fused to obtain fused features, and action prediction parameters and predicted surplus parameters are generated based on the fused features.
8. The method of claim 7, wherein the time step data further includes state parameters; the quantile loss is determined by the following steps: Determine the target time step and the corresponding state parameters, wherein, The target time step is any one of the time steps in the training sequence data; Based on the state parameters corresponding to the target time step, determine the upper limit of the remaining parameters in the training sequence data, and obtain the remaining parameters corresponding to the target time step; The surplus parameters are input into the initial decision model to obtain the upper limit of the predicted surplus parameters output by the initial decision model; Based on the upper limit of the surplus parameter and the upper limit of the predicted surplus parameter, the quantile loss is determined.
9. The method of claim 8, wherein the parameters of the initial decision model are adjusted based on the target loss of the initial decision model to obtain a decision model, comprising: Based on the target loss of the initial decision model, the parameters of the initial decision model are adjusted to obtain the target initial decision model; Based on the target initial decision model and the training sequence data, new sequence data is determined; Based on the newly added sequence data, determine the new loss corresponding to the target initial decision model; Based on the new loss and the target loss, the parameters of the initial decision model for the target are adjusted to obtain the decision model.
10. The method of claim 9, wherein determining new sequence data based on the target initial decision model and the training sequence data includes: Based on the perturbation parameters and the target initial decision model, the residual parameters in the training sequence data are adjusted to obtain the simulation sequence data; The simulated sequence data and the training sequence data are linearly interpolated to obtain the new sequence data.
11. A decision-making method, the method comprising: Determine the decision-making time step and obtain the current state parameters, historical sequence data, and current balance parameter information corresponding to the decision-making time step; Based on the current state parameters, historical sequence data, and current balance parameter information, determine the sequence data corresponding to the time step to be decided; The sequence data is input into the decision model to obtain the predicted action parameters of the next time step associated with the time step to be decided, as output by the decision model. The decision model is trained based on the training method of the decision model according to any one of claims 1 to 10, and the predicted action parameters are determined based on the bid corresponding to the next time step.
12. A computing device, comprising: Memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 11.
13. A computer-readable storage medium storing a computer program or instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.
14. A computer program product comprising a computer program or instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Programmed advertisement putting method, system, device and equipment and storage medium
CN113222656A
Exposure attribution method and device for advertisement return
CN120258909A