Advertisement recommendation method and device, storage medium and electronic equipment
By dividing user and ad features into long-term and short-term based on time attributes, and using learnable concatenation weights and recommendation loss functions for optimization, the problems of exposure bias and static loss weights in the ad recommendation model are solved, resulting in more accurate recommendation effects and improved user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2026-02-13
- Publication Date
- 2026-04-21
AI Technical Summary
Existing advertising recommendation models rely on systematic biases in historical exposure data and static loss weights, making it difficult to dynamically balance multiple business objectives, thus limiting recommendation effectiveness.
By dividing user and advertising features into long-term and short-term features according to time attributes, and fusing them using learnable concatenation weights, a recommendation loss function is designed, which combines exposure correction terms and multi-business loss terms to optimize the network parameters of the recommendation model.
It significantly improved the matching degree between recommendation results and user needs, enhanced the user experience, strengthened the distinguishability of feature expression, dynamically balanced multiple business objectives, and resisted exposure bias.
Smart Images

Figure CN121903702A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and content recommendation technology, and can be applied to fintech scenarios. Specifically, this application relates to an advertising recommendation method, apparatus, storage medium, and electronic device. Background Technology
[0002] The core objective of ad recommendation is to efficiently connect user needs with ad information, thereby improving the matching efficiency and commercial value between the two.
[0003] Recommendation models employing multi-task learning are widely used to simultaneously optimize multiple business objectives such as click-through rate and conversion rate. However, the historical exposure data relied upon for training these models inherently suffers from systematic exposure selection bias, impacting the model's ability to understand and match user needs. Furthermore, the loss weights for each business objective in multi-task learning are typically statically preset, making it impossible for the model optimization process to dynamically balance different objectives or consistently align with users' comprehensive needs or long-term value, thus limiting further improvements in recommendation performance. Summary of the Invention
[0004] This application provides an advertising recommendation method, apparatus, storage medium, and electronic device that improves the matching degree between recommendation results and user needs, thereby improving user experience.
[0005] According to a first aspect of this application, an advertising recommendation method is provided, the method comprising: Based on the ad recommendation request for the target user, determine the candidate ads to be recommended; Based on the time attribute of the business objective, the user characteristics corresponding to the target user and the advertising characteristics corresponding to the candidate advertisement are divided into long-term characteristics and short-term characteristics. The long-term features and the short-term features are concatenated using concatenation weights to obtain comprehensive features, and the comprehensive features are then input into a pre-trained recommendation model. The recommendation model outputs predicted values of the candidate advertisements for each business objective based on the comprehensive features, and determines the target advertisement for the target user from the candidate advertisements based on the predicted values. The network parameters of the recommendation model are obtained by optimizing the recommendation loss function. The recommendation loss function includes an exposure correction term and at least two business loss terms. The exposure correction term is used to correct exposure selection bias and optimize the splicing weights. The weight combination corresponding to the business loss terms is determined based on the correlation between the evaluation indicators of each business objective on offline verification data and the revenue indicators of online business.
[0006] According to a second aspect of this application, an advertising recommendation device is provided, the device comprising: The candidate ad determination module is used to determine candidate ads to be recommended based on ad recommendation requests for target users; The feature segmentation module is used to divide the user features corresponding to the target user and the advertising features corresponding to the candidate advertisement into long-term features and short-term features based on the time attribute of the business objective. The feature concatenation module is used to concatenate the long-term features and the short-term features using concatenation weights to obtain comprehensive features, and input the comprehensive features into a pre-trained recommendation model; The target advertisement determination module is used to output the predicted value of the candidate advertisement on each business objective based on the comprehensive features through the recommendation model, and determine the target advertisement for the target user from the candidate advertisements based on the predicted value; The network parameters of the recommendation model are obtained by optimizing the recommendation loss function. The recommendation loss function includes an exposure correction term and at least two business loss terms. The exposure correction term is used to correct exposure selection bias and optimize the splicing weights. The weight combination corresponding to the business loss terms is determined based on the correlation between the evaluation indicators of each business objective on offline verification data and the revenue indicators of online business.
[0007] According to a third aspect of the present invention, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the advertising recommendation method as described in embodiments of this application.
[0008] According to a fourth aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the advertising recommendation method as described in the embodiments of the present application.
[0009] According to a fifth aspect of this application, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the advertising recommendation method as described in embodiments of this application.
[0010] This technical solution, by dividing user and advertising features into long-term and short-term features based on time attributes and employing learnable concatenation weights for adaptive fusion, can more precisely characterize users' stable interests and real-time intentions, enhancing the discriminative power of feature representation. More importantly, the designed recommendation loss function creatively combines an exposure correction term with a multi-business loss term: the exposure correction term, based on counterfactual inference principles, not only effectively eliminates selection bias in historical exposure data, enabling the recommendation model to learn users' true interests and preferences, but also generates a gradient that serves as the dominant signal driving the optimization of the feature concatenation weights, enabling the recommendation model to learn an optimal feature fusion method that proactively resists exposure bias. Simultaneously, the multi-business loss term is determined and adjusted based on the statistical correlation between offline verification indicators and real online business revenue, ensuring that the multi-objective optimization direction of the recommendation model continuously aligns with the distribution of real online user feedback. In summary, this technical solution, through end-to-end joint optimization, achieves synergistic enhancement of effective feature fusion, data bias correction, and multi-objective dynamic balance, ultimately significantly improving the matching degree between recommendation results and user needs, thereby improving the user experience.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of the advertising recommendation method provided in Embodiment 1; Figure 2 This is a flowchart of the advertising recommendation method provided in Example 2; Figure 3 This is a schematic diagram of the advertising recommendation device provided in Embodiment 3 of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of this application. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0015] It should be noted that the terms "first," "second," "target," and "candidate," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] Example 1 Figure 1 This is a flowchart of the advertising recommendation method provided in Embodiment 1. This embodiment is applicable to the case of advertising recommendation. The method can be executed by an advertising recommendation device, which is implemented in hardware and / or software and can be integrated into an electronic device running this system.
[0017] like Figure 1 As shown, the method includes: S110. Based on the advertising recommendation request for the target user, determine the candidate advertisements to be recommended.
[0018] S120. Based on the time attribute of the business objective, the user characteristics corresponding to the target user and the advertising characteristics corresponding to the candidate advertisement are divided into long-term characteristics and short-term characteristics.
[0019] S130. The long-term features and the short-term features are concatenated using concatenation weights to obtain comprehensive features, and the comprehensive features are input into the pre-trained recommendation model.
[0020] S140. The recommendation model outputs the predicted values of the candidate advertisements for each business objective based on the comprehensive features, and the target advertisement is determined for the target user from the candidate advertisements based on the predicted values.
[0021] The network parameters of the recommendation model are obtained by optimizing the recommendation loss function. The recommendation loss function includes an exposure correction term and at least two business loss terms. The exposure correction term is used to correct exposure selection bias and optimize the splicing weights. The weight combination corresponding to the business loss terms is determined based on the correlation between the evaluation indicators of each business objective on offline verification data and the revenue indicators of online business.
[0022] An ad recommendation request is used to request ad recommendations to a specific target user. The target user refers to the end user corresponding to the ad recommendation request. The candidate ads to be recommended constitute the set of ads entering this recommendation decision-making process. User features describe the attributes and behavioral patterns of the target user, naturally including time-related information such as static user attributes, long-term interest tags, and short-term behavioral sequences. Ad features describe the attributes and contextual state of candidate ads, also including time-related information such as the long-term category affiliation and short-term promotional information of the advertised product.
[0023] Optionally, business objectives include click-through rate, conversion rate, long-term user value, and advertising platform revenue. The time attribute of business objectives refers to the statistical patterns of user feedback models over time, based on which the business objectives to be optimized depend, to guide the granularity and scope of feature segmentation.
[0024] Based on the time-related attributes of business objectives, user characteristics and advertising characteristics can be divided into long-term characteristics and short-term characteristics. Long-term characteristics refer to subsets reflecting stable interests or inherent attributes, such as monthly user interest distribution and advertised product categories. Short-term characteristics, on the other hand, refer to those used to capture recent real-time user intent or the immediate context of advertising, such as click sequences within the current session and real-time contextual information. For business objectives emphasizing real-time feedback, the corresponding short-term characteristics may focus on in-session behavior at the minute or hour level; for business objectives emphasizing long-term value, the long-term characteristics may cover interest accumulation over several weeks or months.
[0025] The concatenation weights are used to weight and combine long-term and short-term features to generate a comprehensive feature. This comprehensive feature effectively integrates the long-term and short-term information represented by both features, serving as a unified feature representation for the input recommendation model. The concatenation weights are a trainable parameter vector for the model, and their introduction allows for the allocation of the proportion of long-term and short-term information in recommendation decisions across different business scenarios. The reason for concatenating long-term and short-term features based on the time attributes of business objectives, rather than directly concatenating user and advertising features, is that directly concatenating user and advertising features fails to distinguish the temporal granularity of the signal, making it difficult for the recommendation model to accurately adapt to different business objectives.
[0026] The recommendation model refers to a computational model built on a neural network used to generate predicted values. The network parameters of the recommendation model have been fixed through offline training and optimization. These network parameters refer to all adjustable weights and biases in the recommendation model. During the inference phase, the recommendation model receives comprehensive features and outputs predicted values. Based on these predicted values, ranking or decision-making is performed, ultimately determining the target advertisement for the target user. Each business objective has a corresponding predicted value, which represents the probability of a candidate advertisement occurring for each business objective, as estimated by the recommendation model.
[0027] The recommendation loss function optimized during the training phase of the recommendation model is crucial for driving model learning. It comprises two synergistic components: a business loss term and an exposure correction term. The business loss term directly calculates the error between the predicted value output by the recommendation model and the feedback labels from real users, aiming to improve prediction accuracy. Its corresponding weight combination is not fixed but determined based on the correlation between the evaluation metrics of each business objective on offline validation data and the revenue metrics of online business. The weight combination quantifies the relative importance of each business loss term. This establishes a statistical mapping between offline model performance and actual online results, thereby dynamically adjusting the training focus to ensure that the model optimization direction remains aligned with the overall online performance goal. The business loss term and its corresponding weight combination establish a statistical mapping between offline model performance and actual online results, thereby dynamically adjusting the training focus to ensure that the model optimization direction remains aligned with the overall online performance goal.
[0028] The recommendation loss function also includes an exposure correction term, which is specifically designed to correct exposure selection bias caused by the non-random exposure of historical logs. The exposure correction term serves a dual optimization purpose: first, it updates the network parameters of the recommendation model to make the predictions closer to the target user's true interests; second, it drives the optimization of the concatenation weights. This makes the learning objective of the feature fusion method, i.e., the concatenation weights, directly target how to better offset data bias, thereby achieving end-to-end joint optimization through feature fusion to remove bias.
[0029] The specific structure of the recommendation model is not limited here; it should be determined based on actual business needs. Optionally, the recommendation model can be a multi-gate hybrid expert model, including a shared expert layer, a gated network layer, and a task tower. The shared expert layer and all gated network layers receive the same input, i.e., the integrated features, in parallel. x Each gated network layer selectively combines expert knowledge for its corresponding task; the task tower layer then makes the final decision based on the combined knowledge. The shared expert layer consists of… N The system consists of several parallel expert subnetworks, each sharing the same input features, i.e., comprehensive features, used to extract diverse low-level feature representations from the comprehensive features; let the first... i Each expert subnetwork has comprehensive features The output vector is Gated network layers do not directly process or transform features, but rather base their processing on the comprehensive features of the current input. To calculate a weight vector for each expert subnetwork, a weight vector is calculated for each corresponding business objective. This weight vector determines the output of each expert subnetwork in the subsequent task tower layer. The degree of contribution to the predicted value.
[0030] The task hierarchy corresponds to each business objective. Each has an independent task tower network; the input of the task tower network consists of all the outputs of the shared expert layer. Based on the corresponding weights output by the gated network layer The weighted summation is used to obtain the task tower network. The weighted feature representation is processed to output the business objective. Predicted value .
[0031] This technical solution, by dividing user and advertising features into long-term and short-term features based on time attributes and employing learnable concatenation weights for adaptive fusion, can more precisely characterize users' stable interests and real-time intentions, enhancing the discriminative power of feature representation. More importantly, the designed recommendation loss function creatively combines an exposure correction term with a multi-business loss term: the exposure correction term, based on counterfactual inference principles, not only effectively eliminates selection bias in historical exposure data, enabling the recommendation model to learn users' true interests and preferences, but also generates a gradient that serves as the dominant signal driving the optimization of the feature concatenation weights, enabling the recommendation model to learn an optimal feature fusion method that proactively resists exposure bias. Simultaneously, the multi-business loss term is determined and adjusted based on the statistical correlation between offline verification indicators and real online business revenue, ensuring that the multi-objective optimization direction of the recommendation model continuously aligns with the distribution of real online user feedback. In summary, this technical solution, through end-to-end joint optimization, achieves synergistic enhancement of effective feature fusion, data bias correction, and multi-objective dynamic balance, ultimately significantly improving the matching degree between recommendation results and user needs, thereby improving the user experience.
[0032] Example 2 Figure 2 This is a flowchart of the advertising recommendation method provided in Embodiment 2. This embodiment further optimizes the above embodiments, specifically by limiting the training method used in the recommendation model.
[0033] like Figure 2 As shown, the method includes: S210. Based on the time attribute of business objectives, the user characteristics of sample users and the advertising characteristics of sample advertisements in the training samples are divided into long-term features and short-term features.
[0034] Training samples refer to historical data records used for model training, where sample users and sample advertisements correspond to user entities and advertisement entities in the historical records, respectively. Each training sample includes one sample user and one sample advertisement. Both user features and advertisement features extracted from the training samples contain time-related information. Based on the time attributes of the business objectives, user features and advertisement features can be divided into long-term features and short-term features.
[0035] S220. The long-term features and short-term features are weighted and concatenated using the concatenation weights to obtain the comprehensive features of the training samples.
[0036] The concatenation weights are used to control the contribution ratio of long-term and short-term features during fusion. Instead of being manually preset to a fixed ratio, the concatenation weights are set as trainable parameter vectors, allowing the resulting integrated features to adaptively balance long-term and short-term information for different users, advertisements, and scenarios, forming a more expressive unified feature representation.
[0037] The comprehensive feature is obtained by concatenating long-term and short-term features according to the concatenation weight. Inputting the comprehensive weight into the recommendation model can provide signal input at different time granularities, enabling the recommendation model to understand both the user's stable preferences and dynamic intentions.
[0038] S230. Input the comprehensive features into the recommendation model to be trained, so as to obtain the predicted value of the recommendation model for each business objective through forward propagation.
[0039] The recommendation model to be trained is a neural network with parameters to be optimized, and the integrated features are the input data of the recommendation model. Forward propagation refers to the unidirectional computation process from the input layer to the output layer of the recommendation model. Its purpose is to perform a series of predefined linear transformations and nonlinear activations on the input data based on all the parameters of the recommendation model, so as to obtain the predicted values of the recommendation model output for each business objective. Forward propagation is a key stage in the training loop for evaluating the current model state and preparing the data needed for parameter updates. Forward propagation, together with the subsequent loss calculation and backpropagation, constitutes a complete training iteration, driving the model parameters to be gradually optimized in the direction of minimizing the loss function.
[0040] S240. Based on the predicted values and sample labels corresponding to each business objective, the short-term features, long-term features, and comprehensive features, calculate the exposure correction term and the at least two business loss terms in the recommendation loss function.
[0041] The sample labels are the real user feedback recorded in the training samples, including at least exposure labels and feedback labels for each business objective. The recommendation loss function is the objective function that guides model optimization. The recommendation loss function includes a business loss term and an exposure correction term.
[0042] The business loss term is determined based on the predicted value corresponding to each business objective and the feedback label of each business objective in the sample labels. The exposure correction term is determined based on the exposure label in the sample labels and the feedback label of each business objective, as well as the short-term feature, long-term feature, and comprehensive feature.
[0043] S250. Based on the recommended loss function, the first gradient generated by the exposure correction term and the second gradient generated by the at least two business loss terms are calculated through backpropagation.
[0044] Backpropagation essentially calculates the gradient direction of the recommendation loss function with respect to all trainable parameters of the recommendation model using the chain rule. Here, all trainable parameters of the recommendation model include the network parameters of the backbone network and the concatenation weights. The first gradient is generated by the exposure correction term, and the second gradient is generated by the business loss term.
[0045] S260. Update the network parameters of the backbone network in the recommendation model using the first gradient and the second gradient, and update the splicing weights using the first gradient.
[0046] The network parameters of the backbone network in the recommendation model are updated by using the first and second gradients together. This means that the network parameters of the backbone network in the recommendation model receive dual signals from both the correction bias and the fitted label. Updating the splicing weights by using the first gradient explicitly binds the learning objective of the splicing weights to the correction bias, so that the learning of feature fusion is directly driven by the correction bias signal. This achieves the joint optimization goal of fundamentally improving the model's robustness to bias by optimizing feature fusion.
[0047] This application's technical solution divides long-term and short-term features based on the time attributes of business objectives, providing the model with a refined information foundation that simultaneously captures users' stable interests and real-time intentions. By introducing learnable concatenation weights for adaptive fusion, the recommendation model can dynamically adjust the contribution ratio of long-term and short-term information for different scenarios, enhancing the adaptability and discriminative ability of feature representation. Crucially, the designed recommendation loss function and specific gradient utilization method are key: the exposure correction term not only directly corrects exposure selection bias in the training data, but its first gradient is also specifically used to update the concatenation weights, forcing the feature fusion method itself to evolve in a direction most resistant to exposure bias; simultaneously, the second gradient generated by the business loss term, together with the first gradient, updates the network parameters of the backbone network in the recommendation model, ensuring the basic accuracy of the recommendation model in multi-objective prediction. This results in a recommendation model that is not only more accurate in prediction, but also whose internal feature representation and decision logic are closer to the user's true interest distribution, thus significantly improving the overall performance of online recommendations.
[0048] In an optional embodiment, based on the predicted values and sample labels corresponding to each business objective, the short-term features, long-term features, and comprehensive features, the exposure correction term in the recommendation loss function is calculated, including: for each business objective, based on the exposure label in the sample labels and the feedback label corresponding to the business objective, determining the counterfactual expectations corresponding to the short-term features, the long-term features, and the comprehensive features respectively, to obtain the short-term counterfactual expectations, long-term counterfactual expectations, and comprehensive counterfactual expectations; for each business objective, calculating the deviation between the predicted value output by the recommendation model for each business objective and the short-term counterfactual expectations, long-term counterfactual expectations, and comprehensive counterfactual expectations respectively, to obtain the short-term deviation, long-term deviation, and comprehensive deviation; weighting the short-term deviation, long-term deviation, and comprehensive deviation, and determining the exposure correction term based on the weighted result.
[0049] The exposure label is a binary label that records whether a sample advertisement in the training sample has been shown to a sample user. The feedback label corresponding to each business objective is a binary label that records whether a sample user has generated the corresponding behavior. The counterfactual expectation refers to the expected probability value of a user generating a certain type of feedback to an advertisement under the ideal assumption that "all relevant advertisements are exposed unbiasedly".
[0050] For each business objective, based on the exposure label in the sample labels and the corresponding feedback label, counterfactual expectations corresponding to short-term features, long-term features, and comprehensive features are determined respectively. This means that for the same business objective, three counterfactual expectations need to be independently estimated using the above two types of labels and based on three different input perspectives: short-term features, long-term features, and comprehensive features. These are referred to as short-term counterfactual expectations, long-term counterfactual expectations, and comprehensive counterfactual expectations, respectively. This approach allows for estimation from different time scales and feature fusion levels, enabling a more detailed identification and removal of exposure biases mixed in with long-term interests, short-term behaviors, and comprehensive decision-making.
[0051] For each business objective, the differences between the predicted value output by the recommendation model for each business objective and the aforementioned three counterfactual expectations are calculated to obtain short-term bias, long-term bias, and comprehensive bias. Short-term bias, long-term bias, and comprehensive bias directly reflect the degree to which the recommendation model is affected by exposure bias in the corresponding feature dimensions.
[0052] The final exposure correction term is formed by weighted summation of the aforementioned short-term, long-term, and combined biases. Bias signals from different feature perspectives can be integrated into a unified correction signal and used as part of the recommendation loss function. During training, minimizing the exposure correction term simultaneously drives the recommendation model to approximate multiple unbiased expectations revealed by long-term, short-term, and combined features, thereby systematically reducing exposure selection bias in predictions and enabling the recommendation model to learn patterns closer to users' true interests.
[0053] The aforementioned technical solution, by independently calculating counterfactual expectations based on short-term, long-term, and comprehensive features, can identify and quantify exposure biases with different time-scale characteristics that are mixed across different levels, such as real-time behavior, stable interests, and overall decision-making. Simultaneously minimizing these biases during model training is equivalent to driving the recommendation model's predictions to simultaneously approximate multiple unbiased targets targeting different sources of bias. This not only reduces the overall prediction bias of the recommendation model but, more importantly, by providing correction signals from different feature perspectives, guides the recommendation model to learn an internal representation capable of simultaneously resisting multiple biases, thereby significantly enhancing the recommendation model's generalization ability and prediction accuracy in truly unbiased environments.
[0054] In an optional embodiment, the at least two business loss terms in the recommendation loss function are calculated based on the predicted values and sample labels corresponding to each business objective, including: calculating the basic loss value corresponding to each business objective based on the predicted values corresponding to each business objective and the feedback labels corresponding to the at least two business objectives in the sample labels; determining the target weight combination by a weight controller based on historical weight combinations, the training state of the recommendation model, the evaluation index of the recommendation model on the validation set, and the reward weight provided by the revenue calibrator in the most recent working cycle; and determining the business loss terms corresponding to each business objective based on the basic loss value and the target weight combination.
[0055] The base loss value is the error between the predicted value and the corresponding feedback label for each business objective, calculated using a specific loss function. The base loss value measures the basic prediction accuracy of the recommendation model on that single objective. The specific type of the specific loss function is not limited here; it is determined based on actual business needs. For example, the specific loss function could be cross-entropy loss.
[0056] The weight controller is an independent decision-making module introduced to achieve automatic balancing among multiple business objectives. Its function is to determine a target weight combination based on multi-source information, that is, to assign a dynamic weight to each business loss item.
[0057] The information used by the weight controller to make decisions includes: historical weight combinations, the training state of the recommendation model, the evaluation metric of the recommendation model on the validation set, and the reward weights provided by the reward calibrator in the most recent working cycle. Among these, historical weight combinations are weight sequences used in previous training to provide empirical reference; the training state of the recommendation model is used to determine the stage of the model, and may include the current training round and loss convergence status; the evaluation metric of the recommendation model on the validation set reflects the current overall capability of the recommendation model.
[0058] The revenue calibrator is another collaborative module that periodically calculates and updates reward weights based on the actual business revenue generated by online advertising. The reward weights reflect the predictive importance of offline metrics for each business objective to the final online revenue. By considering this information when determining the target weight combination, the weight controller ensures that the weight allocation adapts to the current training dynamics and capabilities of the recommendation model while aligning with the ultimate goal of maximizing online business revenue. One work cycle of the revenue calibrator refers to the time period during which it performs a complete data collection, calculation, and parameter update, encompassing multiple training iterations of the recommendation model, with each iteration representing a cycle of model parameter updates. In other words, the cycle in which the revenue calibrator updates reward weights based on online business revenue is much longer than the iteration frequency of model training. This ensures that the reward weights have sufficient stability, providing consistent optimization guidance for multiple consecutive training iterations; simultaneously, the collection and statistical analysis of online revenue data also requires a complete cycle, guaranteeing the reliability of the feedback signals.
[0059] The business loss term is obtained by multiplying the base loss value of each objective by the corresponding weight in the objective weight combination and then summing the results. This approach assigns the dynamic importance determined by the weight controller, which is consistent with the online revenue orientation, to each base loss value. This guides the recommendation model to perform focused joint optimization of multiple business objectives during training, in a way that matches the business results, avoiding the problems of rigid optimization direction or disconnect from business value caused by fixed loss weights.
[0060] The aforementioned technical solution uses a weight controller that integrates historical weights, real-time training status, validation set performance, and reward weights calibrated by online returns to determine the target weight combination. This allows the importance of each business loss item to adapt to different training stages and capability levels of the recommendation model, while maintaining a strong correlation with the overall online return objective. This not only avoids the optimization rigidity caused by fixed weights but also automatically reconciles conflicts between multiple business objectives, guiding the recommendation model to learn efficiently towards maximizing overall returns, thereby significantly improving the overall performance and business value of the recommendation system.
[0061] In an optional embodiment, determining the target weight combination by the weight controller based on historical weight combinations, the training state of the recommendation model, the evaluation metric of the recommendation model on the validation set, and the reward weights provided by the revenue calibrator in the most recent work cycle includes: determining candidate weight combinations by the weight controller based on the training state of the recommendation model; estimating the estimated reward corresponding to the candidate weight combinations based on the evaluation metric of the recommendation model on the validation set and the reward weights provided by the revenue calibrator in the most recent work cycle; wherein one work cycle of the revenue calibrator corresponds to multiple training iterations of the recommendation model; determining the exploration coefficient of the weight controller based on the training state and the historical weight combinations corresponding to the business loss item; selecting the target weight combination from the candidate weight combinations based on the exploration coefficient and the estimated reward, and generating a decision sequence; wherein the decision sequence includes at least: the selected target weight combination, the state vector on which the weight combination selection is based, the estimated reward corresponding to the target weight combination, and the selection probability; the state vector is an encoded representation of the training state of the recommendation model, the evaluation metric of the recommendation model on the validation set, and the historical weight combinations.
[0062] Here, candidate weight combinations refer to a set of potential weight allocation schemes for the weight controller to evaluate and select in the current step, and their generation is strongly correlated with the training state of the recommendation model. This is because the recommendation model's sensitivity and requirements for changes in loss weights differ at different training stages. For example, in the early stages of training, when model parameters are random, a more balanced set of weights may be needed to stabilize learning; during the period of rapid loss decline, it may be necessary to focus on the currently under-optimized objective; and when approaching convergence, fine-tuning the weights may be necessary for precise balance.
[0063] After the candidate weight combinations are determined, the weight controller calculates the estimated reward for each combination. This is estimated by combining the evaluation metric of the recommendation model on the validation set with the reward weights provided by the revenue calibrator. The estimated reward represents the overall revenue signal expected after adopting that weight combination. The evaluation metric of the recommendation model on the validation set refers to the performance measure calculated by the model for each business objective on an independent validation dataset in its current training state. The evaluation metric objectively reflects the level and state of the model's learned capabilities and is the basis for predicting its future optimization potential.
[0064] Simultaneously, the weight controller determines the exploration coefficient based on the training state and historical weight combinations. The training state sets a baseline level for the exploration coefficient by analyzing the training stage of the recommendation model, determining the macro-level exploration intensity. Historical weight combinations, by analyzing the concentration and diversity of weight selection, determine whether the search has fallen into local optima, thereby fine-tuning the exploration coefficient. The exploration coefficient is used to balance exploratory behavior and exploitation behavior in decision-making, where exploratory behavior refers to trying new combinations, and exploitation behavior refers to selecting high-reward combinations.
[0065] The exploration coefficient and the estimated reward jointly determine the final selection of the target weight combination. The estimated reward provides a value orientation for each candidate weight combination, directly reflecting the expected contribution of the combination to online revenue, and is the core basis for selection. The exploration coefficient provides strategy adjustment, injecting uncertainty into the decision-making process: when its value is high, it tends to try candidate weight combinations whose estimated rewards are not the highest but may lead to new discoveries, in order to avoid getting trapped in local optima; when its value is low, it strictly relies on the estimated reward for utilization. The combined effect of the two allows the selection process to focus on high-return directions while maintaining necessary exploration capabilities, thereby achieving a balance between maximizing returns and search efficiency.
[0066] The weight controller integrates exploration coefficients and estimated rewards to select the target weight combination from candidate weight combinations. The entire decision-making process is recorded as a decision sequence, where key elements include: the selected target weight combination, the state vector used for selection, the estimated reward of the combination, and its probability of selection. The state vector is an encoding representation of the training state of the recommendation model, the evaluation metric of the recommendation model on the validation set, and historical weight combinations. This approach models weight selection as a state-based sequential decision problem, using online reward signals for calibration, enabling the allocation of multi-objective weights to adapt to the model's learning dynamics and continuously align with the final business objective.
[0067] The above technical solution involves a weight controller that generates candidate weights and dynamically adjusts exploration coefficients based on the training state. This ensures that the weight selection strategy accurately matches the needs of different training stages, improving optimization efficiency and stability. By introducing reward weights periodically provided by the revenue calibrator to predict the reward of candidate weight combinations, the direction of weight optimization is always strongly correlated with the online comprehensive revenue objective, avoiding a disconnect between offline metrics and online performance. Furthermore, the revenue calibrator operates over a longer period covering multiple training iterations, guaranteeing the reliability of the reward weights as guiding signals and effectively reducing optimization noise caused by online data fluctuations. Finally, the complete record of the decision sequence provides a traceable and analyzable basis for the entire dynamic weight adjustment process, enhancing the system's interpretability and iterative optimization capabilities. This technical solution enables the recommendation model to automatically and intelligently balance multi-objective losses, thereby driving continuous and robust performance improvements.
[0068] In an optional embodiment, determining candidate weight combinations based on the training state of the recommendation model using the weight controller includes: determining whether weight exploration conditions are met based on the current optimization stage and current convergence status in the training state; if the weight exploration conditions are met, determining a basic weight combination based on the estimated reward corresponding to the historical weight combination; interpolating the basic weight combination or adding a perturbation factor to the basic weight combination to obtain a dynamic weight combination; and determining the candidate weight combination based on the dynamic weight combination and the preset weight combination.
[0069] The training state refers to a set of dynamic indicators reflecting the model's training progress, with the current optimization stage and current convergence status being key judgment criteria. Optionally, the current optimization stage includes: early training, mid-training, and late training. The current convergence status can be determined by judging whether the loss is stable. The weight controller uses these two indicators to determine whether the weight exploration condition is met, which is a policy rule that decides whether to try a new weight combination.
[0070] Optionally, in the early stages of training or when the model has not converged, more exploration is needed to find effective directions, which is considered to satisfy the weight exploration condition. In the later stages of training or when it is close to convergence, the focus should be on using known effective weights to stabilize the optimization, which is considered to not satisfy the weight exploration condition.
[0071] When the conditions for weight exploration are met, the weight controller first determines a base weight combination based on historical weight combinations and their corresponding estimated rewards. Historical weight combinations refer to weight vectors that have been selected in the past, and the estimated reward is the expected value assessment of that combination. Optionally, the base weight combination is selected from historical weight combinations with higher estimated rewards, allowing for exploration near high-value areas based on experience, thus improving efficiency.
[0072] Next, the basic weight combination is interpolated or a perturbation factor is added to obtain the dynamic weight combination. Interpolation involves linear interpolation between two historical weight combinations with higher estimated rewards. Adding a perturbation factor can be achieved by introducing random noise. By subjecting the basic weight combination, as a high-quality solution, to controllable random or interpolation perturbations in its neighborhood, the dynamic weight combination can be guaranteed to have some potential while introducing necessary diversity and avoiding getting trapped in local optima.
[0073] Finally, the generated dynamic weight combination is merged with a set of preset weight combinations to form the final candidate weight combination. The preset weight combination refers to a set of fixed weight vectors with clear business meanings, predefined based on domain experience or prior knowledge before model training begins. Preset weight combinations typically represent typical policy tendencies in multi-objective optimization. The candidate weight combination includes both novel choices generated based on current learning dynamics and historical experience, as well as predefined, representative basic policies, balancing the intelligence of exploration with the completeness of policies.
[0074] The above technical solution dynamically triggers exploration based on the training state, matching the exploration behavior with the actual optimization needs of the recommendation model and ensuring training stability. During exploration, the basic weight combination is determined based on the estimated reward of historical weight combinations, and dynamic weight combinations are generated by interpolation or adding perturbation factors, making the exploration both experience-based and innovative, thus improving the efficiency of discovering high-quality combinations. Finally, the candidate weight combination is formed by combining the dynamic weight combination and the preset weight combination, ensuring the completeness of the strategy and the quality of the selection basis, thereby making the optimization process of multi-objective weights adaptive, efficient, and robust.
[0075] In an optional embodiment, the method further includes: acquiring business revenue data and a decision sequence recorded by a weight controller via a revenue calibrator; wherein the business revenue data is actually generated by the target advertisement during the recommendation service and is associated with the at least two business objectives; calculating the revenue metric of the online business based on the business revenue data, and using the revenue metric as an environmental reward signal; calculating the update gradient of the reward weight using a policy gradient method based on the environmental reward signal and the decision sequence; adjusting the reward weight based on the update gradient, and providing the adjusted reward weight to the weight controller.
[0076] The revenue calibrator is a module responsible for calibrating the model's optimization direction based on real online feedback. Its work begins with acquiring business revenue data, which refers to the actual commercial results generated by the target advertisement during its online recommendation service, such as total transaction volume and advertising revenue, and these data must be logically related to at least two business objectives. The purpose of acquiring this data is to directly link the model's optimization performance to ultimate business value.
[0077] The revenue calibrator calculates a comprehensive online business revenue metric, such as normalized daily average profit, based on business revenue data, and uses this metric as the environmental reward signal. The environmental reward signal acts as the global reward role in the reinforcement learning framework, representing the final evaluation of the overall policy performance over a given time period. Simultaneously, the revenue calibrator acquires the decision sequence recorded by the weight controller, which fully documents the weight controller's decision history in the previous period. Based on the environmental reward signal and the decision sequence, the revenue calibrator calculates the update gradient of the reward weights using the policy gradient method. This treats the weight controller's decision-making process as a parameterized policy, with the parameters being the reward weights, and the environmental reward signal representing the reward obtained after the policy is executed. The policy gradient method is a class of optimization algorithms that analyzes the relationship between the policy's decision sequence and the final reward to estimate how to adjust the policy parameters to obtain a higher expected return. The update gradient calculated using this method indicates in which direction the reward weights should be adjusted so that the weight controller is more likely to make decisions that yield higher environmental rewards in the future. Finally, the revenue calibrator adjusts the reward weights based on the update gradient, completing a learning update, and provides the adjusted reward weights to the weight controller.
[0078] The aforementioned technical solution directly uses real business revenue data generated by target advertising to calculate environmental reward signals, ensuring a strong correlation between the ultimate direction of model optimization and business results. This fundamentally avoids the disconnect between offline metric improvements and lackluster online returns. By combining the historical decision sequences of the weight controller with environmental reward signals through the policy gradient method, the contribution of different decisions to the final revenue can be effectively evaluated, and the update gradient of the reward weights can be calculated accordingly. This process ensures that the adjustment of reward weights is not based on heuristic rules but is automatically completed through data-driven learning. Finally, the updated reward weights are fed back to the weight controller, forming a complete closed loop of "decision-execution-feedback-calibration," continuously improving the overall returns of multi-objective optimization.
[0079] Example 3 Figure 3 This is a schematic diagram of the advertising recommendation device provided in Embodiment 3 of this application. This embodiment is applicable to the case of advertising recommendation. The device is implemented by software and / or hardware and can be integrated into electronic devices such as smart terminals.
[0080] like Figure 3 As shown, the advertising recommendation device 300 may include: The candidate advertisement determination module 310 is used to determine candidate advertisements to be recommended based on the advertisement recommendation request for the target user; The feature segmentation module 320 is used to segment the user features corresponding to the target user and the advertising features corresponding to the candidate advertisement into long-term features and short-term features based on the time attribute of the business objective. The feature splicing module 330 is used to splice the long-term features and the short-term features using splicing weights to obtain comprehensive features, and input the comprehensive features into a pre-trained recommendation model; The target advertisement determination module 340 is used to output the predicted value of the candidate advertisement on each business objective based on the comprehensive features through the recommendation model, and determine the target advertisement for the target user from the candidate advertisements based on the predicted value; The network parameters of the recommendation model are obtained by optimizing the recommendation loss function. The recommendation loss function includes an exposure correction term and at least two business loss terms. The exposure correction term is used to correct exposure selection bias and optimize the splicing weights. The weight combination corresponding to the business loss terms is determined based on the correlation between the evaluation indicators of each business objective on offline verification data and the revenue indicators of online business.
[0081] This technical solution, by dividing user and advertising features into long-term and short-term features based on time attributes and employing learnable concatenation weights for adaptive fusion, can more precisely characterize users' stable interests and real-time intentions, enhancing the discriminative power of feature representation. More importantly, the designed recommendation loss function creatively combines an exposure correction term with a multi-business loss term: the exposure correction term, based on counterfactual inference principles, not only effectively eliminates selection bias in historical exposure data, enabling the recommendation model to learn users' true interests and preferences, but also generates a gradient that serves as the dominant signal driving the optimization of the feature concatenation weights, enabling the recommendation model to learn an optimal feature fusion method that proactively resists exposure bias. Simultaneously, the multi-business loss term is determined and adjusted based on the statistical correlation between offline verification indicators and real online business revenue, ensuring that the multi-objective optimization direction of the recommendation model continuously aligns with the distribution of real online user feedback. In summary, this technical solution, through end-to-end joint optimization, achieves synergistic enhancement of effective feature fusion, data bias correction, and multi-objective dynamic balance, ultimately significantly improving the matching degree between recommendation results and user needs, thereby improving the user experience.
[0082] Optionally, the apparatus further includes: a model training module for training a recommendation model; the model training module includes: a feature segmentation submodule for segmenting user features of sample users and advertising features of sample advertisements in the training samples into long-term features and short-term features based on the time attributes of business objectives; a feature concatenation submodule for weighted concatenation of the long-term features and short-term features using the concatenation weights to obtain comprehensive features of the training samples; a feature input submodule for inputting the comprehensive features into the recommendation model to be trained, so as to obtain the predicted values output by the recommendation model for each business objective through forward propagation; a loss function calculation submodule for calculating the exposure correction term and the at least two business loss terms in the recommendation loss function based on the predicted values and sample labels corresponding to each business objective, the short-term features, the long-term features, and the comprehensive features; a gradient calculation submodule for calculating the first gradient generated by the exposure correction term and the second gradient generated by the at least two business loss terms through backpropagation based on the recommendation loss function; and a weight update submodule for updating the network parameters of the backbone network in the recommendation model using the first gradient and the second gradient, and updating the concatenation weights using the first gradient.
[0083] Optionally, the loss function calculation submodule includes: a counterfactual expectation determination unit, used to determine the counterfactual expectations corresponding to the short-term feature, the long-term feature, and the comprehensive feature for each business objective, based on the exposure label in the sample label and the feedback label corresponding to the business objective, to obtain the short-term counterfactual expectation, the long-term counterfactual expectation, and the comprehensive counterfactual expectation; a bias determination unit, used to calculate the bias between the predicted value output by the recommendation model for each business objective and the short-term counterfactual expectation, the long-term counterfactual expectation, and the comprehensive counterfactual expectation, to obtain the short-term bias, the long-term bias, and the comprehensive bias; and an exposure correction term determination unit, used to weight the short-term bias, the long-term bias, and the comprehensive bias, and determine the exposure correction term based on the weighted result.
[0084] Optionally, the loss function calculation submodule includes: a basic loss value calculation unit, used to calculate the basic loss value corresponding to each business objective based on the predicted value corresponding to each business objective and the feedback labels corresponding to at least two business objectives in the sample labels; a weight combination determination unit, used to determine the target weight combination through the weight controller based on historical weight combinations, the training state of the recommendation model, the evaluation index of the recommendation model on the validation set, and the reward weight provided by the revenue calibrator in the most recent working period; and a business loss item determination unit, used to determine the business loss item corresponding to each business objective based on the basic loss value and the target weight combination.
[0085] Optionally, the weight combination determination unit includes: a candidate combination determination subunit, used to determine candidate weight combinations based on the training state of the recommendation model by the weight controller; a predicted reward determination subunit, used to predict the predicted reward corresponding to the candidate weight combination based on the evaluation index of the recommendation model on the validation set and the reward weight provided by the revenue calibrator in the most recent working cycle; wherein, one working cycle of the revenue calibrator corresponds to multiple training iterations of the recommendation model; an exploration coefficient determination subunit, used to determine the exploration coefficient of the weight controller based on the historical weight combinations corresponding to the training state and the business loss item; and a weight combination determination subunit, used to select a target weight combination from the candidate weight combinations based on the exploration coefficient and the predicted reward, and generate a decision sequence; wherein, the decision sequence includes at least: the selected target weight combination, the state vector on which the weight combination selection is based, the predicted reward corresponding to the target weight combination, and the selection probability; the state vector is an encoded representation of the training state of the recommendation model, the evaluation index of the recommendation model on the validation set, and the historical weight combinations.
[0086] Optionally, the candidate combination determination sub-unit is specifically used for: determining whether the weight exploration condition is met based on the current optimization stage and current convergence status in the training state by the weight controller; if the weight exploration condition is met, determining the basic weight combination based on the estimated reward corresponding to the historical weight combination; performing interpolation processing on the basic weight combination or adding a perturbation factor to the basic weight combination to obtain a dynamic weight combination; and determining the candidate weight combination based on the dynamic weight combination and the preset weight combination.
[0087] Optionally, the apparatus further includes: a data acquisition module, configured to acquire business revenue data and a decision sequence recorded by a weight controller via a revenue calibrator; wherein the business revenue data is actually generated by the target advertisement during the recommendation service and is associated with the at least two business objectives; a reward signal determination module, configured to calculate the revenue index of the online business based on the business revenue data and use the revenue index as an environmental reward signal; an update gradient determination module, configured to calculate the update gradient of the reward weight using a policy gradient method based on the environmental reward signal and the decision sequence; and a reward weight adjustment unit, configured to adjust the reward weight based on the update gradient and provide the adjusted reward weight to the weight controller.
[0088] The advertising recommendation device provided in the embodiments of the invention can execute the advertising recommendation method provided in any embodiment of this application, and has the corresponding performance modules and beneficial effects for executing the advertising recommendation method.
[0089] In the technical solution of this application, the user data involved, such as the user characteristics corresponding to the target user, is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0090] Example 4 According to embodiments of this application, this application also provides an electronic device, a readable storage medium, and a computer program product.
[0091] Figure 4 A schematic diagram of an electronic device 410, which can be implemented using an embodiment, is shown. The electronic device 410 includes at least one processor 411 and a memory, such as a read-only memory 412 or a random access memory 413, communicatively connected to the at least one processor 411. The memory stores computer programs executable by the at least one processor. The processor 411 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 412 or loaded from storage unit 418 into the random access memory 413. The random access memory 413 may also store various programs and data required for the operation of the electronic device 410. The processor 411, read-only memory 412, and random access memory 413 are interconnected via a bus 414. An input / output interface 415 is also connected to the bus 414.
[0092] Multiple components in electronic device 410 are connected to input / output interface 415, including: input unit 416, such as keyboard, mouse, etc.; output unit 417, such as various types of monitors, speakers, etc.; storage unit 418, such as disk, optical disk, etc.; and communication unit 419, such as network card, modem, wireless transceiver, etc. Communication unit 419 allows electronic device 410 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0093] Processor 411 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 411 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 411 performs the various methods and processes described above, such as advertising recommendation methods.
[0094] In some embodiments, the advertising recommendation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 418. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 410 via read-only memory 412 and / or communication unit 419. When the computer program is loaded into random access memory 413 and executed by processor 411, one or more steps of the advertising recommendation method described above may be performed. Alternatively, in other embodiments, processor 411 may be configured to perform the advertising recommendation method by any other suitable means (e.g., by means of firmware).
[0095] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), complex programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0096] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable advertising recommendation device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0097] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory or flash memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0098] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0099] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as an advertising recommendation server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0100] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to address the shortcomings of traditional physical hosts and virtual private servers, such as high management difficulty and weak business scalability.
[0101] This application also discloses a computer program product, which includes a computer program that, when executed by a processor, implements the advertising recommendation method provided in any embodiment of this application. This program product and the advertising recommendation methods disclosed in the embodiments of this application belong to the same inventive concept, and therefore will not be described in detail here.
[0102] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0103] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. An advertising recommendation method, characterized in that, The method includes: Based on the ad recommendation request for the target user, determine the candidate ads to be recommended; Based on the time attribute of the business objective, the user characteristics corresponding to the target user and the advertising characteristics corresponding to the candidate advertisement are divided into long-term characteristics and short-term characteristics. The long-term features and the short-term features are concatenated using concatenation weights to obtain comprehensive features, and the comprehensive features are then input into a pre-trained recommendation model. The recommendation model outputs predicted values of the candidate advertisements for each business objective based on the comprehensive features, and determines the target advertisement for the target user from the candidate advertisements based on the predicted values. The network parameters of the recommendation model are obtained by optimizing the recommendation loss function. The recommendation loss function includes an exposure correction term and at least two business loss terms. The exposure correction term is used to correct exposure selection bias and optimize the splicing weights. The weight combination corresponding to the business loss terms is determined based on the correlation between the evaluation indicators of each business objective on offline verification data and the revenue indicators of online business.
2. The method according to claim 1, characterized in that, The recommendation model was trained in the following manner: Based on the time attribute of business objectives, the user characteristics of sample users and the advertising characteristics of sample advertisements in the training samples are divided into long-term features and short-term features. The long-term and short-term features are weighted and concatenated using the concatenation weights to obtain the comprehensive features of the training samples; The comprehensive features are input into the recommendation model to be trained, so that the predicted values of the recommendation model for each business objective are obtained through forward propagation. Based on the predicted values and sample labels corresponding to each business objective, the short-term features, long-term features, and comprehensive features, the exposure correction term and the at least two business loss terms in the recommendation loss function are calculated respectively. Based on the recommendation loss function, the first gradient generated by the exposure correction term and the second gradient generated by the at least two business loss terms are calculated through backpropagation. The network parameters of the backbone network in the recommendation model are updated using the first gradient and the second gradient, and the splicing weights are updated using the first gradient.
3. The method according to claim 2, characterized in that, Based on the predicted values and sample labels corresponding to each business objective, the short-term features, long-term features, and comprehensive features, the exposure correction term in the recommendation loss function is calculated, including: For each business objective, based on the exposure label in the sample label and the feedback label corresponding to the business objective, the counterfactual expectations corresponding to the short-term feature, the long-term feature and the comprehensive feature are determined respectively, so as to obtain the short-term counterfactual expectation, the long-term counterfactual expectation and the comprehensive counterfactual expectation; For each business objective, the deviations between the predicted values output by the recommendation model for each business objective and the short-term counterfactual expectations, the long-term counterfactual expectations, and the comprehensive counterfactual expectations are calculated to obtain the short-term deviation, the long-term deviation, and the comprehensive deviation. The short-term deviation, the long-term deviation, and the combined deviation are weighted, and the exposure correction term is determined based on the weighted result.
4. The method according to claim 2, characterized in that, Based on the predicted values and sample labels corresponding to each business objective, calculate the at least two business loss terms in the recommendation loss function, including: Based on the predicted values corresponding to each business objective and the feedback labels in the sample labels that correspond to at least two business objectives, calculate the basic loss value corresponding to each business objective. The target weight combination is determined by the weight controller based on historical weight combinations, the training state of the recommendation model, the evaluation metric of the recommendation model on the validation set, and the reward weight provided by the reward calibrator in the most recent working cycle. Based on the combination of the base loss value and the target weight, the business loss item corresponding to each business objective is determined.
5. The method according to claim 4, characterized in that, The process of determining the target weight combination through a weight controller based on historical weight combinations, the training state of the recommendation model, the evaluation metrics of the recommendation model on the validation set, and the reward weights provided by the reward calibrator in the most recent working cycle includes: The weight controller determines candidate weight combinations based on the training state of the recommendation model. Based on the evaluation metrics of the recommendation model on the validation set and the reward weights provided by the reward calibrator in the most recent working cycle, the predicted reward corresponding to the candidate weight combination is estimated; wherein, one working cycle of the reward calibrator corresponds to multiple training iterations of the recommendation model. Based on the historical weight combinations corresponding to the training state and the business loss item, the exploration coefficient of the weight controller is determined; Based on the exploration coefficients and the estimated reward, a target weight combination is selected from the candidate weight combinations, and a decision sequence is generated; wherein, the decision sequence includes at least: the selected target weight combination, the state vector on which the weight combination selection is based, the estimated reward and selection probability corresponding to the target weight combination; the state vector is an encoded representation of the training state of the recommendation model, the evaluation index of the recommendation model on the validation set, and the historical weight combinations.
6. The method according to claim 5, characterized in that, The step of determining candidate weight combinations based on the training state of the recommendation model using the weight controller includes: The weight controller determines whether the weight exploration conditions are met based on the current optimization stage and current convergence status in the training state. If the weight exploration conditions are met, the basic weight combination is determined based on the estimated reward corresponding to the historical weight combination. The basic weight combination is interpolated or a perturbation factor is added to the basic weight combination to obtain a dynamic weight combination; The candidate weight combination is determined based on the dynamic weight combination and the preset weight combination.
7. The method according to claim 5, characterized in that, The method further includes: Business revenue data and decision sequences recorded by a weight controller are obtained through a revenue calibrator; wherein, the business revenue data is actually generated by the target advertisement during the recommendation service and is associated with the at least two business objectives; The revenue indicators of the online business are calculated based on the business revenue data, and the revenue indicators are used as environmental reward signals. Based on the environmental reward signal and the decision sequence, the update gradient of the reward weight is calculated using the policy gradient method; The reward weights are adjusted based on the update gradient, and the adjusted reward weights are provided to the weight controller.
8. An advertising recommendation device, characterized in that, The device includes: The candidate ad determination module is used to determine candidate ads to be recommended based on ad recommendation requests for target users; The feature segmentation module is used to divide the user features corresponding to the target user and the advertising features corresponding to the candidate advertisement into long-term features and short-term features based on the time attribute of the business objective. The feature concatenation module is used to concatenate the long-term features and the short-term features using concatenation weights to obtain comprehensive features, and input the comprehensive features into a pre-trained recommendation model; The target advertisement determination module is used to output the predicted value of the candidate advertisement on each business objective based on the comprehensive features through the recommendation model, and determine the target advertisement for the target user from the candidate advertisements based on the predicted value; The network parameters of the recommendation model are obtained by optimizing the recommendation loss function. The recommendation loss function includes an exposure correction term and at least two business loss terms. The exposure correction term is used to correct exposure selection bias and optimize the splicing weights. The weight combination corresponding to the business loss terms is determined based on the correlation between the evaluation indicators of each business objective on offline verification data and the revenue indicators of online business.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the advertising recommendation method as described in any one of claims 1-7.
10. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the advertising recommendation method as described in any one of claims 1-7.