Method and device for training synergy model
By introducing smooth indication function and performance evaluation index optimization training process in the integrated model, the deviation between the model training objectives and actual applications is solved, and more accurate intervention incremental prediction and user sorting is achieved, which improves the prediction effect of the model.
Patent Information
- Application Number
- CN202510399119.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
There is a deviation between the training objectives used by the existing integrated model and the actual application, which leads to the poor performance of the model in actual use, making it impossible to effectively identify users who are more sensitive to intervention and conduct differentiated interventions.
By introducing a smooth indicator function, the ranking relative value between samples is determined, and based on performance evaluation indicators such as improving the area under the curve and Top-K increment, the parameter training process of the efficiency model can be optimized so that it can directly predict the intervention increment and update the model parameters using the gradient descent method.
The performance of the efficiency enhancement model is improved, and users' intervention increments can be more accurately identified and sorted, thereby achieving better intervention effects and resource optimization.
Smart Images

Figure CN120278219A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of machine learning, and in particular, to methods and apparatuses for training uplift models. Background Art
[0002] An uplift model is a model used to measure the impact of an intervention on an individual's response, and it has extensive applications in scenarios or fields such as e-commerce platforms, content platforms, medical treatments, and education. For example, in an e-commerce platform, merchants often apply marketing-related interventions (such as pushing messages, sending coupons, etc.) to users to increase the probability of response behaviors such as clicks / purchases. To this end, merchants hope to identify the group of users with the highest marketing effect and target these users with interventions to increase the overall revenue.
[0003] There is a deviation between the training objective used in training the uplift model in the related art and the actual application. Specifically, although performance evaluation metrics related to the actual application are used in the performance evaluation stage of these models, these performance evaluation metrics are not directly considered during model modeling and training. This difference makes the performance of these models in actual use unsatisfactory. Therefore, a model training method is needed to improve the performance of the uplift model. Summary of the Invention
[0004] One or more embodiments of this specification describe methods and apparatuses for training uplift models to improve the performance of uplift models.
[0005] In a first aspect, a method for training an uplift model is provided, including:
[0006] Obtain a sample set, where any sample includes user-related features, a user group indicating whether the user is intervened, and a user response;
[0007] Input the user-related features of each sample into the uplift model to obtain the intervention increment of each sample, which indicates the improvement of the intervention on the user response;
[0008] For any target sample, determine the ranking factor of the target sample according to the relative ranking value of the target sample and at least one other sample. Among them, the relative ranking value of the target sample and any other sample is obtained by inputting the difference between the target intervention increment of the target sample and the other intervention increment of the other sample into a predetermined indication function, and the indication function is a monotonically increasing smooth function;
[0009] Update the parameters of the uplift model in the direction of increasing the function value of the objective function, where the function value of the objective function is determined according to the user group, user response, and the ranking factor of each sample.
[0010] In some possible embodiments, the user-related features include user attribute features and scenario features where the user is located.
[0011] In some possible embodiments, the relative ranking value of the target sample with respect to any other sample is determined through the following steps:
[0012] Use the difference multiplied by a preset first coefficient as the independent variable and input it into a preset smoothing function to obtain the relative ranking value.
[0013] In some possible embodiments, the smoothing function includes at least one of the following: Sigmoid function, tanh function.
[0014] In some possible embodiments, the function value of the indicator function is between 0 and 1.
[0015] In some possible embodiments, the at least one other sample is each sample other than the target sample; determining the ranking factor of the target sample includes:
[0016] Add 1 to the sum of the relative ranking values of the target sample with respect to each sample other than the target sample to obtain the ranking factor of the target sample.
[0017] In some possible embodiments, the objective function is the area under the lift curve function, and its function value is determined through the following steps:
[0018] Determine the lift share of the target sample. Among them, if the target sample belongs to the intervention group, its lift share is the user response of the target sample divided by the total number of samples in the intervention group; if the target sample belongs to the control group, its lift share is the user response of the target sample divided by the total number of samples in the control group and take the opposite value;
[0019] Use the ranking factor as the weight to perform a weighted sum of the lift shares of each sample to determine the function value of the objective function.
[0020] In some possible embodiments, the at least one other sample is the Kth sample among all samples sorted in descending order of intervention increment; determining the ranking factor of the target sample includes:
[0021] Input the difference between the target intervention increment and the intervention increment of the Kth sample into the indicator function to determine the ranking factor of the target sample.
[0022] In some possible embodiments, the objective function is the top-K intervention effect function, and its function value is determined through the following steps:
[0023] Determine the first average intervention increment of the intervention group according to the ranking factors and user responses of each intervention group sample in the intervention group;
[0024] Determine the second average intervention increment of the control group according to the ranking factors and user responses of each control group sample in the control group;
[0025] Determine the function value of the objective function according to the difference between the first average intervention increment and the second average intervention increment.
[0026] In some possible implementation manners, determining the first average intervention increment of the intervention group includes:
[0027] Taking the ranking factor as the weight, perform weighted summation on the user responses of each intervention group sample to obtain the first weighted sum;
[0028] Sum up the ranking factors of each intervention group sample to obtain the first smoothing quantity;
[0029] Determine the first average intervention increment according to the ratio of the first weighted sum to the first smoothing quantity.
[0030] In some possible implementation manners, determining the second average intervention increment of the control group includes:
[0031] Taking the ranking factor as the weight, perform weighted summation on the user responses of each control group sample to obtain the second weighted sum;
[0032] Sum up the ranking factors of each control group sample to obtain the second smoothing quantity;
[0033] Determine the second average intervention increment according to the ratio of the second weighted sum to the second smoothing quantity.
[0034] In some possible implementation manners, in the direction of increasing the function value of the objective function, update the parameters of the synergy model, including:
[0035] Take the opposite number of the objective function as the loss function, and update the parameters of the synergy model based on the gradient descent method.
[0036] In a second aspect, there is provided an apparatus for training a synergy model, including:
[0037] An acquisition unit, configured to acquire a sample set, where any sample includes user-related features, a user group indicating whether the user is intervened, and a user response;
[0038] An intervention increment prediction unit, configured to input the user-related features of each sample into the synergy model to obtain the intervention increment of each sample, which indicates the improvement of the intervention on the user response;
[0039] a ranking factor determination unit configured to determine, for any target sample, a ranking factor of the target sample according to a relative ranking value between the target sample and at least one other sample, wherein the relative ranking value between the target sample and any other sample is obtained by inputting a difference between a target intervention increment of the target sample and other intervention increments of the other sample into a predetermined indicator function, wherein the indicator function is a monotonically increasing smooth function;
[0040] The model training unit is configured to update the parameters of the synergy model in the direction of increasing the function value of the objective function, wherein the function value of the objective function is determined according to the user group, user response and ranking factor of each sample.
[0041] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.
[0042] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.
[0043] The method and device for training a efficiency-enhancing model proposed in the embodiments of this specification, the efficiency-enhancing model trained by the method receives user-related features and predicts the corresponding intervention increment, and the intervention increment indicates the improvement of the intervention on the user response. The method uses a smooth indicator function to determine the relative ranking value between each sample according to the intervention increment of each sample, and determines the ranking factor of each sample based on the relative ranking value. Then, the objective function is determined according to the user group, user response and ranking factor of each sample, and the objective function can be determined based on the performance evaluation index of the efficiency-enhancing model. Since the indicator function is smooth and differentiable, the objective function determined based on the indicator function is also differentiable, and the objective function can be used to directly update the parameters of the efficiency-enhancing model.
[0044] The method proposed in the embodiment of this specification can directly train a synergy model that can predict intervention increments based on performance evaluation indicators by introducing a smooth indicator function, thereby effectively improving the performance of the synergy model. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0046] Figure 1Schematic diagram showing the relationship between intervention and user response according to an example;
[0047] Figure 2 Schematic diagram showing the lift curve according to an example;
[0048] Figure 3 Flowchart showing a method for training a lift model according to an embodiment;
[0049] Figure 4 Schematic block diagram showing an apparatus for training a lift model according to an embodiment. Detailed implementation manners
[0050] The solutions provided in this specification will be described below with reference to the accompanying drawings.
[0051] A lift model is a type of model used to measure the difference in the impact of intervention measures on individuals. Different from traditional classification models or regression models, the goal of a lift model is not to predict the absolute probability of user behavior, but to predict the change in user response (i.e., intervention increment) caused by a certain intervention (such as marketing activities, promotional behaviors, treatment behaviors, etc.).
[0052] For example, in the field of commercial marketing, using a lift model can identify users who are sensitive to marketing means and conduct targeted marketing on these users to control costs and improve marketing effects. In the medical field, for a specific treatment plan, using a lift model can predict which patients will have a positive treatment effect with this treatment plan, and which patients will have no effect or even a reverse effect. In the field of education, using a lift model can predict the improvement in learning outcomes brought by a certain education plan to different student groups.
[0053] Exemplarily, in a specific example, for a commodity, the relationship between intervention (such as advertising placement, pushing to the home page directionally, issuing coupons, etc.) and user response (whether to purchase the commodity) can be examined, as Figure 1 shown. Figure 1 Schematic diagram showing the relationship between intervention and user response according to an example. In Figure 1 , users can be divided into four types according to whether they will purchase under the condition of whether there is an intervention. Among them, the determiners in the upper right corner are users who will purchase both without intervention and with intervention; the rejecters in the lower left corner are users who will not purchase both without intervention and with intervention; the persuadables in the upper left corner are users who will not purchase without intervention, but will purchase with intervention; the counteractives in the lower right corner are users who will purchase without intervention, but will not purchase with intervention. It can be seen that the goal of the lift model is to find the persuadables in the upper left corner and intervene in the persuadables to improve the marketing effect.
[0054] For a single user, assume that the user response when not intervened (the user is in the control group) is y 0 , and the user response when intervened (the user is in the intervention group) is y 1 . Then the true intervention increment τ of this user can be denoted as τ = y 1 - y 0 . The larger the true intervention increment τ of a certain user, the better the intervention effect on this user. However, in practical applications, the same user cannot be in the control group and the intervention group at the same time, that is, the values of y 0 and y 1 cannot be obtained simultaneously. Therefore, the true intervention increment τ cannot be used as the ground truth to directly perform supervised training on the uplift model.
[0055] In related technical solutions, models are usually used to fit the user responses of the control group and the intervention group respectively, and the difference between the user responses of the two user groups is used to obtain an estimated value of the true intervention increment τ (It should be noted that since the true intervention increment τ cannot be obtained, the following paragraphs all describe the estimated value of the true intervention increment τ and it is simply referred to as the intervention increment ). However, there is a natural deviation between the model training objective and practical applications in these technical solutions. Specifically, during the training process, the training objective of the model is to fit the user responses under different intervention conditions. Therefore, the models trained according to these solutions will pay more attention to the features related to predicting the user response y (0,1) , rather than the features related to predicting the intervention increment .
[0056] However, in practical applications, the ultimate goal of using the uplift model is to identify those users who are more sensitive to the intervention, and according to the sorting results of their intervention increments , different interventions are taken for users in different sorting positions. Correspondingly, the model performance evaluation metrics used are usually used to evaluate the sorting effect of the uplift model according to the predicted intervention increments , such as the area under the uplift curve (AUUC) and Top-K lift.
[0057] The area under the uplift curve is the area covered under the uplift curve. Figure 2 Fig. shows a schematic diagram of the uplift curve according to an example. As Figure 2As shown in the figure, the horizontal axis represents the cumulative coverage of users, that is, the first X% of users ranked in a certain order, and the vertical axis represents the cumulative intervention increment. The solid line represents the lift curve, and any point on the lift curve represents the cumulative intervention increment. The cumulative intervention increment after intervention for the first X% of users after sorting from largest to smallest. The dotted line is the baseline curve, which represents the cumulative intervention increment when randomly selecting users for intervention. In theory, due to random selection, the baseline curve is a straight line segment sloping upward, and intersects with the lift curve when the horizontal axis is 100%, that is, no matter what method is used to sort users, when all users are intervened, the cumulative intervention increment should be the same.
[0058] Since the lift curve is based on the intervention increment The calculation and plotting are done in descending order, so the cumulative increment of intervention at the beginning should increase at a faster rate and finally intersect with the baseline curve at 100%. That is, for an effective efficiency model, its improvement curve should always be above the baseline curve, and the area under the improvement curve ( Figure 2 The larger the area of the area enclosed by the lifting curve, the vertical dotted line at the 100% position and the horizontal axis, the greater the lifting curve is relative to the benchmark curve, indicating that the prediction effect of the synergistic model is better.
[0059] Top-K increment is the intervention increment predicted by the efficiency model for users After sorting from largest to smallest, the intervention effect after the intervention on the top K users is calculated using the user response y of the users in the intervention group among the top K users. 1 The average value of the user response y of the users belonging to the control group 0 The difference between the average values of . Similar to the area under the lift curve, the larger the Top-K increment, the better the prediction effect of the synergy model. For example, in commercial marketing, marketing intervention is carried out on the top K users (i.e., the aforementioned persuaders) ranked according to the synergy model to achieve better marketing results. It can be seen that the Top-K increment is also a performance evaluation indicator that is close to the needs of actual applications.
[0060] It can be seen that the performance evaluation indicators used in actual applications, such as the area under the lift curve and the Top-K increment, are directly related to the intervention increment. The model training goal in the related technical solution is to fit the user response. The deviation between this training goal and the actual application will lead to poor effect of the final trained model.
[0061] Based on the above analysis, the embodiments of this specification propose a method for training an incremental effect model, taking the performance evaluation index as the direct optimization objective, and training an incremental effect model that can directly predict the intervention increment. The incremental effect model.
[0062] First, describe the sample set used for training the incremental effect model in the embodiments of this specification, the specific form of the incremental effect model, and the loss function for training the incremental effect model determined based on the performance evaluation index.
[0063] The sample set can be denoted as where there are a total of n samples. For the i-th sample, x i represents the user attribute features, that is, the features related to the user himself, such as the region where the user is located, the category preference of the user, and so on; c i represents the scenario features where the user is located, such as the device model used by the user, the network IP address, and so on; t i ∈{0, 1} represents the user group that identifies whether the user is intervened, t i =0 represents that the user belongs to the control group (not intervened), t i =1 represents that the user belongs to the intervention group (intervened); y i represents the user response, which can be a discrete variable (such as whether to click a button or a page), or a continuous variable (such as the total consumption amount of the current marketing activity).
[0064] As mentioned above, since for the user corresponding to a single sample, its true intervention increment cannot be obtained, so the embodiments of this specification use the incremental effect model to estimate it, obtaining an estimated value and abbreviating the estimated value as the intervention increment
[0065] The incremental effect model trained in the embodiments of this specification can be any type of neural network, such as a deep neural network, which is not limited here. The incremental effect model receives user-related features (including the user attribute feature x i and the user scenario feature c i ) and predicts the intervention increment of this user
[0066] The performance evaluation indexes for evaluating the performance of the incremental effect model in the related art include the area under the lift curve and the Top-K increment. For the area under the lift curve AUUC, its calculation formula can be written in the form shown in formula (1):
[0067]
[0068] Among them, is the average intervention increment, which is the difference between the average user response of the intervention group and the average user response of the control group, as shown in formula (2):
[0069]
[0070] where n t is the number of samples in the intervention group in the sample set, and n c is the number of samples in the control group in the sample set.
[0071] rank(i) in formula (1) represents the ranking position of sample i after sorting all samples according to the intervention increment predicted by the synergy model from large to small, and the value range is from 1 to n. g(i) represents the gain of the cumulative intervention increment of sample i for the vertical axis of the lift curve, as shown in formula (3):
[0072]
[0073] The part (n - rank(i) + 1) in formula (1) represents the number of times sample i is cumulatively calculated when calculating the area under the lift curve, that is, the higher the ranking (the smaller the rank() value) of the sample, the more times it is calculated. The sum of the products of the calculation times of each sample and the corresponding g(i) is the actual area under the lift curve. is the normalization coefficient.
[0074] For and Top-K increment lift@K, its calculation formula can be written in the form shown in formula (4):
[0075]
[0076] Only samples with rank(i) ≤ K are considered in formula (4), that is, the first K samples after sorting all samples according to the intervention increment predicted by the synergy model from large to small. The first term is the average user response of the intervention group among the first K samples, and the second term is the average user response of the control group among the first K samples.
[0077] According to formula (1) and formula (4), it can be seen that in the process of calculating the area under the lift curve and the Top-K increment, it is necessary to sort all samples according to the intervention increment predicted by the synergy model from large to small. The existence of this sorting process makes these two performance evaluation indicators unable to be directly used to optimize the synergy model. Specifically, due to the existence of the sample ranking rank(i) in formula (1) and formula (4), and rank(i) itself is a non-differentiable function, the area under the lift curve and the Top-K increment determined based on rank(i) are also non-differentiable and cannot directly use the gradient descent method to adjust the parameters in the synergy model.
[0078] Based on the above analysis, in the embodiments of this specification, rank(i) is smoothed to make it differentiable. The conventional rank(i) in the related art can be written in the form of formula (5).
[0079]
[0080] Among them, As shown in formula (6):
[0081]
[0082] That is, it counts the number of samples in the sample set whose intervention increment is less than the intervention increment of sample i, and then subtracts this number from the total number n of samples in the sample set.
[0083] The function shown in formula (6) is not differentiable, resulting in the non-differentiability of the conventional rank(i) in the related art, and further resulting in the non-differentiability of both the area under the lift curve and the Top-K increment as functions. According to the research results, the inventor of the present invention smooths the function in formula (6) to obtain the function shown in formula (7).
[0084]
[0085] Among them, e is the natural constant, and α is a preset hyperparameter. The larger the value of α, the closer ψ(i, j) is to φ(i, j). It can be seen from formula (7) that ψ(i, j) is a smooth differentiable function with a value range between 0 and 1, and The difference between is larger, the closer the value of ψ(i, j) is to 1, and vice versa, the closer it is to 0, and it is similar to φ(i, j) in formula (6) in the mapping relationship. ψ(i, j) can represent the relative ranking value between sample i and sample j, and ψ(i, j) can be used to replace φ(i, j) to calculate rank(i) to obtain the optimized ranking function As shown in formula (8).
[0086]
[0087] Using the optimized ranking function can obtain the differentiable optimized area under the lift curve and the optimized Top-K increment. Specifically, for the area under the lift curve AUUC, after using ψ(i, j) to replace φ(i, j), the optimized area under the lift curve is obtained as shown in formula (9):
[0088]
[0089] Optimize the area under the curve is a differentiable function and can be calculated based on the intervention increment User response i and user group i Calculation, and then the area under the curve can be improved based on optimization The negative value of is used to update the parameters in the enhancement model using the gradient descent method.
[0090] For the Top-K incremental lift@K, by introducing φ(i,j) in formula (6), the condition i∈{i|rank(i)≤K} in formula (4) can be replaced, and formula (4) can be rewritten as shown in formula (10):
[0091]
[0092] Among them, K represents the intervention increment predicted by the efficiency model for each sample. The Kth sample after sorting from largest to smallest. Using ψ(i,j) instead of φ(i,j) gives the optimized Top-K increment As shown in formula (11):
[0093]
[0094] Optimizing Top-K increment is a differentiable function and can be calculated based on the intervention increment User response i and user group i Calculation, and then according to the optimization of Top-K increment The negative value of is used to update the parameters in the enhancement model using the gradient descent method.
[0095] The above optimization improvement area under the curve and optimization Top-K increment can be collectively referred to as optimization performance evaluation indicators.
[0096] It is understandable that in other embodiments, the function φ(i, j) in formula (6) may be subjected to other smoothing processing, as long as the function after processing is smooth and differentiable. For example, in one embodiment, the function in formula (6) may be subjected to smoothing processing to obtain a function as shown in formula (12).
[0097]
[0098] Where tanh() is the hyperbolic tangent function, and β is a preset hyperparameter. From formula (12), we can see that δ(i,j) is a smooth differentiable function with a value range between 0 and 1, and and The larger the difference between them, the closer the value of δ(i,j) is to 1, and vice versa, the closer it is to 0, and the mapping relationship is similar to φ(i,j) in formula (6). φ(i,j) can represent the relative ranking value between sample i and sample j. δ(i,j) can be used to replace φ(i,j) to calculate rank(i), and the corresponding optimized ranking function can be obtained. The area under the optimization improvement curve and / or the optimization Top-K increment are further calculated, and the parameters in the efficiency improvement model are updated accordingly, which will not be described in detail here.
[0099] Based on the above analysis results, the embodiment of this specification proposes a method for training a synergy-enhancing model. Figure 3 A flowchart of a method for training a synergy-enhancing model according to an embodiment is shown. The execution subject of the method can be any platform or server or device cluster with computing and processing capabilities. Figure 3 As shown, the method at least includes: step 302, obtaining a sample set, wherein any sample includes user-related features, a user group that identifies whether the user is intervened, and a user response; step 304, inputting the user-related features of each sample into the synergy model to obtain the intervention increment of each sample, which indicates the improvement of the user response by the intervention; step 306, for any target sample, determining the ranking factor of the target sample according to the relative ranking value of the target sample and at least one other sample, wherein the relative ranking value of the target sample and any other sample is obtained by inputting the difference between the target intervention increment of the target sample and the other intervention increments of the other samples into a predetermined indicator function, and the indicator function is a monotonically increasing smooth function; step 308, updating the parameters of the synergy model in the direction of increasing the function value of the target function, wherein the function value of the target function is determined according to the user group, the user response and the ranking factor of each sample.
[0100] The specific execution process of each of the above steps is described below.
[0101] First, in step 302, a sample set is obtained, wherein any sample includes user-related features, a user group identifying whether the user is intervened, and a user response.
[0102] User-related features can be various features related to the user. In one embodiment, the user-related features include user attribute features x i , and the user's scene characteristics c i .
[0103] It is understandable that the user-related features may be in the form of representation vectors (embeddings), which are obtained by processing the original features of the user through a representation learning method (eg, an encoder).
[0104] Then, in step 304, the user-related features of each sample are input into the synergistic model to obtain the intervention increment of each sample, which indicates the improvement of the intervention on the user response.
[0105] It can be understood that the intervention increment of each sample predicted by the synergistic model in step 304 is an estimated value of the true intervention increment of each sample.
[0106] Next, in step 306, for any target sample, the ranking factor of the target sample is determined according to the relative ranking value between the target sample and at least one other sample. Among them, the relative ranking value between the target sample and any other sample is obtained by inputting the difference between the target intervention increment of the target sample and the other intervention increment of the other sample into a predetermined indication function, and the indication function is a monotonically increasing smooth function.
[0107] Finally, in step 308, in the direction of increasing the function value of the objective function, the parameters of the synergistic model are updated, where the function value of the objective function is determined according to the user group, user response of each sample, and the ranking factor.
[0108] The target sample can be, for example, sample i. The target intervention increment of the target sample can be The other intervention increment of the other sample j can be The difference between the two can be the value obtained by subtracting the other intervention increment from the target intervention increment, that is By inputting this difference into a predetermined indication function to obtain the relative ranking value between the target sample i and the other sample j.
[0109] In one embodiment, the relative ranking value between the target sample and any other sample in step 306 is determined through the following steps:
[0110] Multiply the difference by a preset first coefficient as the independent variable and input it into a preset smooth function to obtain the relative ranking value.
[0111] Specifically, the smooth function includes at least one of the following: Sigmoid function, tanh function.
[0112] In a more specific embodiment, when the smooth function is the Sigmoid function, the relative ranking value between the target sample i and the other sample j determined according to the Sigmoid function can be as shown in formula (7).
[0113] In this embodiment, the first coefficient can be α, and the predetermined indication function in step 306 can be Sigmoid(αx). The difference Input into the predetermined indicator function to obtain the relative ranking value of the target sample i and the other sample j
[0114] In another more specific embodiment, when the smoothing function is a tanh function, the relative ranking value of the target sample i and the other sample j determined by the tanh function can be as shown in formula (12).
[0115] In this embodiment, the first coefficient may be β, and the predetermined indicator function in step 306 may be The difference Input into the predetermined indicator function to obtain the relative ranking value of the target sample i and the other sample j
[0116] In other embodiments, the difference can also be directly It is directly input as an independent variable into the preset smoothing function to obtain the relative ranking value.
[0117] In one embodiment, in order to make the optimized performance evaluation index obtained based on the indicator function more consistent with the original performance evaluation index, the function value of the indicator function in step 306 is between 0 and 1.
[0118] After obtaining the relative ranking value between the target sample i and at least one other sample, the ranking factor of the target sample may be determined according to the relative ranking value between the target sample i and at least one other sample.
[0119] For the convenience of description, the following ranking relative values are all represented by ψ(i, j) as shown in formula (7), but this does not constitute a limitation on the protection scope of the embodiments of this specification. It is understandable that at least δ(i, j) as shown in (12) can also be used as the ranking relative value.
[0120] In one embodiment, the loss function used in the training synergy model is based on the area under the lifting curve. In this embodiment, the at least one other sample in step 306 is each sample j other than the target sample, j∈n, j≠i. Determining the ranking factor of the target sample in step 306 includes:
[0121] The ranking factor ∑ of the target sample is obtained by adding 1 to the sum of the relative ranking values of the target sample i and each sample j except the target sample. j∈n,j≠i ψ(i,j)+1.
[0122] In this embodiment, the objective function in step 308 is the area under the lifting curve function, and its function value is determined by the following steps:
[0123] Determine the uplift share g(i) of the target sample. If the target sample belongs to the intervention group, i.e., t i = 1, its uplift share is the user response of the target sample divided by the total number of samples in the intervention group, i.e., If the target sample belongs to the control group, i.e., t i = 0, its uplift share is the user response of the target sample divided by the total number of samples in the control group and then taking the opposite, i.e.,
[0124] Then, using the ranking factor ∑ j∈n,j≠i ψ(i,j)+1 as the weight, perform a weighted sum of the uplift shares g(i) of each sample to determine the function value of the objective function.
[0125] In a more specific embodiment, the result of the weighted sum can also be normalized, specifically by multiplying to improve the stability during model training. This can obtain the objective function as shown in formula (9).
[0126] In another embodiment, the loss function used to train the uplift model is based on Top-K increment. In this embodiment, the at least one other sample in step 306 is the K-th sample among all samples sorted in descending order of intervention increment, i.e., sample K. Determining the ranking factor of the target sample in step 306 includes:
[0127] Subtract the intervention increment of the K-th sample from the target intervention increment to obtain the difference and input it into the indicator function to determine the ranking factor ψ(i,K) of the target sample.
[0128] In this embodiment, the objective function in step 308 is the top-K intervention effect function (i.e., Top-K increment), and its function value is determined through the following steps:
[0129] Based on the ranking factor ψ(i,K) and user response y i of each intervention group sample in the intervention group (t i = 1), determine the first average intervention increment of the intervention group.
[0130] Based on the ranking factor ψ(i,K) and user response y i of each control group sample in the control group (t i = 0), determine the second average intervention increment of the control group.
[0131] Based on the difference between the first average intervention increment and the second average intervention increment, determine the function value of the objective function.
[0132] Specifically, the first average intervention increment of the intervention group is determined, including:
[0133] Using the ranking factor as the weight, the user responses of each intervention group sample are weighted and summed to obtain the first weighted sum
[0134] Sum the ranking factors of samples in each intervention group to obtain the first smoothed quantity
[0135] Determine a first average intervention increment based on the ratio of the first weighted sum to the first smoothed quantity
[0136] Specifically, the second average intervention increment of the control group is determined, including:
[0137] Using the ranking factor as the weight, the user responses of each control group sample are weighted and summed to obtain the second weighted sum
[0138] Sum the ranking factors of each control group sample to get the second smoothing quantity
[0139] Determine a second average intervention increment based on the ratio of the second weighted sum to the second smoothed quantity
[0140] The function value of the objective function is determined according to the difference between the first average intervention increment and the second average intervention increment, as shown in formula (11).
[0141] After determining the objective function, the parameters of the synergy model may be adjusted according to the objective function. In one embodiment, step 308 shows updating the parameters of the synergy model in the direction of increasing the function value of the objective function, including:
[0142] The inverse of the objective function is used as the loss function, and the parameters of the synergistic model are updated based on the gradient descent method.
[0143] After the efficiency enhancement model is trained according to steps 302 to 308, the efficiency enhancement model can be used to predict each user in a user data set to obtain the intervention increment of each user. Then, the users are sorted in descending order according to the intervention increment, and the top ranked users are selected for intervention to achieve better intervention effect.
[0144] The method for training an efficiency-enhancing model proposed in the embodiments of this specification introduces loss functions derived from the area under the lift curve and the Top-K increment. These two loss functions can be directly optimized using common gradient descent methods in a deep learning framework. Through these two loss functions, the ranking performance or the head improvement performance can be directly optimized during the training of the efficiency-enhancing model. This method is applicable to classification and regression problems and can be combined with other functions in related technologies to achieve better training and prediction effects.
[0145] According to an embodiment of another aspect, there is also provided an apparatus for training an efficiency-enhancing model. Figure 4 A schematic block diagram showing an apparatus for training an efficiency-enhancing model according to an embodiment. This apparatus can be deployed in any device, platform, or cluster of devices with computing and processing capabilities. As Figure 4 shown, the apparatus 400 includes:
[0146] An acquisition unit 402, configured to acquire a sample set, where any sample includes user-related features, a user group indicating whether the user is intervened, and a user response;
[0147] An intervention increment prediction unit 404, configured to input the user-related features of each sample into the efficiency-enhancing model to obtain an intervention increment for each sample, which indicates the improvement of the intervention on the user response;
[0148] A ranking factor determination unit 406, configured to determine a ranking factor for any target sample according to the relative ranking value of the target sample and at least one other sample. Among them, the relative ranking value of the target sample and any other sample is obtained by inputting the difference between the target intervention increment of the target sample and the other intervention increment of the other sample into a predetermined indication function, and the indication function is a monotonically increasing smooth function;
[0149] A model training unit 408, configured to update the parameters of the efficiency-enhancing model in the direction in which the function value of the target function increases, where the function value of the target function is determined according to the user group, the user response, and the ranking factor of each sample.
[0150] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in any of the above embodiments.
[0151] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor. Among them, an executable code is stored in the memory, and when the processor executes the executable code, the method described in any of the above embodiments is implemented.
[0152] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the corresponding description in the method embodiments.
[0153] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0154] It can be understood that before or when using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner in accordance with relevant laws and regulations, and the user's authorization will be obtained.
[0155] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to the software or hardware such as an electronic device, application program, server, or storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0156] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0157] It can be understood that the above processes of notifying and obtaining the user's authorization are only illustrative and do not limit the implementation manners of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0158] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0159] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0160] The specific embodiments described above have further elaborated on the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for training an efficacy enhancement model, comprising: Obtaining a sample set, where any sample includes user-related features, a user group indicating whether the user is intervened, and a user response; Inputting the user-related features of each sample into the efficacy enhancement model to obtain the intervention increment of each sample, which indicates the improvement of the user response by the intervention; For any target sample, determining the ranking factor of the target sample according to the relative ranking value between the target sample and at least one other sample, where the relative ranking value between the target sample and any other sample is obtained by inputting the difference between the target intervention increment of the target sample and the other intervention increment of the other sample into a predetermined indication function, and the indication function is a monotonically increasing smoothing function; Updating the parameters of the efficacy enhancement model in the direction of increasing the function value of the objective function, where the function value of the objective function is determined according to the user group, user response, and the ranking factor of each sample.
2. The method according to claim 1, wherein, The user-related features include User attribute features and scene features where the user is located.
3. The method according to claim 1, wherein, The relative ranking value between the target sample and any other sample is determined through the following steps: Using the difference multiplied by a preset first coefficient as the independent variable and inputting it into a preset smoothing function to obtain the relative ranking value.
4. The method according to claim 3, wherein, The smoothing function includes at least one of the following: Sigmoid function, tanh function.
5. The method according to claim 1, wherein The function value of the indication function is between 0 and 1.
6. The method according to claim 1, wherein, The at least one other sample is each sample except the target sample; Determining the ranking factor of the target sample includes: Adding 1 to the sum of the relative ranking values between the target sample and each sample except the target sample to obtain the ranking factor of the target sample.
7. The method according to claim 6, wherein, The objective function is the area under the lift curve function, and its function value is determined through the following steps: Determining the lift share of the target sample, where if the target sample belongs to the intervention group, its lift share is the user response of the target sample divided by the total number of samples in the intervention group; if the target sample belongs to the control group, its lift share is the user response of the target sample divided by the total number of samples in the control group and taking the opposite number; Using the ranking factor as the weight, performing a weighted sum of the lift shares of each sample to determine the function value of the objective function.
8. The method according to claim 1, wherein The at least one other sample is the Kth sample among all samples sorted in descending order of intervention increment; determining the ranking factor of the target sample includes: Inputting the difference between the target intervention increment and the intervention increment of the Kth sample into the indication function to determine the ranking factor of the target sample.
9. The method according to claim 8, wherein The objective function is the top K intervention effect function, and its function value is determined through the following steps: Determining the first average intervention increment of the intervention group according to the ranking factor and user response of each intervention group sample in the intervention group; Determining the second average intervention increment of the control group according to the ranking factor and user response of each control group sample in the control group; Determining the function value of the objective function according to the difference between the first average intervention increment and the second average intervention increment.
10. The method according to claim 9, wherein, Determining the first average intervention increment of the intervention group includes: Using the ranking factor as the weight, performing a weighted sum of the user responses of each intervention group sample to obtain the first weighted sum; The ranking factors of samples in each intervention group were summed to obtain the first smoothed quantity; A first average intervention increment is determined based on a ratio of the first weighted sum to the first smoothed quantity.
11. The method according to claim 9, wherein, Determine the second mean intervention increment for the control group, including: Taking the ranking factor as the weight, perform weighted summation on the user responses of each control group sample to obtain a second weighted sum; Sum the ranking factors of each control group sample to obtain the second smoothed quantity; A second average intervention increment is determined based on a ratio of the second weighted sum to the second smoothed quantity.
12. The method according to claim 1, wherein, In the direction where the function value of the objective function increases, updating the parameters of the synergy model includes: The inverse of the objective function is used as the loss function, and the parameters of the synergistic model are updated based on the gradient descent method.
13. A device for training a synergistic model, comprising: An acquisition unit configured to acquire a sample set, wherein any sample includes user-related features, a user group identifying whether the user is interfered with, and a user response; An intervention increment prediction unit configured to input the user-related features of each sample into the efficiency enhancement model to obtain the intervention increment of each sample, which indicates the improvement of the user response by the intervention; a ranking factor determination unit configured to determine, for any target sample, a ranking factor of the target sample according to a relative ranking value between the target sample and at least one other sample, wherein the relative ranking value between the target sample and any other sample is obtained by inputting a difference between a target intervention increment of the target sample and other intervention increments of the other sample into a predetermined indicator function, wherein the indicator function is a monotonically increasing smooth function; The model training unit is configured to update the parameters of the synergy model in the direction of increasing the function value of the objective function, wherein the function value of the objective function is determined according to the user group, user response and ranking factor of each sample.
14. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 12.
15. A computing device, comprising a memory and a processor, wherein, The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 12 is implemented.