Deviation correction inference method and device, equipment and storage medium

By calculating the exposure probability of recommended content in the randomized controlled trial and eliminating the bias in grabbing effect, the problem of large deviation in causal inference statistics is solved, and the accuracy of causal inference is improved.

CN120020851APending Publication Date: 2025-05-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311544854.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

In randomized controlled trials, when there is a grab effect, the deviation of the causal inference statistics is large and the variance is high, resulting in a decrease in the accuracy of causal inference.

Method used

By calculating the first exposure probability and the second exposure probability corresponding to the recommended content in the recall pool corresponding to the consumer account, the index result data is obtained, and the grab effect deviation in the causal effect estimator is eliminated based on these probabilities, and the causal effect estimator after eliminating the deviation is obtained.

Benefits of technology

Effectively eliminate the bias in grabbing effect in randomized controlled trials and improve the accuracy and stability of causal inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020851A_ABST
    Figure CN120020851A_ABST
Patent Text Reader

Abstract

The invention discloses a deviation correction inference method and device, equipment and a storage medium, and belongs to the field of causal inference. The method comprises the following steps: calculating a first exposure probability and a second exposure probability respectively corresponding to K recommended contents in a recall pool corresponding to a first consumer account; obtaining index result data of the first consumer account for the K recommended contents; based on a first exposure probability and a second exposure probability corresponding to the K recommended contents respectively, eliminating a snatching effect deviation in a causal effect estimator corresponding to the index result data to obtain a causal effect estimator corresponding to the first consumer account after the deviation is eliminated; and obtaining a causal inference statistic corresponding to the random control test based on the causal effect estimators corresponding to the plurality of consumer accounts after the deviation is eliminated. According to the method and the device, the causal effect estimator after deviation elimination is obtained by eliminating the robbery effect deviation in the random control test, and the causal inference statistic corresponding to the random control test can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of causal inference, and particularly to a bias correction inference method, device, equipment, and storage medium. Background Art

[0002] A randomized controlled trial is a commonly used experimental design method for comparing the effect differences of an experimental group and a control group on a certain variable. In a randomized controlled experiment, participants are randomly assigned to the experimental group and the control group, and the performance of the participants on a certain index is observed and measured, so as to evaluate the causal effect between different strategies. By comparing the differences between the experimental group and the control group, a causal inference can be obtained, that is, the influence of a certain treatment on the variable.

[0003] In the related art, a randomized controlled experiment is a standard for evaluating the causal effect of a strategy. The mean difference of the indexes in the experimental group and the control group can be used for causal inference, and the mean variance of the control group and the mean variance of the experimental group are used to estimate the variance of the causal inference statistic. In the absence of a poaching effect, it can be ensured that the causal inference statistic is an unbiased estimate of the overall causal effect.

[0004] When there is a poaching effect, the causal inference statistic has the poor properties of large bias and high variance. How to eliminate the bias of the causal inference statistic so as to improve the accuracy of causal inference is an urgent problem to be solved currently. Summary of the Invention:

[0005] This application provides a bias correction inference method, device, equipment, and storage medium, and the technical solutions are as follows:

[0006] According to one aspect of this application, a bias correction inference method is provided, and the method includes:

[0007] Calculate the first exposure probability and the second exposure probability respectively corresponding to K recommended contents in the recall pool corresponding to the first consumer account. The first exposure probability is used to indicate the exposure probability of the current recommended content when assuming that the K recommended contents are all recommended contents in the experimental group of a randomized controlled experiment, and the second exposure probability is used to indicate the exposure probability of the current recommended content when assuming that the K recommended contents are all recommended contents in the control group of a randomized controlled experiment;

[0008] Obtain the index result data of the first consumer account for the K recommended contents, and the first consumer account is any one of multiple consumer accounts;

[0009] Based on the first exposure probability and the second exposure probability respectively corresponding to the K recommended contents, eliminate the poaching effect bias in the causal effect estimator corresponding to the index result data, and obtain the causal effect estimator after bias elimination corresponding to the first consumer account;

[0010] Based on the debiased causal effect estimators corresponding to the multiple consumer accounts, obtain the causal inference statistic corresponding to the randomized controlled trial.

[0011] According to another aspect of the present application, a bias correction inference device is provided, and the device includes:

[0012] A calculation module, configured to calculate a first exposure probability and a second exposure probability corresponding to each of the K recommended contents in the recall pool of the first consumer account, where the first exposure probability is used to indicate the exposure probability of the current recommended content when assuming that all the K recommended contents are in the experimental group of the randomized controlled trial, and the second exposure probability is used to indicate the exposure probability of the current recommended content when assuming that all the K recommended contents are in the control group of the randomized controlled trial;

[0013] An acquisition module, configured to acquire the metric result data of the first consumer account for the K recommended contents, where the first consumer account is any one of the multiple consumer accounts;

[0014] An elimination module, configured to eliminate the poaching effect bias in the causal effect estimator corresponding to the metric result data based on the first exposure probability and the second exposure probability corresponding to each of the K recommended contents, and obtain the debiased causal effect estimator corresponding to the first consumer account;

[0015] A processing module, configured to obtain the causal inference statistic corresponding to the randomized controlled trial based on the debiased causal effect estimators corresponding to the multiple consumer accounts.

[0016] According to another aspect of the present application, a computer device is provided, and the computer device includes a processor and a memory, and a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the bias correction inference method as described above.

[0017] According to another aspect of the present application, a computer-readable storage medium is provided, and an executable instruction is stored in the computer-readable storage medium, and the executable instruction is loaded and executed by the processor to implement the bias correction inference method as described above.

[0018] According to another aspect of the present application, a computer program product is provided, and the computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium, and the processor reads and executes the computer instructions from the computer-readable storage medium to implement the bias correction inference method as described above.

[0019] The beneficial effects brought by the technical solution provided by the present application at least include:

[0020] The randomized controlled trial is divided into an experimental group and a control group: by calculating the first exposure probability and the second exposure probability corresponding to the recommended content in the recall pool corresponding to the first consumer account; obtaining the index result data of the first consumer account for the recommended content; based on the first exposure probability and the second exposure probability, eliminating the poaching effect bias in the causal effect estimator corresponding to the index result data to obtain the causal effect estimator after bias elimination corresponding to the first consumer account; based on the causal effect estimators after bias elimination corresponding to multiple consumer accounts, obtaining the causal inference statistic corresponding to the randomized controlled trial. In this application, by calculating the exposure probability of the recommended content in the recall pool, combining the difference between the first exposure probability and the second exposure probability with the index result data of the recommended content, eliminating the poaching effect bias in the randomized controlled trial, obtaining the causal effect estimator after bias elimination, and based on the causal effect estimator, the causal inference statistic corresponding to the randomized controlled trial can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0022] Figure 1 Shows a schematic diagram of a computer system provided by an exemplary embodiment of the present application;

[0023] Figure 2 Shows a flowchart of a bias correction inference method provided by an exemplary embodiment of the present application;

[0024] Figure 3 Shows a flowchart of a bias correction inference method provided by an exemplary embodiment of the present application;

[0025] Figure 4 Shows a flowchart of a bias correction inference method provided by an exemplary embodiment of the present application;

[0026] Figure 5 Shows a flowchart of a bias correction inference method provided by an exemplary embodiment of the present application;

[0027] Figure 6 Shows a flowchart of a bias correction inference method provided by an exemplary embodiment of the present application;

[0028] Figure 7 Shows a flowchart of a bias correction inference method provided by an exemplary embodiment of the present application;

[0029] Figure 8Shows a flowchart of a deviation correction inference method provided by an exemplary embodiment of the present application;

[0030] Figure 9 Shows a flowchart of a deviation correction inference method provided by an exemplary embodiment of the present application;

[0031] Figure 10 Shows a flowchart of a deviation correction inference method provided by an exemplary embodiment of the present application;

[0032] Figure 11 Shows a flowchart of a deviation correction inference method provided by an exemplary embodiment of the present application;

[0033] Figure 12 Shows a flowchart of a deviation correction inference method provided by an exemplary embodiment of the present application;

[0034] Figure 13 Shows a flowchart of a deviation correction inference method provided by an exemplary embodiment of the present application;

[0035] Figure 14 Shows a schematic diagram of a deviation correction inference method provided by an exemplary embodiment of the present application;

[0036] Figure 15 Shows a structural block diagram of a deviation correction inference device provided by an exemplary embodiment of the present application;

[0037] Figure 16 Shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0038] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0039] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0040] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0041] It should be noted that before and during the collection of relevant user data in this application, a prompt interface, a pop-up window, or voice prompt information can be displayed. The prompt interface, pop-up window, or voice prompt information is used to prompt the user that their relevant data is currently being collected, so that this application only starts to execute the relevant steps for obtaining the user's relevant data after obtaining the confirmation operation issued by the user for the prompt interface or pop-up window. Otherwise (that is, when the confirmation operation issued by the user for the prompt interface or pop-up window is not obtained), the relevant steps for obtaining the user's relevant data are ended, that is, the relevant data of the user is not obtained. In other words, all user data collected by this application is collected with the consent and authorization of the user, and the collection, use, and processing of relevant user data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0042] First, introduce the relevant terms involved in this application:

[0043] AB test: That is, a randomized controlled trial, which is a commonly used experimental design method for comparing the effects of two or more different strategies, treatments, or interventions on a certain indicator or outcome. In an AB test, the experimental subjects are randomly divided into two or more groups. One group receives the new strategy or treatment, called the experimental group (Group A), while the other group continues to use the old strategy or treatment, called the control group (Group B).

[0044] Snatching effect: It refers to the mutual influence between the experimental group and the control group. For example, the authors in the experimental group may attract the users in the control group, thus affecting the experimental results of the control group.

[0045] Causal inference statistic: It is a statistic used to estimate the causal effect, and it infers the impact of the causal effect by comparing the results of the experimental group and the control group. Exemplarily, assume to evaluate the impact of a new recommendation algorithm of an e-commerce platform on the user's purchase behavior. The causal inference statistic can be the difference in purchase rates between the experimental group and the control group. The difference between the purchase rate of the authors in the experimental group and the purchase rate of the authors in the control group can be calculated as the causal inference statistic.

[0046] Utility function: A function used to measure the user's preference or satisfaction with the recommended content. The utility function can evaluate the effect of the video according to different factors and characteristics, such as the user's click-through rate, viewing duration, etc. For example, when studying the effect of a video recommendation system, there are two strategies to choose from. Strategy 1: Add a video cover with a decorative frame to the recommended content. The utility function corresponding to Strategy 1 can be expressed as θ 1 (u, v), which is evaluated according to the user's click-through rate. Strategy 2: Use a video cover without adding a decorative frame for the recommended content. The utility function corresponding to Strategy 2 can be expressed as θ2 (u, v) can be evaluated based on the user's click-through rate. By calculating the utility function under different strategies, the effects of different strategies on the user-video combination can be compared. If θ 1 (u, v) > θ 2 (u, v), it indicates that the recommendation under Strategy 1 is more in line with the user's interests and preferences.

[0047] Recommendation system: An information filtering system that provides personalized recommended content for consumer accounts by analyzing the historical behaviors, interests, and preferences of consumer accounts.

[0048] Recall pool: In a recommendation system, it refers to a set of content that is of interest to consumer accounts selected from the candidate video set according to the recall strategy and is used for further processing and sorting. Exemplarily, in a recommendation system, some recommended content that consumers may consume is recalled to form a recall pool, and the recommended content in the recall pool is sorted according to preset rules.

[0049] Two-sided market: It refers to a situation in a certain platform or scenario where the performance of two different groups or wholes needs to be simultaneously concerned. Exemplarily, in the video account scenario of WeChat, the two-sided market refers to the performance of two groups: the viewers (consumers) of the video account and the creators of the video account or the content they create (producers).

[0050] Figure 1 The architecture schematic diagram of a computer system provided by an embodiment of the present application is shown. The computer system may include: a terminal 100 and a server 200.

[0051] The terminal 100 may be an electronic device such as a mobile phone, a tablet computer, an in-vehicle terminal (car computer), a wearable device, a personal computer (PC), an in-vehicle terminal, etc. A client for running a target application may be installed in the terminal 100. The target application may be an application for simulating tomographic inversion, and the present application does not make any limitations in this regard. Additionally, the present application does not make any limitations on the form of the target application, including but not limited to applications (Apps), applets, etc. installed in the terminal 100, and may also be in the form of a web page.

[0052] The server 200 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, Content Delivery Network (CDN), and a cloud server providing basic cloud computing services such as a big data and artificial palm image recognition platform. The server 200 can be the background server of the above target application program, and is used to provide background services for the client of the target application program.

[0053] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or a local area network to realize data calculation, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, and can form a resource pool, which can be used on demand, flexibly and conveniently. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system backup support, which can only be achieved through cloud computing.

[0054] In some embodiments, the above server can also be implemented as a node in a blockchain system. Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain is essentially a decentralized database, a string of data blocks generated by using cryptographic methods, and each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0055] Communication can be carried out between the terminal 100 and the server 200 through a network, such as a wired or wireless network.

[0056] In the correction and inference method provided by the embodiments of the present application, the execution subject of each step can be a computer device, and a computer device refers to an electronic device with data calculation, processing, and storage capabilities. Figure 1Taking the solution implementation environment shown as an example, the deviation correction inference method can be executed by the terminal 100, or by the server 200, or by the interaction and cooperation of the terminal 100 and the server 200. This application does not make any limitations in this regard.

[0057] In the embodiments of the present application, when there is a snatching effect, an exposure probability model is created to predict the exposure probability of each video in the recall pool; then a result model is created to calculate the estimated value of the result model, and the deviation is eliminated through the exposure probability model and the result model; by calculating the expected value of the estimated value of the result model, the causal inference statistic after eliminating the deviation is obtained; based on the causal inference statistic after eliminating the deviation, a preset method is used for variance inference to obtain the variance of the causal inference statistic.

[0058] Step 1, create an exposure probability model;

[0059] When conducting an AB test on the author side in a two-sided market, a part of the authors, as the experimental group, will receive the new strategy, and another part of the authors, as the control group, will not receive the new strategy. The new strategy may cause some authors to snatch the resources of other authors, such as the attention of the audience, viewing time, total consumption, etc.

[0060] The snatching effect is mainly reflected in the impact on the exposure probability of videos in the same recall pool. The recall pool for each request of consumer u is V u , V u contains K videos. Use W u ∈{0,1} K to represent the strategy status of each video in the recall pool. Among them, 0 and 1 indicate whether the video is selected. 1 is the experimental group and receives the new strategy, and 0 is the control group and does not receive the new strategy. Use e u ∈{0,1} K to represent the exposure status of the videos in the recall pool. 1 means the video is exposed, and 0 means the video is not exposed.

[0061] In some embodiments, use θ l (u,v) to represent the utility function under strategy l. Its input is a consumer u and a video v, and the output is the utility of this group of user-video. The utility function is a function used to measure the preference or satisfaction of a user for a certain video. The utility function can evaluate the effect of a video according to different factors and characteristics, such as the click-through rate and viewing duration of the user.

[0062] In some embodiments, when there is only one experimental group and one control group, the selection model in the recommendation system will create an exposure probability model for the videos in the recall pool, specifically:

[0063] Formula (1)

[0064]

[0065] Among them, W v ∈ {0, 1} represents the policy-affected situation of video v. W v being 1 indicates that video v is affected by a new policy, and W v being 0 indicates that video v is not affected by the new policy. p(v|u, V u , W u ) is a conditional probability, which represents the probability that video v is exposed under the conditions of consumer u, recall pool V u and the policy-affected situation W u . This probability can be used to measure the exposure degree and competition situation of the video in the recommendation system. u represents the consumer, V u represents the recall pool, v represents the video in the recall pool, and W u represents the policy-affected situations of each video in the recall pool. 1(W v = 1) is an indicator function, and indicator functions are usually used to represent whether a certain condition is met. 1(W v = 1) is used to judge the policy-affected situation of video v in the recall pool of consumer u. When W v = 1, the value of the indicator function is 1.

[0066] In some embodiments, there can be multiple experimental groups. When there are L experimental groups, W v ∈ {0, 1, 2, 3, …, L} K is used to represent whether the video is affected by the policy of the l-th experimental group.

[0067] For the first video in the recall pool:

[0068] Formula (2)

[0069]

[0070] Among them, this is a probability model for the first video in the recall pool. p(1|u, V u , W u ; θ) is a conditional probability, which represents the probability that consumer u watches the first video under the conditions of the parameters of the θ model, consumer u, recall pool V u and the policy-affected situation W u . This probability can be used to measure the exposure degree and competition situation of the video in the recommendation system. Among them, θ is the parameter of the model, which is used to represent the influence of different experimental group policies on consumer u and the first video in recall pool V u . 1[W u,1 = l] is an indicator function, and 1[W u,1= l] is used to determine the policy application situation of the first video in the recall pool, W u,1 When = l, it means that when the first video in the recall pool is subject to policy l, the value of the indicator function is 1. θ l (u, V u,1 ) represents the effect function of the first video in the recall pool being subject to policy l.

[0071] For the k = 2, …, K-th videos:

[0072] Formula (3)

[0073]

[0074] Among them, this is the probability model for the k-th video in the recall pool, p(k|u, V u , W u ; θ) is a conditional probability, which represents the probability that consumer u watches the k-th video under the parameters of the θ model, consumer u, recall pool V u and the policy application situation W u . Among them, θ is the parameter of the model, used to represent the influence of different experimental group policies on consumer u and the k videos in recall pool V u 1[W u,k = l] is an indicator function, 1[W u,k = l] is used to determine the policy application situation of the k-th video in the recall pool, W u,k When = l, it means that when the k-th video in the recall pool is subject to policy l, the value of the indicator function is 1. θ l (u, V u,1 ) represents the effect function of the k-th video in the recall pool under policy l. θ 0 (u, V u,k ) represents the effect function of the k-th video in the recall pool not being subject to policy l.

[0075] Step 2, create a result model;

[0076] In some embodiments, a result model can be created to estimate the results of consumers watching different videos. Assume that each consumer u has a potential result vector y u on the K videos in the recall pool V u ∈R K , where μ(u, v) is used to represent the index result performance of consumer u on video v.

[0077] In some embodiments, a result model can be created to indicate the result of consumer u watching video v. The result model is a vector function for estimating the potential results of consumers on different videos. The result model is represented by the vector function:

[0078] Formula (4)

[0079] H(X; θ, μ) = (h(X; θ 0 , θ 1 , μ), …, h(X; θ 0 , θ L , μ)) T

[0080] where X = (u, V i ), representing the combined vector of consumer u and the recall pool V u h(X; θ 0 , θ l , μ) represents the estimated value of the result model for consumer u on each video in the recall pool V u Specifically, h(X; θ 0 , θ l , μ) can be expressed as:

[0081] Formula (5)

[0082]

[0083] In this formula, μ(u, v u,k ) represents the performance of the metric results of consumer u on the k-th video in the recall pool. p(k|u, V u , W u = l; θ) is a conditional probability, representing the probability that consumer u watches the k-th video under policy l, while p(k|u, V u , W u = 0; θ) represents the probability that consumer u watches the k-th video without being affected by policy l. Among them, u represents the consumer, v u,k represents the k-th video in the recall pool. V u represents the recall pool, and W u represents the policy application situation of each video in the recall pool.

[0084] By calculating the difference between the exposure probabilities of the video under policy l and without policy l, the unfair exposure probability caused by the snatching effect can be eliminated, and an estimate of the causal effect can be obtained. The estimated value of the result model is used to evaluate and compare the impacts of different videos on consumers in the recommendation system, and the recommendation system can be optimized based on the estimated value of the result model.

[0085] Step 3: Obtain the causal inference statistic after eliminating the bias;

[0086] In some embodiments, due to the existence of the poaching effect, the causal inference statistic may have problems of large bias and high variance. The result of causal inference may be inaccurate, and the estimated variance is large, resulting in instability in inference. If we want to improve the accuracy of the causal inference statistic, we need to eliminate the bias problem caused by the poaching effect.

[0087] The causal inference statistic after bias elimination is as follows:

[0088] Formula (6)

[0089] Δ = E X [H(X; θ, μ)]

[0090] where the causal inference statistic after bias elimination is obtained by calculating the expectation of H(X; θ, μ). Δ is the expected value for all samples X, H(X; θ, μ) is a vector function, X = (u, V u ), representing the combined vector of consumer u and the recall pool V u . θ is the model parameter, and μ(u, v) represents the performance of the index of consumer u on video v. Each sample X can be substituted into the H function to obtain the corresponding estimated value, and then these estimated values are averaged to obtain the causal inference statistic after bias elimination.

[0091] Step 4, perform variance inference using a preset method.

[0092] In some embodiments, variance inference can be performed to estimate the variance of the causal inference statistic. Specifically, according to the theory in semi-parametric statistics, the Neyman Orthogonal Score can be used to calculate the variance, and the variance of the Neyman Orthogonal Score is used as the estimate of the variance of the causal effect statistic.

[0093] For the case of one experimental group and one control group:

[0094] When there is only one experimental group and one control group, θ 1 (u, v) represents the utility function affected by the new strategy 1, and θ 0 (u, v) represents the utility function not affected by the new strategy 1 or the utility function affected by strategy 0 (not affected by the new strategy);

[0095] The exposure probability of each video in the K recommended contents in the recall pool predicted by the exposure probability model (part of them belong to the experimental group and part belong to the control group):

[0096]

[0097] Assume that all the recommended contents in the recall pool are affected by the new strategy 1, and the exposure probability of the k-th recommended content in the recall pool:

[0098] p(k|u, V u , W u = l; θ) can be expressed as:

[0099] Assume that all the recommended contents in the recall pool are not affected by the new policy 1. The exposure probability of the k-th recommended content in the recall pool:

[0100] p(k|u, V u , W u = 0; θ) can be expressed as:

[0101] In the case of multiple experimental groups (such as L) and a control group:

[0102] When there are multiple experimental groups, θ l (u, v) represents the utility function affected by the new policy l, and θ 0 (u, v) represents the utility function not affected by the new policy l or the utility function affected by policy 0 (not affected by the new policy);

[0103] Assume that all the recommended contents in the recall pool are affected by the new policy l. The exposure probability of the k-th recommended content in the recall pool:

[0104] For the first video in the recall pool:

[0105] The above formula (2) can be expressed as:

[0106]

[0107] For the k = 2,..., K-th videos:

[0108] The above formula (3) can be expressed as:

[0109]

[0110] Assume that all the recommended contents in the recall pool are not affected by the new policy l. The exposure probability of the k-th recommended content in the recall pool:

[0111] For the k-th video in the recall pool:

[0112] The above formula (2) and formula (3) can be expressed as:

[0113]

[0114] Next, the bias correction inference method provided in the embodiments of the present application will be described.

[0115] Figure 2The flowchart of the deviation correction inference method provided by an exemplary embodiment of the present application is shown. Taking the method being used in a computer device as an example, the method includes at least some of the following steps:

[0116] Step 210: Calculate the first exposure probability and the second exposure probability corresponding to each of the K recommended contents in the recall pool corresponding to the first consumer account;

[0117] The exposure probability refers to the probability that multimedia content (such as advertisements, recommended contents, videos, etc.) is presented to consumers. In a recommendation system, the exposure probability is a measure of the chance that multimedia content is seen by consumers.

[0118] Among them, the first exposure probability is used to indicate the exposure probability of the current recommended content assuming that the recommended contents are all in the experimental group of a randomized controlled trial, and the second exposure probability is used to indicate the exposure probability of the current recommended content assuming that the recommended contents are all in the control group of a randomized controlled trial. When actually conducting a randomized controlled trial, the recall pool contains K recommended contents, K is greater than or equal to 2. In most cases, among the K recommended contents in the recall pool corresponding to the same consumer, some recommended contents belong to the experimental group and some belong to the control group.

[0119] The first exposure probability and the second exposure probability are respectively used to measure the exposure situation of the recommended contents in the recall pool indicated in the experimental group and the control group. By comparing the exposure probabilities of the experimental group and the control group, the impact of the strategy on the exposure of the recommended content can be evaluated.

[0120] In some embodiments, when conducting an AB test on the author side in a two-sided market, some authors will use a new strategy as the experimental group, and some other authors will use the old strategy or not use the new strategy as the control group. Using the new strategy in the experimental group may cause some authors to snatch resources from other authors, such as consumers' attention, viewing time, total consumption, etc.

[0121] In some embodiments, the snatching effect is mainly reflected in the impact on the exposure probability of the recommended contents in the same recall pool. The recall pool for each request of the first consumer account is V u , V u contains K recommended contents. Use W u ∈{0,1} K to represent the situation of each recommended content in the recall pool being affected by the strategy, where 1 is the experimental group, indicating being affected by the new strategy, and 0 is the control group, indicating not being affected by the new strategy. Use e u ∈{0,1} K to represent the exposure situation of the recommended contents in the recall pool. 1 means the recommended content is exposed, and 0 means the recommended content is not exposed.

[0122] In some embodiments, the first exposure probability of the k-th recommended content in the recall pool can be expressed as: p(k|u, V u , W u = l; θ), and the second exposure probability of the k-th recommended content in the recall pool can be expressed as: p(k|u, V u , W u = 0; θ).

[0123] Wherein, u represents the first consumer account, V u represents within the same recall pool, W u represents the policy application status of each recommended content in the recall pool. W u = l indicates that the recommended content in the experimental group is affected by policy l, and W u = 0 indicates that the recommended content in the control group is not affected by policy l. θ represents the parameters of the model. In the expressions of the first exposure probability and the second exposure probability of the k recommended contents in the recall pool, θ can be used to represent the parameters in the model for calculating and estimating the first exposure probability and the second exposure probability. p(k|u, V u , W u = l; θ) represents the probability that the first consumer account selects to view the k-th recommended content in the recall pool under experimental group policy l. p(k|u, V u , W u = 0; θ) represents the probability that the first consumer account selects to view the k-th recommended content in the recall pool under the control group not being affected by policy l or being affected by policy 0.

[0124] Step 220: Obtain the metric result data of the first consumer account for K recommended contents;

[0125] The metric result data refers to the data metrics used to measure and evaluate the performance of the recommendation system. In the recommendation system, the metric result data is a metric used to measure the results generated when the first consumer account views the recommended content. Optionally, the metric result data includes multiple different measurement dimensions, such as whether the first consumer u clicks on the recommended content, the viewing duration of the first consumer account for the recommended content, the consumption amount of the first consumer account in the recommended content, etc. These metric result data can be used to evaluate the interest level, engagement, or purchase intention of the first consumer account towards the recommended content.

[0126] In some embodiments, when the first consumer account views the recommended content in the recall pool, a result is generated, and this result can be represented by the metric result data. Let μ(u, v) represent the metric result data of the first consumer account in the recommended content. Wherein, u represents the first consumer account, and v represents the v-th recommended content in the recall pool.

[0127] Step 230: Based on the first exposure probability and the second exposure probability corresponding to each of the K recommended contents, eliminate the poaching effect bias in the causal effect estimator corresponding to the metric result data, and obtain the causal effect estimator after bias elimination for the first consumer;

[0128] A causal effect estimator is a statistic used to measure the magnitude of a causal effect. The causal effect estimator describes the degree of influence of one causal variable on another causal variable.

[0129] The poaching effect bias refers to the situation where the causal effect estimator is biased due to the interference of other factors. To eliminate the poaching effect bias, it is necessary to compare the exposure probabilities between the experimental group and the control group to determine the impact of the poaching effect.

[0130] Exemplarily, in the scenario of recommended content recommendation, a part of the authors, as the experimental group, will use a new strategy, and another part of the authors, as the control group, will not use the new strategy. The use of the new strategy by the experimental group may cause some authors to poach the resources of other authors, resulting in the emergence of the poaching effect bias.

[0131] In some embodiments, eliminating the poaching effect bias means that in causal inference, by controlling or adjusting relevant variables, the causal effect estimator can more accurately reflect the true causal impact of the first consumer on the recommended content without being interfered by the poaching effect.

[0132] Optionally, evaluate the causal effect of the experimental group using the new strategy compared to the control group not using the new strategy on the duration of consumers watching the recommended content. First, calculate the first exposure probability corresponding to the recommended content in the experimental group and the second exposure probability corresponding to the recommended content in the control group, and use the duration data of the first consumer account watching the recommended content in the recommended content of the experimental group and the control group as the metric result data. Due to the existence of the poaching effect, the difference in exposure probabilities between the experimental group and the control group may lead to bias. To eliminate the poaching effect bias, the exposure probability corresponding to the first consumer account can be used to adjust the causal effect estimator.

[0133] Step 240: Based on the causal effect estimators after bias elimination corresponding to multiple consumer accounts, obtain the causal inference statistic corresponding to the randomized controlled trial.

[0134] A causal inference statistic is a statistic used to evaluate the causal relationship between the treatment group and the control group in a randomized controlled trial. The causal inference statistic is obtained by averaging the causal effect estimators of multiple consumer accounts. The causal effect estimator is a part of the causal inference statistic and is used to calculate and estimate the causal inference statistic. The causal inference statistic includes the expected value and variance of the causal effect estimator after bias elimination.

[0135] In some embodiments, a result model can be created to indicate the results generated when the first consumer account views the recommended content. The result model is a vector function for estimating the results generated by the first consumer account for different recommended contents. The result model is represented by the vector function:

[0136] H(X; θ, μ) = (h(X; θ 0 , θ 1 , μ), …, h(X; θ 0 , θ L , μ)) T

[0137] where h(X; θ 0 , θ l , μ) represents the debiased causal effect estimator corresponding to the first consumer account. In the vector function H, the input is X = (u, V u ), u represents the first consumer account, and V u represents the set of recommended contents viewed by the first consumer account. The output of the vector function H is a vector containing L elements, and each element is calculated through the function h(X; θ 0 , θ l , μ). Among them, X = (u, V u ) represents the combined vector of the first consumer account and the recall pool V u , θ 0 and θ 1 are model parameters, and μ is the parameter of the index result data.

[0138] In some embodiments, based on the debiased causal effect estimators corresponding to multiple consumer accounts, the causal inference statistic corresponding to the randomized controlled trial can be obtained. Optionally, the causal inference statistic can be obtained by calculating the expected value of the vector function H(X; θ, μ), and the causal inference statistic includes at least one of the mean difference and the variance.

[0139] where the expected value of the vector function H(X; θ, μ) can be expressed as:

[0140] Δ = E X [H(X; θ, μ)]

[0141] where H(X; θ, μ) is the vector function, X = (u, v u ), represents the combined vector of the first consumer account and the recall pool, θ is the model parameter, and μ is the parameter of the index result data. Each sample X can be brought into the H function to obtain the corresponding estimated value, and then these estimated values are averaged to obtain the debiased causal inference statistic.

[0142] In summary, for the method provided in this embodiment, by calculating the first exposure probability and the second exposure probability corresponding to the recommended content in the recall pool of the first consumer account respectively; obtaining the index result data of the first consumer account for the recommended content; based on the first exposure probability and the second exposure probability, eliminating the poaching effect bias in the causal effect estimator corresponding to the index result data, to obtain the causal effect estimator after eliminating bias corresponding to the first consumer account; based on the causal effect estimators after eliminating bias corresponding to multiple consumer accounts, obtaining the causal inference statistic corresponding to the randomized controlled trial. In this application, by calculating the exposure probability of the recommended content in the recall pool, combining the difference between the first exposure probability and the second exposure probability with the index result data of the recommended content, eliminating the poaching effect bias in the randomized controlled trial, obtaining the causal effect estimator after eliminating bias, and then performing variance inference on the causal effect estimator after eliminating bias, the causal inference statistic corresponding to the randomized controlled trial can be obtained.

[0143] The following embodiments take the recommended content as videos and the first consumer account as the first consumer as an example for illustration:

[0144] Regarding the exposure probability difference:

[0145] In the optional embodiment based on Figure 2 , Figure 3 FIG. shows the flowchart of the bias correction inference method provided by an exemplary embodiment of the present application. In this embodiment, step 230 is replaced and implemented as steps 231, 232, and 233:

[0146] Step 231: Calculate the difference between the first exposure probability and the second exposure probability corresponding to each of the K recommended contents, to obtain the exposure probability difference corresponding to each of the K recommended contents;

[0147] The exposure probability difference refers to the difference between the first exposure probability of the video in the experimental group and the second exposure probability of the video in the control group between the experimental group and the control group.

[0148] In some embodiments, the same video in the recall pool is compared, and the exposure probability difference is used to evaluate the difference in the impact of the new strategy on the randomized controlled trial. The exposure probability difference can be expressed as:

[0149] p(k|u,V u ,W u =l;θ)-p(k|u,V u ,W u =0;θ)

[0150] where p(k|u,V u ,W u =l;θ) represents the assumption that all K videos are subject to the experimental group strategy l, and the first consumer chooses to watch the recall pool Vu The probability of the k-th video in u , W u = 0; θ) represents the probability that the first consumer chooses to watch the k-th video in the recall pool when assuming that none of the K videos are affected by policy l or are affected by policy 0. Among them, u represents the first consumer, and V u represents the recall pool, and W u represents the policy application situation of each video in the recall pool. W u = l indicates that the video is affected by policy l, and W u = 0 indicates that the video is not affected by policy l, and θ is the model parameter.

[0151] By calculating the difference between the first exposure probability and the second exposure probability, the exposure probability difference corresponding to the k-th video can be obtained, and this exposure probability difference reflects the difference in the exposure probabilities of the k-th video in the recall pool between the experimental group and the control group.

[0152] Step 232: For each of the K recommended contents, calculate the product of the metric result data of each recommended content and the exposure probability difference;

[0153] The metric result data is a metric used to measure the result generated by the first consumer when watching a video. The first consumer will generate a result when watching the videos in the recall pool, and this result can be represented by the metric result data. Let μ(u, v) represent the metric result data of the first consumer in video v. Among them, u represents the first consumer, and v represents the v-th video in the recall pool V u of.

[0154] In some embodiments, for each of the K videos, it is necessary to calculate the product of the metric result data of each video and the exposure probability difference. This product represents the contribution degree of each video to the causal effect.

[0155] The product of the metric result data of the k-th video and the exposure probability difference can be expressed as:

[0156] μ(u, v)(p(k|u, V u , W u = l; θ) - p(k|u, V u , W u = 0; θ))

[0157] Among them, μ(u, v) represents the metric result data of each video, and (p(k|u, V u , W u = l; θ) - p(k|u, V u , W u = 0; θ) represents the exposure probability difference, that is, the difference between the first exposure probability and the second exposure probability.

[0158] Step 233: Sum the products corresponding to the K recommended contents to obtain an unbiased causal effect estimator corresponding to the first consumer account.

[0159] In some embodiments, sum the products corresponding to the K videos in the recall pool to obtain an unbiased causal effect estimator. The causal effect estimator is a statistic used to measure the magnitude of the causal effect.

[0160] The unbiased causal effect estimator corresponding to the first consumer can be expressed as:

[0161]

[0162] where μ(u, v u,k ) represents the metric result data of the first consumer on the K videos viewed in the recall pool, p(k|u, V u , W u = l; θ) represents the first exposure probability, p(k|u, V u , W u = 0; θ) represents the second exposure probability, h(X; θ 0 , θ l , μ) represents the unbiased causal effect estimator. Where u represents the first consumer, V u represents the recall pool, W u represents the policy application status of each video in the recall pool. W u = l indicates that the video is affected by policy l, W u = 0 indicates that the video is not affected by the policy, and θ is the model parameter. For the first consumer, multiply the metric result data corresponding to the K videos viewed in the recall pool by the difference between the first exposure probability and the second exposure probability, and then sum to obtain the unbiased causal effect estimator. By calculating the contribution degree of each of the K videos to the causal effect and calculating the overall causal effect estimator, by calculating the product of the metric result data and the exposure probability difference of each of the K videos and then summing, the products of each of the K videos can be added to obtain a comprehensive causal effect estimator.

[0163] In summary, the method provided in this embodiment first calculates the exposure probability difference. By calculating the difference between the first exposure probability and the second exposure probability corresponding to the k-th video respectively, the exposure probability differences corresponding to the K videos are obtained; for each video, calculate the product of its corresponding metric result data and the exposure probability difference; sum the products corresponding to the K videos to obtain the unbiased causal effect estimator corresponding to the first consumer. In this way, the unbiased causal effect estimator corresponding to the first consumer can be calculated.

[0164] For the first exposure probability:

[0165] Figure 4 The flowchart of the deviation correction inference method provided by an exemplary embodiment of the present application is shown. Taking the method used in a terminal device as an example, the method includes at least some of the following steps:

[0166] Step 310: Obtain the first utility function and the second utility function of each recommended content among the K recommended contents;

[0167] The utility function is a function used to measure the preference degree of the first consumer for the videos in the recall pool. The utility function can evaluate the effect of the videos according to different factors and features. For example, it can evaluate the effect of the videos in the recall pool according to the click-through rate, viewing duration, etc. of the first consumer for the videos.

[0168] Among them, the first utility function is the utility function of the first consumer when each video belongs to the experimental group, and the second utility function is the utility function of the first consumer when each video belongs to the control group. The first utility function and the second utility function are used to measure the preference degree of the first consumer for different videos in the experimental group and the control group.

[0169] In some embodiments, the first utility function is expressed as: θ 1 (u, v), and the second utility function is expressed as: θ 0 (u, v), where the first utility function θ 1 (u, v) represents the utility function affected by the new policy, and the second utility function θ 0 (u, v) represents the utility function not affected by the new policy. Among them, 1 represents the new policy, u represents the first consumer, and v represents the video viewed by the first consumer.

[0170] Step 320: Calculate the first exposure probability corresponding to the v-th recommended content based on the first utility function and the second utility function of the v-th recommended content among the K recommended contents, and the first utility functions and the second utility functions of each recommended content among the K recommended contents.

[0171] Among them, the v-th video is any one of the K videos.

[0172] In some embodiments, for the v-th video, the first utility function and the second utility function of the v-th video are obtained. The first utility function represents the utility function of the first consumer when the v-th video belongs to the experimental group, denoted as θ 1 (u, v), and the second utility function represents the utility function of the first consumer when the v-th video belongs to the control group, denoted as θ 0 (u, v); the first utility functions θ of each video among the K videos are obtained1 (u, v′) and the second utility function θ 0 (u, v′), and each of the K videos is represented as v′. The first exposure probability corresponding to the v-th video can be expressed as: p(v|u, V u , W u = 1). Among them, p(v|u, V u , W u = 1) is a conditional probability, which represents the probability that video v is exposed under the condition of the first consumer, the recall pool V u and the policy situation W u = 1. This probability can be used to measure the exposure degree and competition situation of the video in the recommendation system. u represents the first consumer, V u represents the recall pool, v represents the video in the recall pool, and W u represents the policy situation of each video in the recall pool.

[0173] In summary, the method provided in this embodiment calculates the first exposure probability corresponding to the v-th video by obtaining the first utility function and the second utility function of each video and based on the v-th video among the K videos and the first utility function and the second utility function of each video among the K videos. This calculation method calculates the second exposure probability corresponding to each of the K videos in the recall pool by obtaining the utility function of the video and combining the utility function.

[0174] In an alternative embodiment based on Figure 4 , Figure 5 shows a flowchart of a bias correction inference method provided by an exemplary embodiment of the present application. In this embodiment, step 320 is replaced and implemented as steps 321, 322, and 323:

[0175] Step 321: Calculate a first value based on the first utility function and the second utility function of the v-th recommended content;

[0176] The first value refers to a value calculated according to the given first utility function and the second utility function. Specifically, the first value can be expressed as:

[0177] Among them, θ 0 (u, v) and θ 1 (u, v) are parameters of the model, used to measure the relationship between the first consumer u and the video v. θ 0 (u, v) represents the second utility function when the video is not under policy l, and θ 1 (u, v) represents the first utility function when the video is under policy l.

[0178] On the basis of Figure 5 ,Figure 6 The flowchart of the deviation correction inference method provided by an exemplary embodiment of the present application is shown. Step 321 can be replaced by step 321-1:

[0179] Step 321-1: Calculate the exponential function value with the sum of the first utility function and the second utility function of the v-th recommended content as the power and the natural constant as the base, and use it as the first value.

[0180] In some embodiments, the method for calculating the first value is to use the sum of the first utility function and the second utility function as the power of the exponent and the natural constant e as the base of the exponent, and the obtained value is used as the first value.

[0181] Step 322: Calculate the second value based on the first utility function and the second utility function of each recommended content among the K recommended contents;

[0182] The second value is calculated based on the first utility function and the second utility function of each video among the k videos. Specifically, the second value can be expressed as:

[0183] where v' represents each video among the k videos in the recall pool, and θ 0 (u, v') and θ 1 (u, v') are the parameters of the model, used to measure the relationship between the first consumer and the video v'. θ 0 (u, v') represents the second effect function of the video v' without the strategy, and θ 1 (u, v') represents the first effect function of the video v' under the strategy 1.

[0184] On the basis of Figure 5 On the basis of Figure 7 The flowchart of the deviation correction inference method provided by an exemplary embodiment of the present application is shown. Step 322 can be replaced by step 322-1 and 322-2:

[0185] Step 322-1: Calculate the sum of the first utility function and the second utility function of the v'-th recommended content as the power value corresponding to the v'-th recommended content;

[0186] In some embodiments, the sum of the first utility function θ 1 (u, v') and the second utility function θ 0 (u, v') of the v'-th video is used as the power value corresponding to the v'-th video, that is, calculate the value of θ 1 (u, v') + θ 0 (u, v').

[0187] Step 322-2: Calculate the exponential function value with the sum of the power values corresponding to the K recommended contents as the power and the natural constant as the base, and use it as the second value;

[0188] In some embodiments, for each video in the experimental group and the control group, calculate its corresponding power value, and add up all the power values to obtain the power sum. Then, use the natural constant e as the base and the power sum as the exponent to calculate the second value. The specific calculation method is as follows:

[0189]

[0190] Among them, v′ represents each video in the K videos of the recall pool, and V u represents the recall pool, and θ 0 (u, v′) and θ 1 (u, v′) are the parameters of the model, used to measure the relationship between the first consumer u and the video v′. θ 0 (u, v′) represents the second effect function of the video v′ without the policy l, and θ 1 (u, v′) represents the first effect function of the video v′ under the new policy.

[0191] By summing the power values of each video, using the power sum as the exponent and the natural constant e as the base, the second value is obtained.

[0192] Step 323: Use the ratio of the first value and the second value as the first exposure probability corresponding to the vth recommended content.

[0193] The first exposure probability is used to indicate the exposure probability when the video belongs to the experimental group in the randomized controlled trial.

[0194] In some embodiments, calculate the first value corresponding to the vth video as the numerator part of the ratio. The calculation method of the numerator part is to use the sum of the first utility function θ 1 (u, v) and the second utility function θ 0 (u, v) as the power, and the value of the exponential function with the natural constant e as the base, that is, calculate as the numerator part of the ratio. Calculate the second value corresponding to the v′th video as the denominator part of the ratio. The calculation method of the denominator part is to calculate the power value of each video for all videos v′ viewed by the first consumer u, and then add them up. The calculation method of the power value is to use the sum of the first utility function θ 1 (u, v′) and the second utility function θ 0 (u, v′) as the power, and the value of the exponential function with the natural constant e as the base, that is, calculate:

[0195]

[0196] Next, divide the numerator part by the denominator part to obtain the first exposure probability corresponding to the v-th video. That is, the first exposure probability is:

[0197]

[0198] where p(v|u, V u , W u = 1) is a conditional probability, which represents the probability that video v is exposed under the condition of the first consumer u, the recall pool V u and the policy situation W u = 1. This probability can be used to measure the exposure degree and competition situation of the video in the recommendation system. θ 1 (u, v) represents the first utility function, and θ 0 (u, v) represents the second utility function, and θ 1 (u, v′) represents the first effect function of video v′ under the new policy situation, and θ 0 (u, v′) represents the second effect function of video v′ without the new policy situation.

[0199] In some embodiments, divide the utility function of the v-th video by the utility functions of all videos to obtain the first exposure probability of the v-th video. By this way, the ratio is normalized so that the sum of the probabilities of all videos is 1. This calculation method determines the first exposure probability based on the ratio of the first value and the second value. The larger the ratio, the relatively larger the first value, and the higher the possibility that the video in the experimental group is recommended and exposed to the first consumer.

[0200] In summary, for the method provided in this embodiment, when calculating the first exposure probability corresponding to the v-th video, first calculate the first utility function and the second utility function of the v-th video as the numerator part of the first exposure probability, and then calculate the first utility function and the second utility function of each video among the K videos as the denominator part of the first exposure probability. Calculate the first exposure probability corresponding to the v-th video through this modular calculation method.

[0201] Regarding the second exposure probability:

[0202] Figure 8 The flowchart of the correction inference method provided by an exemplary embodiment of the present application is shown. Taking this method used in a terminal device as an example, the method includes at least some of the following steps:

[0203] Step 410: Obtain the second utility function of each recommended content among the K recommended contents;

[0204] The utility function is a function used to measure the preference degree of the first consumer for the videos in the recall pool. The utility function can evaluate the effect of videos according to different factors and characteristics. For example, it can evaluate the effect of the videos in the recall pool based on the click-through rate, viewing duration, etc. of the first consumer for the videos.

[0205] Among them, the second utility function is the utility function of the first consumer when each video belongs to the control group. The second utility function is used to measure the preference degree of the first consumer for different videos in the experimental group and the control group.

[0206] In some embodiments, the second utility function is expressed as: θ 0 (u, v), where the second utility function θ 0 (u, v) represents the utility function not affected by the new strategy or the utility function affected by strategy 0. Among them, 1 represents being affected by the new strategy, u represents the first consumer, and v represents the video viewed by the first consumer.

[0207] Step 420: Calculate the second exposure probability corresponding to the v-th recommended content based on the second utility function of the v-th recommended content among the K recommended contents and the second utility functions of each recommended content among the V recommended contents.

[0208] Among them, the v-th recommended content is any one of the K recommended contents.

[0209] In some embodiments, for the v-th video, the second utility function of the v-th video is obtained. The second utility function represents the utility function of the first consumer when the v-th video belongs to the control group, denoted as θ 0 (u, v); the second exposure probability corresponding to the v-th video can be expressed as: p(v|u, V u , W u = 0). Among them, p(v|u, V u , W u = 0) is a conditional probability, which represents the probability that video v is exposed under the conditions of the first consumer, the recall pool V u and the policy application situation W u = 0. This probability can be used to measure the exposure degree and competition situation of the video in the recommendation system. u represents the first consumer, V u represents the recall pool, v represents the video in the recall pool, and W u represents the policy application situation of each video in the recall pool.

[0210] In summary, for the method provided in this embodiment, by obtaining the first utility function and the second utility function of each video, based on the second utility function of the v-th video among the k videos and the v-th video among the K videos, as well as the first utility function and the second utility function of each video among the K videos, the first exposure probability corresponding to the v-th video is calculated. This calculation method obtains the utility function of the video and calculates the second exposure probability corresponding to each of the K videos in the recall pool in combination with the utility function.

[0211] In an alternative embodiment based on Figure 8 the following Figure 9 shows a flowchart of a bias correction inference method provided by an exemplary embodiment of the present application. In this embodiment, step 420 is replaced and implemented as steps 421, 422, and 423:

[0212] Step 421: Calculate a third value based on the second utility function of the v-th recommended content;

[0213] The third value is a value calculated according to the given second utility function. Specifically, the third value can be expressed as:

[0214] where θ 0 (u, v) are the parameters of the model, used to measure the relationship between the first consumer u and the video v. θ 0 (u, v) represents the second function when the video is not subject to policy l or is subject to policy 0.

[0215] In an alternative embodiment based on Figure 9 the following Figure 10 shows a flowchart of a bias correction inference method provided by an exemplary embodiment of the present application. Step 421 can be replaced by step 421-1:

[0216] Step 421-1: Calculate the value of the exponential function with the sum of the second utility functions of the v-th recommended content as the exponent and the natural constant as the base, as the third value;

[0217] In some embodiments, the method for calculating the first value is to calculate the sum of the second utility functions of the v-th video as the exponent of the power, and use the natural constant e as the base of the exponent, and the obtained value as the third value.

[0218] Step 422: Calculate a fourth value based on the second utility functions of each of the K recommended contents;

[0219] The fourth value is calculated based on the first utility function and the second utility function of each of the k videos. Specifically, the fourth value can be expressed as:

[0220] Among them, v′ represents each of the k videos in the recall pool, and θ 0 (u, v′) are the parameters of the model, used to measure the relationship between the first consumer u and the video v′. θ 0 (u, v′) represents the second effect function of the video v′ without the new policy.

[0221] Based on Figure 9 on the basis of Figure 11 FIG. shows a flowchart of a deviation correction inference method provided by an exemplary embodiment of the present application. Step 422 can be replaced by step 422-1 and step 422-2:

[0222] Step 422-1: Calculate the second utility function of the v′-th recommended content as the power value corresponding to the v′-th recommended content;

[0223] In some embodiments, the second utility function θ 0 (u, v′) corresponding to the v′-th video is used as the power value corresponding to the v′-th video, that is, calculate the value of θ 0 (u, v′).

[0224] Step 422-2: Calculate the exponential function value with the sum of the power values corresponding to the K recommended contents as the power and the natural constant as the base as the fourth value;

[0225] In some embodiments, for each video in the experimental group and the control group, calculate its corresponding power value, and add up all the power values to obtain the power sum. Then, using the natural constant e as the base and the power sum as the exponent, calculate the fourth value. The specific calculation method is as follows:

[0226]

[0227] Among them, v′ represents each of the K videos in the recall pool, and V u represents the recall pool, and θ 0 (u, v′) are the parameters of the model, used to measure the relationship between the first consumer u and the video v′. θ 0 (u, v′) represents the second effect function of the video v′ without the new policy.

[0228] By summing the power values of each video, using the power sum as the exponent and the natural constant e as the base, the fourth value is obtained.

[0229] Step 423: Use the ratio of the third value to the fourth value as the second exposure probability corresponding to the v-th recommended content.

[0230] The second exposure probability is used to indicate the exposure probability when the video belongs to the control group in the randomized controlled trial.

[0231] In some embodiments, the third value corresponding to the v-th video is calculated as the numerator part of the ratio. The calculation method of the numerator part is to use the second utility function θ 0 (u, v) as the power, and the value of the exponential function with the natural constant e as the base, that is, calculate as the numerator part of the ratio. The fourth value corresponding to the v'-th video is calculated as the denominator part of the ratio. The calculation method of the denominator part is that for all videos v' viewed by the first consumer u, calculate the power value of each video, and then add them up. The calculation method of the power value is to use the sum of the second utility functions θ 0 (u, v') as the power, and the value of the exponential function with the natural constant e as the base, that is, calculate:

[0232]

[0233] Then divide the numerator part by the denominator part to obtain the second exposure probability corresponding to the v-th video. That is, the second exposure probability is:

[0234]

[0235] where p(v|u, V u , W u = 0) is a conditional probability, which represents the probability that the video v is exposed under the conditions of the first consumer u, the recall pool V u and the policy situation W u = 0. This probability can be used to measure the exposure degree and competition situation of the video in the recommendation system. θ 0 (u, v) represents the second utility function, and θ 0 (u, v') represents the second effect function of the video v' without the new policy.

[0236] In some embodiments, the utility function of the v-th video is divided by the utility functions of all videos to obtain the second exposure probability of the v-th video. By this way, the ratio is normalized so that the sum of the probabilities of all videos is 1. This calculation method determines the second exposure probability based on the ratio of the third value and the fourth value. The larger the ratio, the relatively larger the third value, and the higher the possibility that the video in the control group is recommended and exposed to the first consumer.

[0237] In summary, for the method provided in this embodiment, when calculating the first exposure probability corresponding to the v-th video, first calculate the second utility function of the v-th video as the numerator part of the second exposure probability, and then calculate the first utility function and the second utility function of each video among the k videos as the denominator part of the second exposure probability. The second exposure probability corresponding to the v-th video is calculated by this modular calculation method.

[0238] For the causal inference statistic "mean difference":

[0239] Figure 12 The flowchart of the bias correction inference method provided by an exemplary embodiment of the present application is shown. Taking the method used in a terminal device as an example, the method includes at least some of the following steps:

[0240] Step 510: Calculate the mean difference or expected value of the causal effect estimators after removing bias corresponding to multiple consumer accounts, and obtain the causal inference statistic corresponding to the randomized controlled trial.

[0241] The mean difference refers to the difference between the averages of the experimental group and the control group. Exemplarily, the mean difference can be used to measure the impact of a strategy on the causal effect.

[0242] Exemplarily, assume studying the impact of a new strategy on video recommendations in a recommendation system. The randomized controlled trial is divided into an experimental group and a control group. The experimental group receives the new strategy, and the control group does not receive the new strategy. Then the video exposure situations of the two groups are respectively recorded. X represents the video exposure situation of the experimental group, and Y represents the video exposure situation of the control group. If you want to know whether the new strategy has an impact on video recommendations, you can calculate the mean μ_X of the video exposure situation in the experimental group and calculate the mean μ_Y of the video exposure situation in the control group. Then, by calculating the mean difference μ X -μ_Y, the difference in video exposure situations between the experimental group and the control group can be obtained.

[0243] The expected value refers to the average value or expected result of a random variable. The expected value is the result of a weighted average of the values of a random variable, where the weight of each value is the probability of its occurrence. It can be considered that the expected value is the average result of the random variable in a large number of repeated experiments. Exemplarily, the expected value can be used to represent the average result of the causal effect estimators of multiple consumers.

[0244] The causal effect estimator is a statistic used to measure the magnitude of the causal effect. The causal effect estimator describes the degree of influence of one causal variable on another causal variable. The causal inference statistic is a statistic used to evaluate the causal relationship between the treatment group and the control group in a randomized controlled trial. The causal inference statistic is obtained by averaging the causal effect estimators of multiple consumers. The causal effect estimator is a part of the causal inference statistic and is used to calculate and estimate the causal inference statistic.

[0245] In some embodiments, due to the existence of the poaching effect, the causal inference statistic may have problems such as large bias and high variance. The result of causal inference may be inaccurate. If you want to improve the accuracy of the causal inference statistic, it is necessary to eliminate the bias problem caused by the poaching effect.

[0246] Optionally, by calculating the expected value of the debiased causal effect estimators of multiple consumers, a causal inference statistic can be obtained for causal inference in a randomized controlled trial. By taking the expected value of the causal effect estimators of multiple consumers, the influence of bias can be reduced, and more accurate causal inference results can be obtained. At the same time, an overall causal effect estimator can be obtained, which can be used to infer the overall causal inference statistic.

[0247] The causal inference statistic corresponding to the randomized controlled trial is as follows:

[0248] Δ = E X [H(X; θ, μ)]

[0249] where the causal inference statistic corresponding to the randomized controlled trial is obtained by calculating the expectation of H(X; θ, μ). H(X; θ, μ) is a vector function, X = (u, V u ), representing the combined vector of consumer u and the recall pool V u . θ is the model parameter, and μ(u, v) represents the index result data of the first consumer u on video v. Each sample X can be substituted into the H function to obtain the corresponding causal effect estimator, and then these causal effect estimators are processed by calculating the expectation to obtain the causal inference statistic corresponding to the randomized controlled trial.

[0250] In summary, the method provided in this embodiment obtains the causal inference statistic corresponding to the randomized controlled trial by calculating the mean difference or expected value of the debiased causal effect estimators corresponding to multiple consumers. Summarizing the causal effect estimators of multiple consumers to obtain a causal inference statistic to represent the causal inference result of the randomized controlled trial helps to more comprehensively evaluate the effect of the randomized controlled trial.

[0251] Regarding the "variance" of the causal inference statistic:

[0252] Figure 13 The flowchart of the bias-corrected inference method provided by an exemplary embodiment of the present application is shown. Taking the method used in a terminal device as an example, the method includes at least some of the following steps:

[0253] Step 610: Calculate the Neyman orthogonal score corresponding to the randomized controlled trial based on the debiased causal effect estimators corresponding to multiple consumer accounts;

[0254] In some embodiments, if the results of the experimental group and the control group vary greatly, the statistical variance will be high, which means the reliability of causal inference is low. On the contrary, if the results of the experimental group and the control group vary little, the statistical variance will be low, and the reliability of causal inference is high. The statistical variance in a randomized controlled trial is an indicator to measure the uncertainty of causal inference, and the statistical variance can help understand the stability and reliability degree of experimental results.

[0255] The Neyman orthogonal score is a method for eliminating confounding factors. When conducting causal inference, the Neyman orthogonal score is used to reduce the bias in the causal inference statistic. By performing Neyman orthogonal analysis on the data of the experimental group and the control group, the Neyman orthogonal score corresponding to the randomized controlled trial is obtained.

[0256] In some embodiments, the Neyman orthogonal score constructs a multiple linear regression model. By calculating the variance of the Neyman orthogonal score, the variance of the Neyman orthogonal score can be used as an estimate of the causal inference statistic.

[0257] The Neyman orthogonal score can be expressed as: Ψ(X, W, Y; θ, μ)

[0258] where X = (u, V u ) represents the combined vector of consumer u and the recall pool V u , W represents the condition whether the video is subject to the policy, Y is the potential outcome of the consumer on the video, θ is the model parameter, and μ is the metric result performance of the consumer on the video.

[0259] Ψ(X, W, Y; θ, μ) = H(X; θ, μ) - the first subtraction part × the second subtraction part × the third subtraction part, that is

[0260] where

[0261] the above first subtraction part is the first-order derivative operation of the h(X; θ 0 , θ 1 , μ) function; the second subtraction part is the inverse operation of the Hessian matrix, that is, the second-order derivative operation of the function, and the third subtraction part is the first-order derivative operation of the loss function.

[0262] · The first subtraction part

[0263] For k = 2, …, K:

[0264]

[0265] where the left side of the above equal sign is a function of θ 0 (u, v u,k) partial derivative expression, representing the partial derivative of the function h(X; θ 0 , θ 1 , μ) with respect to θ 0 (u, v u,k ) The partial derivative of p(k|u, V u , W u ≡l) is a conditional probability, representing the exposure probability of k given u, V u and W u ≡l. p(k|u, V u , W u ≡0) is a conditional probability, representing the exposure probability of k given u, V u and W u ≡0. μ(u, v u,k ) is an index result data, and E[y|u, V u , W u ≡l] represents the expected value of y given v, V u and W u ≡l. E[y|u, V u , W u ≡0] respectively represents the expected value of y under the given conditions u, V u and W u ≡0. This expression represents that under the given conditions of u, V u , W u , by comparing the probability distributions of k in the video with and without the policy l, and the expected values of y in the video with and without the policy l, the prediction performance of the model in the experimental group and the control group is measured.

[0266] For k = 1,..., K:

[0267]

[0268] Among them, the left side of the above equal sign is a partial derivative expression with respect to θ l (u, v u,k ) representing the partial derivative of the function h(X; θ 0 , θ 1 , μ) with respect to θ l (u, v u,k ) p(k|u, V u , W u ≡l) is a conditional probability, representing the exposure probability of k given u, V u and W u ≡l, and μ(u, v u,k ) is an index result data, and E[y|u, V u , W u≡l] represents the expected value of y given u, V u and W u ≡l.

[0269]

[0270] Where the left side of the above equal sign is a partial derivative expression with respect to μ, representing the partial derivative of the function h(X; θ 0 , θ 1 , μ) with respect to μ. p(k|u, V u , W u ≡l) is a conditional probability, representing the exposure probability of k given u, V u and W u ≡l, and p(k|u, V u , W u ≡0) is a conditional probability, representing the exposure probability of k given u, V u and W u ≡0.

[0271] ·The second part of the subtraction

[0272]

[0273] Where the left side of the above equal sign is a partial derivative expression with respect to θ 0 , representing the square of the partial derivative of the function with respect to . p(k|u, V u , W u ; θ) is a conditional probability, representing the exposure probability of k given u, V u and W u under the θ model parameters. When the exposure probability of k is p(k|u, V u , W u ; θ), and the non-exposure probability of k is 1 - p(k|u, V u , W u ; θ). Among them, k ∈ {2:K} represents the k = 2 to Kth videos in the recall pool.

[0274]

[0275] Where the left side of the above equal sign is a partial derivative expression with respect to , representing the square of the partial derivative of the function with respect to . w k =l is an indicator function, which takes the value of 1 when w k =l, and 0 otherwise. The indicator function plays a selection role. p(k|u, V u , Wu ; θ) is a conditional probability, representing the exposure probability of k under the condition of given u, V u and W u under the θ model parameters, and 1 - p(k|u, V u , W u ; θ) represents the non - exposure probability of k under the condition of given u, V u and W u under the θ model parameters. Here, k ∈ {1:K} represents the k - th video from k = 1 to K in the recall pool.

[0276]

[0277] Among them, the left - hand side of the above - mentioned equal sign is the product of the partial derivatives of θ 0 (u, v u,k ) and θ l (u, v u,k ), representing the partial derivatives of the function with respect to θ l (u, v u,k ) and θ 0 (u, v u,k ). w k = l is an indicator function, which takes the value of 1 when w k = l and 0 otherwise. The indicator function plays a selection role. p(k|u, V u , W u ; θ) is a conditional probability, representing the exposure probability of k under the condition of given u, V u and W u under the θ model parameters, and 1 - p(k|u, V u , W u ; θ) represents the non - exposure probability of k under the condition of given u, V u and W u under the θ model parameters. Here, k ∈ {2:K} represents the k - th video from k = 2 to K in the recall pool.

[0278]

[0279] Among them, the left - hand side of the above - mentioned equal sign is the second - order partial derivative of the mixed partial derivatives of θ 0 (u, v u,k1 ) and θ l (u, v u,k2 ), representing the 1 partial derivatives of the e l (X, W, Y; θ) function with respect to θ u,k2 (u, v 0 (u, v u,k1 ) and θ 1 |u, Vu ,W u ; θ) is a conditional probability, representing the exposure probability of k given u, W under the θ model parameters u and W u under the condition of 1 k, p(k 2 |u, V u ,W u ; θ) is a conditional probability, representing the exposure probability of k given u, V under the θ model parameters u and W u under the condition of 2 k. Among them, k 1 ≠k 2 means that k 1 and k 2 are different videos in the recall pool, and both are the k = 2 to K-th videos.

[0280]

[0281] Among them, the left side of the above equation is the second-order derivative of the mixed partial derivative of θ 0 (u, v u,k1 ) and θ l (u, v u,k2 ), representing the partial derivatives of the e 1 (X, W, Y; θ) function with respect to θ l (u, v u,k2 ) and θ 0 (u, v u,k1 ). w k2 =l is an indicator function, which takes the value of 1 when w k =l, and 0 otherwise. The indicator function plays a selection role. p(k 1 |u, V u ,W u ; θ) is a conditional probability, representing the exposure probability of k given u, V under the θ model parameters u and W u under the condition of 1 k, p(k 2 |u, V u ,W u ; θ) is a conditional probability, representing the exposure probability of k given u, V under the θ model parameters u and W u under the condition of 2 k. Among them, k 1 ≠k 2 means that k 1 and k 2 are different videos in the recall pool, and k 1 is the k = 2 to K-th video.

[0282]

[0283] Among them, on the left side of the above equal sign is a partial derivative with respect to θ l2 (u, v u,k2 ) and θ 11 (u, v u,k1 ) for the second-order derivative of the mixed partial derivative, representing the partial derivative of the e 1 (X, W, Y; θ) function with respect to θ l2 (u, v u,k2 ) and θ l1 (u, v u,k1 ). w k1 = l1 and w k2 = l2 are indicator functions. l1 and l2 represent different strategies respectively. When w k1 = l1, w k2 = l2, the value is 1, otherwise it is 0. The indicator function plays a selection role. p(k 1 |u, V u , W u ; θ) is a conditional probability, representing the exposure probability of k u under the condition of given u, V u and W 1 under the θ model parameters. p(k 2 |u, V u , W u ; θ) is a conditional probability, representing the exposure probability of k u under the condition of given u, V u and w 2 under the θ model parameters. Among them, k 1 ≠ k 2 means that k 1 and k 2 are different videos in the recall pool, and both k 1 and k 2 are the k = 1 to K-th videos.

[0284]

[0285] Among them, on the left side of the above equal sign is a partial derivative expression with respect to , representing the square of the partial derivative of the function with respect to . Among them, k u = k is an indicator function. When k u = k, the value is 1, otherwise it is 0. The indicator function plays a selection role. Among them, k ∈ {1:K} represents for the k = 1 to K-th videos in the recall pool.

[0286]

[0287] Among them, the left side of the above equal sign is a second-order derivative for the mixed partial derivative of μ(u, v u,k2 ) and μ(u, v u,k1 ), which represents is the partial derivative of μ(u, v u,k2 ) and μ(u, v u,k1 ). Among them, k 1 ≠k 2 ∈{1:K} means that k 1 and k 2 are different videos in the recall pool, and both k 1 and k 2 are the k = 1 to K-th videos.

[0288] ·For the third part from the bottom

[0289] For k = 2, …, K:

[0290]

[0291] Among them, the left side of the above equal sign is a partial derivative expression with respect to θ 0 (u, v u,k ), which represents the partial derivative of the e 1 (X, W, Y; θ) function with respect to θ 0 (u, v u,k ). p(k|u, V u , W u ; θ) is a conditional probability, which represents the exposure probability of k under the condition of given u, V u and W u . k u =k is an indicator function, which takes the value of 1 when k u =k, and 0 otherwise.

[0292] For k = 1, …, K:

[0293]

[0294] Among them, the left side of the above equal sign is a partial derivative expression with respect to θ l (u, v u,k ), which represents the partial derivative of the e 1 (X, W, Y; θ) function with respect to θ l (u, v u,k ). W k =l is an indicator function, which takes the value of 1 when W k =l, and 0 otherwise. p(k|u, V u , W u; θ) is a conditional probability, representing the exposure probability of k given u, V u and W u under the condition of θ model parameters.

[0295]

[0296] Among them, the left side of the above equal sign is a partial derivative expression with respect to θ l (u, v u,k ), representing the partial derivative of the e 2 (X, W, Y; μ) function with respect to θ l (u, v u,k ). k u = k is an indicator function, which takes the value of 1 when k u = k, and 0 otherwise. μ(u, v u,k ) is an index result data, and y is the expected value.

[0297] Step 620: Calculate the difference between the Neyman orthogonal score corresponding to the randomized controlled trial and the causal inference statistic corresponding to the randomized controlled trial, as the debiased Neyman orthogonal score;

[0298] In some embodiments, calculate the difference between the Neyman orthogonal score corresponding to the randomized controlled trial and the causal inference statistic to obtain the debiased Neyman orthogonal score. This difference can be used to measure the causal effect after eliminating confounding factors.

[0299] The debiased Neyman orthogonal score is:

[0300] Ψ(X, W, Y; θ, μ) - E X [H(X; θ, μ)]

[0301] Among them, Ψ(X, W, y; θ, μ) represents the Neyman orthogonal score corresponding to the randomized controlled trial, and E X [H(X; θ, μ)] represents the causal inference statistic corresponding to the randomized controlled trial, X = (u, V u ) represents the combined vector of consumer u and recall pool V u , W represents the condition of whether the video is subject to the policy, y is the potential result of the consumer on the video, θ is the model parameter, and μ is the index result performance of the consumer on the video.

[0302] Among them, Ψ(X, W, Y; θ, μ) = H(X; θ, μ) - the first subtraction part × the second subtraction part × the third subtraction part, as detailed in step 610 and will not be elaborated here.

[0303] Step 630: Calculate the variance of the debiased Neyman orthogonal score as the statistical variance corresponding to the randomized controlled trial.

[0304] In some embodiments, the variance of the debiased Neyman orthogonal score is calculated as the statistical variance corresponding to the randomized controlled trial. This statistical variance can be used to evaluate the dispersion or uncertainty of the debiased Neyman orthogonal score, and further evaluate the stability and reliability of causal inference.

[0305] In summary, the method provided in this embodiment can construct a multiple linear regression model by calculating the Neyman orthogonal score of a randomized controlled trial; by calculating the difference between the Neyman orthogonal score of the randomized controlled trial and the causal inference statistic, the debiased Neyman orthogonal score can be obtained, and the debiased Neyman orthogonal score reflects the change of the Neyman orthogonal score of the randomized controlled trial after eliminating the bias; by calculating the variance of the debiased Neyman orthogonal score, the statistical variance of the randomized controlled trial can be obtained, and this variance can be used to evaluate the stability and reliability of the results of the randomized controlled trial.

[0306] Beneficial effects of this solution:

[0307] This solution will have smaller bias and variance compared to the mean difference between the experimental group and the control group in the A / B test as the causal effect estimate in terms of bias and variance dimensions, so there will be more accurate causal effect inference.

[0308] In the experiment on WeChat, a simulated experimental scenario was constructed. Table 1 shows the advantages of the method of this embodiment of the present application compared to the traditional method in terms of the bias, variance, and bias^2 + variance of the statistic:

[0309] Table 1

[0310]

[0311] Combined with reference Figure 14 , Figure 14 in which the shaded area a represents the distribution of the statistic of this embodiment of the present application, and the shaded area b represents the distribution of the statistic of the A / B test in the traditional method, Figure 14 and c in

[0312] Figure 15 shows the structural block diagram of a bias correction inference device provided by an embodiment of the present application. This device has the functions of implementing the above-mentioned bias correction inference method example, and the functions can be implemented by hardware or by hardware executing corresponding software. This device can be the server introduced above, or can be set in the server. As Figure 15 shown, the device 1100 may include: a calculation module 1110, an acquisition module 1120, an elimination module 1130, and a processing module 1140;

[0313] A calculation module 1110, configured to calculate a first exposure probability and a second exposure probability respectively corresponding to K recommended contents in a recall pool corresponding to a first consumer account, where the first exposure probability is used to indicate an exposure probability of a current recommended content when assuming that the K recommended contents are all recommended contents in an experimental group of a randomized controlled trial, and the second exposure probability is used to indicate an exposure probability of the current recommended content when assuming that the K recommended contents are all recommended contents in a control group of the randomized controlled trial;

[0314] An acquisition module 1120, configured to acquire index result data of the K recommended contents for the first consumer account, where the first consumer account is any one of multiple consumer accounts;

[0315] An elimination module 1130, configured to eliminate a poaching effect bias in a causal effect estimator corresponding to the index result data based on the first exposure probability and the second exposure probability respectively corresponding to the K recommended contents, to obtain an elimination bias causal effect estimator corresponding to the first consumer account;

[0316] A processing module 1140, configured to obtain a causal inference statistic corresponding to the randomized controlled trial based on the elimination bias causal effect estimators corresponding to the multiple consumer accounts.

[0317] In some optional embodiments, the elimination module 1130 further includes a calculation sub-module and a summation sub-module.

[0318] In an optional embodiment, the calculation sub-module is configured to calculate a difference between the first exposure probability and the second exposure probability respectively corresponding to the K recommended contents, to obtain an exposure probability difference respectively corresponding to the K recommended contents; the calculation sub-module is configured to, for each of the K recommended contents, calculate a product of the index result data of each recommended content and the exposure probability difference; the summation sub-module is configured to sum the products corresponding to the K recommended contents, to obtain an elimination bias causal effect estimator corresponding to the first consumer account.

[0319] In some optional embodiments, the calculation module 1110 further includes an acquisition sub-module and a calculation sub-module.

[0320] In an optional embodiment, an acquisition sub-module is configured to acquire a first utility function and a second utility function for each of the K recommended contents, where the first utility function is the utility function of the first consumer account when each of the recommended contents belongs to the experimental group, and the second utility function is the utility function of the first consumer account when each of the recommended contents belongs to the control group; a calculation sub-module is configured to calculate a first exposure probability corresponding to the v-th recommended content based on the first utility function and the second utility function of the v-th recommended content among the K recommended contents, and the first utility functions and the second utility functions of the respective recommended contents among the K recommended contents.

[0321] Wherein, the v-th recommended content is any one of the K recommended contents.

[0322] In some optional embodiments, the calculation sub-module further includes a calculation unit.

[0323] In an optional embodiment, the calculation unit is configured to calculate a first value based on the first utility function and the second utility function of the v-th recommended content; the calculation unit calculates a second value based on the first utility functions and the second utility functions of the respective recommended contents among the K recommended contents; the calculation unit uses the ratio of the first value and the second value as the first exposure probability corresponding to the v-th recommended content.

[0324] In some optional embodiments, the calculation unit further includes a calculation sub-unit.

[0325] In an optional embodiment, the calculation sub-unit is configured to calculate the value of an exponential function with the sum of the first utility function and the second utility function of the v-th recommended content as the power and the natural constant as the base as the first value.

[0326] In some optional embodiments, the calculation unit further includes a calculation sub-unit.

[0327] In an optional embodiment, the calculation sub-unit is configured to calculate the sum of the first utility function and the second utility function of the v'-th recommended content as the power value corresponding to the v'-th recommended content; the calculation sub-unit is configured to calculate the value of an exponential function with the sum of the power values corresponding to the K recommended contents as the power and the natural constant as the base as the second value.

[0328] In some optional embodiments, the calculation module 1110 further includes an acquisition sub-module and a calculation sub-module.

[0329] In an optional embodiment, an acquisition sub-module is configured to acquire a second utility function of each of the K recommended contents, where the second utility function is the utility function of the first consumer account when each of the recommended contents belongs to a control group; a calculation sub-module is configured to calculate a second exposure probability corresponding to the v-th recommended content based on the second utility function of the v-th recommended content among the K recommended contents and the second utility functions of the respective recommended contents among the K recommended contents.

[0330] Wherein, the v-th recommended content is any one of the recommended contents.

[0331] In some optional embodiments, the calculation sub-module further includes a calculation unit.

[0332] In an optional embodiment, the calculation unit is configured to calculate a third value based on the second utility function of the v-th recommended content; the calculation unit is configured to calculate a fourth value based on the second utility functions of the respective recommended contents among the K recommended contents; the calculation unit is configured to use the ratio of the third value to the fourth value as the second exposure probability corresponding to the v-th recommended content.

[0333] In some optional embodiments, the calculation unit further includes a calculation subunit.

[0334] In an optional embodiment, the calculation subunit is configured to calculate an exponential function value with the sum of the second utility functions of the v-th recommended content as the power and the natural constant as the base as the third value.

[0335] In some optional embodiments, the calculation unit further includes a calculation subunit.

[0336] In an optional embodiment, the calculation subunit is configured to calculate the second utility function of the v'-th recommended content as the power value corresponding to the v'-th recommended content; the calculation subunit is configured to calculate an exponential function value with the sum of the power values corresponding to the K recommended contents as the power and the natural constant as the base as the fourth value.

[0337] In some optional embodiments, the processing module 1140 further includes a calculation sub-module.

[0338] In an optional embodiment, the calculation sub-module is configured to calculate the mean difference or expected value of the debiased causal effect estimators corresponding to the multiple consumer accounts.

[0339] In some optional embodiments, the processing module 1140 further includes a calculation sub-module.

[0340] In an optional embodiment, a calculation sub-module is configured to calculate a Neyman orthogonal score corresponding to the randomized controlled trial based on the debiased causal effect estimators corresponding to the multiple consumer accounts; a calculation sub-module is configured to calculate a difference between the Neyman orthogonal score corresponding to the randomized controlled trial and the causal inference statistic corresponding to the randomized controlled trial as a debiased Neyman orthogonal score; a calculation sub-module is configured to calculate a variance of the debiased Neyman orthogonal score as the statistical variance corresponding to the randomized controlled trial.

[0341] In summary, the device provided in this embodiment calculates the first exposure probability and the second exposure probability corresponding to the recommended content in the recall pool corresponding to the first consumer account; obtains the indicator result data of the first consumer account for the recommended content; based on the first exposure probability and the second exposure probability, eliminates the poaching effect bias in the causal effect estimator corresponding to the indicator result data, and obtains a debiased causal effect estimator corresponding to the first consumer account; based on the debiased causal effect estimators corresponding to the multiple consumer accounts, obtains the causal inference statistic corresponding to the randomized controlled trial. This application calculates the exposure probability of the recommended content in the recall pool, combines the difference between the first exposure probability and the second exposure probability with the indicator result data of the recommended content, eliminates the poaching effect bias in the randomized controlled trial, obtains a debiased causal effect estimator, and then performs variance inference on the debiased causal effect estimator to obtain the causal inference statistic corresponding to the randomized controlled trial.

[0342] Figure 16 FIG. shows a structural block diagram of a computer device 1500 shown in an exemplary embodiment of the present application. The computer device can be used to implement the bias correction inference method provided in the above embodiment. The computer device 1500 includes a central processing unit (CPU) 1501, a system memory 1504 including a random access memory (RAM) 1502 and a read-only memory (ROM) 1503, and a system bus 1505 connecting the system memory 1504 and the central processing unit 1501. The computer device 1500 further includes a basic input / output system (Input / Output system, I / O system) 1506 for facilitating information transmission between various components within the computer device, and a mass storage device 1507 for storing an operating system 1513, application programs 1514, and other program modules 1515.

[0343] The basic input / output system 1506 includes a display 1508 for displaying information and input devices 1509 such as a mouse and a keyboard for user input of information. Both the display 1508 and the input devices 1509 are connected to the central processing unit 1501 through an input / output controller 1510 connected to the system bus 1505. The basic input / output system 1506 may further include an input / output controller 1510 for receiving and processing inputs from a plurality of other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1510 also provides outputs to a display screen, a printer, or other types of output devices.

[0344] The mass storage device 1507 is connected to the central processing unit 1501 through a mass storage controller (not shown) connected to the system bus 1505. The mass storage device 1507 and its associated computer-readable storage medium provide non-volatile storage for the terminal device 1500. That is, the mass storage device 1507 may include a computer-readable storage medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.

[0345] Without loss of generality, the computer-readable storage medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable storage instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only registers (EPROM), electrically-erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage media is not limited to the above several types. The above-mentioned system memory 1504 and mass storage device 1507 may be collectively referred to as memory.

[0346] The memory stores one or more programs, the one or more programs are configured to be executed by one or more central processing units 1501, the one or more programs include instructions for implementing the above method embodiments, and the central processing unit 1501 executes the one or more programs to implement the methods provided by the above various method embodiments.

[0347] According to various embodiments of the present application, the computer device 1500 can also operate by connecting to a remote terminal device on the network through a network such as the Internet. That is, the computer device 1500 can be connected to the network 1512 through the network interface unit 1511 connected to the system bus 1505. Or rather, the network interface unit 1511 can also be used to connect to other types of networks or remote terminal device systems (not shown).

[0348] The memory further includes one or more programs. The one or more programs are stored in the memory, and the one or more programs include steps for performing the methods provided in the embodiments of the present application that are executed by the terminal device.

[0349] Embodiments of the present application further provide a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the deviation correction and inference method provided in each of the above method embodiments.

[0350] Embodiments of the present application further provide a computer program product. The computer program product includes a computer program. The computer program is stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device from the computer-readable storage medium, so that the computer device executes to implement the deviation correction and inference method provided in each of the above method embodiments.

[0351] It can be understood that in the specific implementation of the present application, for data, historical data, portraits, and other user data processing related to user identity or characteristics, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0352] It should be noted that unless otherwise clearly defined herein, all terms used in the claims are interpreted according to their ordinary meanings in the technical field. Unless otherwise clearly stated, all references to "an element, device, component, equipment, step, etc." will be interpreted openly as referring to at least one instance of the element, device, component, equipment, step, etc. Unless clearly stated, the steps of any method disclosed herein are not necessarily executed in the exact order disclosed.

[0353] It should be understood that the "plurality" mentioned herein refers to two or more. "And / or", which describes the associated relationship of associated objects, indicates that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

Claims

1. A deviation correction inference method, characterized in that: The method comprises: Calculate a first exposure probability and a second exposure probability respectively corresponding to K recommended contents in the recall pool corresponding to the first consumer account, wherein the first exposure probability is used to indicate the exposure probability of the current recommended content when assuming that the K recommended contents are all recommended contents in the experimental group of the randomized controlled experiment, and the second exposure probability is used to indicate the exposure probability of the current recommended content when assuming that the K recommended contents are all recommended contents in the control group of the randomized controlled experiment; Obtaining indicator result data of the first consumer account for the K recommended contents, where the first consumer account is any one of the multiple consumer accounts; Eliminating the snatch effect bias in the causal effect estimate corresponding to the indicator result data based on the first exposure probability and the second exposure probability respectively corresponding to the K recommended contents, and obtaining the causal effect estimate corresponding to the first consumer account after eliminating the bias; Based on the bias-eliminated causal effect estimates corresponding to the multiple consumer accounts, a causal inference statistic corresponding to the randomized controlled trial is obtained.

2. The method according to claim 1, characterized in that The method of eliminating the snatch effect bias in the causal effect estimator corresponding to the indicator result data based on the first exposure probability and the second exposure probability respectively corresponding to the K recommended contents, and obtaining the causal effect estimator corresponding to the first consumer account after eliminating the bias, includes: Calculating the difference between the first exposure probability and the second exposure probability respectively corresponding to the K recommended contents, to obtain the exposure probability differences respectively corresponding to the K recommended contents; For each of the K recommended contents, calculating the product of the indicator result data of each recommended content and the exposure probability difference; The products corresponding to the K recommended contents are summed to obtain a bias-eliminated causal effect estimate corresponding to the first consumer account.

3. The method according to claim 2, characterized in that The calculating the first exposure probabilities respectively corresponding to the K recommended contents in the recall pool corresponding to the first consumer account includes: Obtaining a first utility function and a second utility function of each of the K recommended contents, wherein the first utility function is a utility function of the first consumer account when each of the recommended contents belongs to the experimental group, and the second utility function is a utility function of the first consumer account when each of the recommended contents belongs to the control group; Calculating a first exposure probability corresponding to the vth recommended content based on the first utility function and the second utility function of the vth recommended content among the K recommended contents, and the first utility function and the second utility function of each recommended content among the K recommended contents; The vth recommended content is any one of the K recommended contents.

4. The method according to claim 3, characterized in that The calculating, based on the first utility function and the second utility function of the vth recommended content among the K recommended contents, and the first utility function and the second utility function of each recommended content among the K recommended contents, a first exposure probability corresponding to the vth recommended content includes: Calculating a first value based on the first utility function and the second utility function of the vth recommended content; Calculating a second value based on the first utility function and the second utility function of each of the K recommended contents; The ratio of the first value to the second value is used as the first exposure probability corresponding to the vth recommended content.

5. The method according to claim 4, characterized in that The calculating the first value based on the first utility function and the second utility function of the vth recommended content includes: An exponential function value with the sum of the first utility function and the second utility function of the v-th recommended content as the power and a natural constant as the base is calculated as the first value.

6. The method according to claim 4, characterized in that The calculating the second value based on the utility function actually corresponding to each of the K recommended contents includes: Calculating the sum of the first utility function and the second utility function of the v'-th recommended content as the power value corresponding to the v'-th recommended content; An exponential function value with the sum of the power values ​​corresponding to the K recommended contents as the power and a natural constant as the base is calculated as the second value.

7. The method according to claim 2, characterized in that The calculating the second exposure probabilities respectively corresponding to the recommended contents in the recall pool corresponding to the first consumer account includes: Obtaining a second utility function of each of the K recommended contents, where the second utility function is a utility function of the first consumer account when each of the recommended contents belongs to a control group; Calculating a second exposure probability corresponding to the vth recommended content based on the second utility function of the vth recommended content among the K recommended contents and the second utility function of each recommended content among the K recommended contents; The vth recommended content is any one of the recommended contents.

8. The method according to claim 7, characterized in that The calculating, based on the second utility function of the vth recommended content among the K recommended contents and the second utility function of each recommended content among the K recommended contents, a first exposure probability corresponding to the vth recommended content includes: Calculating a third value based on the second utility function of the vth recommended content; Calculating a fourth value based on a second utility function of each of the K recommended contents; A ratio of the third value to the fourth value is used as the second exposure probability corresponding to the vth recommended content.

9. The method according to claim 8, characterized in that The calculating a third value based on the second utility function of the vth recommended content includes: An exponential function value with the sum of the second utility functions of the v-th recommended content as the power and a natural constant as the base is calculated as the third value.

10. The method according to claim 8, characterized in that The calculating the fourth value based on the second utility function of each recommended content in the recommended content includes: Calculating a second utility function of the v'th recommended content as a power value corresponding to the v'th recommended content; An exponential function value with the sum of the power values ​​corresponding to the K recommended contents as the power and a natural constant as the base is calculated as the fourth value.

11. The method according to any one of claims 1 to 10, characterized in that: The causal effect estimates after eliminating bias corresponding to the multiple consumer accounts are used to obtain the causal inference statistics corresponding to the randomized controlled trial, including: Calculate the mean difference or expected value of the bias-eliminated causal effect estimates corresponding to the multiple consumers.

12. The method according to any one of claims 1 to 10, characterized in that: The causal effect estimates after eliminating bias corresponding to the multiple consumer accounts are used to obtain the causal inference statistics corresponding to the randomized controlled trial, including: Calculate the Neyman orthogonal score corresponding to the randomized controlled trial based on the bias-eliminated causal effect estimates corresponding to the multiple consumers; Calculate the difference between the Neyman orthogonal score corresponding to the randomized controlled trial and the causal inference statistic corresponding to the randomized controlled trial as the debiased Neyman orthogonal score; The variance of the debiased Neyman quadrature score is calculated as the statistical variance corresponding to the randomized controlled trial.

13. A deviation correction inference device, characterized in that: The device comprises: a calculation module, configured to calculate a first exposure probability and a second exposure probability respectively corresponding to K recommended contents in a recall pool corresponding to the first consumer account, wherein the first exposure probability is used to indicate an exposure probability of the current recommended content when assuming that the K recommended contents are all recommended contents in an experimental group of a randomized controlled experiment, and the second exposure probability is used to indicate an exposure probability of the current recommended content when assuming that the K recommended contents are all recommended contents in a control group of the randomized controlled experiment; an acquisition module, configured to acquire the indicator result data of the first consumer account for the K recommended contents, wherein the first consumer account is any one of the plurality of consumer accounts; An elimination module, configured to eliminate the snatch effect bias in the causal effect estimate corresponding to the indicator result data based on the first exposure probability and the second exposure probability respectively corresponding to the K recommended contents, and obtain the causal effect estimate corresponding to the first consumer account after the bias is eliminated; A processing module is used to obtain the causal inference statistic corresponding to the randomized controlled trial based on the bias-eliminated causal effect estimates corresponding to the multiple consumers.

14. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the correction inference method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores executable instructions, and the executable instructions are loaded and executed by a processor to implement the correction inference method as described in any one of claims 1 to 12.