Strategy evaluation method and device, electronic equipment and storage medium

By constructing an objective function and solving for the objective weights, and using the weighted index data of the control experimental group, the problem of low policy evaluation accuracy was solved, and higher evaluation accuracy was achieved.

CN121390291APending Publication Date: 2026-01-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511522743.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in strategy evaluation, making it difficult to accurately assess the effectiveness of strategies.

Method used

By constructing an objective function, using the weighted results and variable intercepts of the first run-through index data from multiple control experimental groups, and combining them with preset constraints, the objective weights are solved, and then the control index data are integrated to evaluate the strategy's effectiveness.

Benefits of technology

This improves the accuracy of strategy evaluation and ensures the accuracy and reliability of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121390291A_ABST
    Figure CN121390291A_ABST
Patent Text Reader

Abstract

The invention discloses a strategy evaluation method and device, electronic equipment and a storage medium. The method comprises the following steps: determining a target function of a target experimental group based on first idle running index data of at least three experimental groups in a reference time period; under a preset constraint condition, solving the target function to obtain a target weight corresponding to the target experiment group; the target weights comprise respective weights of the plurality of control experiment groups; the constraint conditions comprise that the weight of each control experiment group is a real number and the intercept is not zero; according to the target weight, fusing the second idle running index data of the plurality of control experiment groups in the strategy time period to obtain the control index data of the target experiment group in the strategy time period; and evaluating the target strategy based on the contrast index data to obtain a strategy evaluation result of the target strategy. According to the method provided by the invention, the evaluation accuracy of target strategy evaluation is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electronic information, and more particularly, to a policy evaluation method and device, an electronic device, and a storage medium. BACKGROUND

[0002] The synthetic control method is a kind of causal inference method. The synthetic control method refers to constructing a "synthetic control group" highly similar to a target experimental group affected by a policy by weighting and combining a plurality of experimental groups not affected by the policy. The index data of the control group under the condition of not being affected by the policy is the weighted result of the index data of the plurality of experimental groups not affected by the policy. Then, the effect of the policy is evaluated based on the index data of the target experimental group under the condition of being affected by the policy and the index data of the synthetic control group under the condition of not being affected by the policy.

[0003] However, when the method is used to evaluate the policy, the evaluation accuracy is low. SUMMARY

[0004] Therefore, the embodiments of the present application provide a policy evaluation method and device, an electronic device, and a storage medium.

[0005] In a first aspect, the embodiments of the present application provide a policy evaluation method, which comprises: determining a target function of a target experimental group based on first empty running index data of at least three experimental groups in a reference period; no target policy is applied to each experimental group in the reference period; the at least three experimental groups include the target experimental group; the target function includes a weighted result of first empty running index data of a plurality of control experimental groups and an intercept for adjusting the weighted result; the plurality of control experimental groups are experimental groups other than the target experimental group in the at least three experimental groups; the solution target of the target function is to minimize the target function by adjusting the weights of the plurality of control experimental groups and the intercept; under a preset constraint condition, the target function is solved to obtain a target weight corresponding to the target experimental group; the target weight includes the weights of the plurality of control experimental groups; the constraint condition includes that the weights of the control experimental groups are real numbers and the intercept is not zero; the second empty running index data of the plurality of control experimental groups in a policy period is fused by using the target weight to obtain control index data of the target experimental group in the policy period; no target policy is applied to each control experimental group in the policy period, and the policy period is different from the reference period.

[0006] Secondly, embodiments of this application provide a strategy evaluation device, comprising: a determination module, configured to determine an objective function for a target experimental group based on first no-load index data of at least three experimental groups within a reference time period; no target strategy was applied to any of the experimental groups during the reference time period, and the at least three experimental groups include the target experimental group; the objective function includes a weighted result of the first no-load index data of multiple control experimental groups and an intercept adjusted for the weighted result; the multiple control experimental groups are the experimental groups other than the target experimental group among the at least three experimental groups; the objective of solving the objective function is to minimize the objective function by adjusting the weights and intercepts of the multiple control experimental groups; The solution module is used to solve the objective function under preset constraints to obtain the target weights corresponding to the target experimental group. The target weights include the weights of each of the multiple control experimental groups. The constraints include that the weights of each control experimental group are real numbers and the intercept is not zero. The fusion module is used to fuse the second run index data of multiple control experimental groups during the policy period using the target weights to obtain the control index data of the target experimental group during the policy period. The target policy was not applied to each control experimental group during the policy period, and the policy period is different from the reference period. The evaluation module is used to evaluate the target policy based on the control index data to obtain the policy evaluation result of the target policy.

[0007] Optionally, the reference time period includes reference time periods under different time period categories; the determination module is further configured to determine the objective function corresponding to the target experimental group under the time period category for each time period category, based on the first no-load indicator data of at least three experimental groups within the reference time period of the time period category; the solution module is further configured to solve the objective function corresponding to the target experimental group under the time period category under constraints, and obtain the target weight corresponding to the target experimental group under the time period category; the fusion module is further configured to obtain the time period category to which the strategy time period belongs as the target time period category; and based on the target weight corresponding to the target experimental group under the target time period category, fuse the second no-load indicator data of multiple control experimental groups to obtain the control indicator data of the target experimental group within the strategy time period.

[0008] Optionally, the determination module is also used to obtain a reference regularization term; the reference regularization term is associated with the weights of multiple control experimental groups; and the objective function is determined based on the first run index data of at least three experimental groups and the reference regularization term.

[0009] Optionally, the reference regularization term is multiple; the objective function includes an objective function of the target experiment group for the multiple reference regularization terms; the solving module is further configured to solve, for each reference regularization term, the objective function of the target experiment group for the reference regularization term under the constraint condition, to obtain a candidate weight of the target experiment group corresponding to the reference regularization term; the candidate weight includes a weight of each of the multiple control experiment groups; the first empty running index data of the multiple control experiment groups are fused by referring to the candidate weight of the target experiment group corresponding to the reference regularization term, to obtain candidate index data of the target experiment group corresponding to the regularization term; and the target weight of the target experiment group is determined from the candidate weight of the target experiment group corresponding to the multiple reference regularization terms, based on the candidate index data of the target experiment group corresponding to the multiple reference regularization terms.

[0010] Optionally, the solving module is further configured to determine an error index of the target experiment group corresponding to the reference regularization term based on a difference between the candidate index data of the target experiment group corresponding to the reference regularization term and the first empty running index data; determine the target regularization term from the multiple reference regularization terms based on the error index of the target experiment group corresponding to the multiple reference regularization terms; and obtain the candidate weight of the target experiment group corresponding to the target regularization term as the target weight of the target experiment group.

[0011] Optionally, the solving module is further configured to determine the reference regularization term based on at least one of a square sum of the weights of the multiple control experiment groups and an absolute value sum of the weights of the multiple control experiment groups.

[0012] Optionally, the evaluation module is further configured to obtain multiple improvement amplitudes and strategy index data of the target experiment group in a strategy period; the target experiment group is subjected to a target strategy in the strategy period; adjust the first empty running index data of the target experiment group by each improvement amplitude respectively, to obtain adjusted index data of the target experiment group under each improvement amplitude; for each improvement amplitude, determine an evaluation result of a test efficiency of a reference test algorithm under the improvement amplitude based on the adjusted index data of the at least three experiment groups and the first empty running index data of the at least three experiment groups by the reference test algorithm; the reference test algorithm is any one of multiple candidate test algorithms; determine a target test algorithm from the multiple candidate test algorithms based on the test efficiency evaluation result of each of the multiple candidate test algorithms under the multiple improvement amplitudes; and determine a strategy evaluation result of the target strategy based on the control index data and the strategy index data of the target experiment group in the strategy period by the target test algorithm.

[0013] Optionally, the evaluation module is further configured to, for each of the boosting amplitudes, determine, by referring to the reference test algorithm, a probability that the reference test algorithm detects a type I error at the boosting amplitude based on a difference between the adjustment indicator data of the at least three experimental groups at the boosting amplitude and the first empty run indicator data of the at least three experimental groups, as a type I error rate of the reference test algorithm at the boosting amplitude; and determine, based on the type I error rate of the reference test algorithm at the boosting amplitude, the test efficacy evaluation result of the reference test algorithm at the boosting amplitude.

[0014] Optionally, the adjustment indicator data comprises single-dimension adjustment indicator data of each of a plurality of dimensions, and the first empty run indicator data comprises single-dimension empty run indicator data of each of the plurality of dimensions; the evaluation module is further configured to, for each of the boosting amplitudes, determine, by referring to the reference test algorithm, a P value between the single-dimension adjustment indicator data of a reference dimension of a reference experimental group and the single-dimension empty run indicator data of the reference dimension of the reference experimental group at the boosting amplitude, as a P value of the reference experimental group at the reference dimension for the reference test algorithm at the boosting amplitude, wherein the reference experimental group is any one of the at least three experimental groups, and the reference dimension is any one of the plurality of dimensions; and determine, based on P values of the at least three experimental groups at the plurality of dimensions for the reference test algorithm at the boosting amplitude, the type I error rate of the reference test algorithm at the boosting amplitude.

[0015] Optionally, the first empty run indicator data comprises single-dimension empty run indicator data of each of a plurality of dimensions; the evaluation module is further configured to determine, based on the type I error rate of the reference test algorithm at the boosting amplitude, a dimension number of the plurality of dimensions, and an experimental group number of the at least three experimental groups, a type I error controllable test result of the reference test algorithm at the boosting amplitude; the type I error controllable test result of the reference test algorithm is used to indicate whether a type I error detected by the reference test algorithm is controllable; and determine, based on the type I error rate of the reference test algorithm at the boosting amplitude and the type I error controllable test result, the test efficacy evaluation result of the reference test algorithm at the boosting amplitude.

[0016] Optionally, the evaluation module is further configured to determine, based on a preset significant threshold, a dimension number of the plurality of dimensions, and an experimental group number of the at least three experimental groups, a controllable interval; and determine, based on a comparison result of the type I error rate of the reference test algorithm at the boosting amplitude and the controllable interval, the type I error controllable test result of the reference test algorithm at the boosting amplitude.

[0017] Optionally, the evaluation module is further configured to, for each of the candidate test algorithms, determine, based on a number of key boosting amplitudes corresponding to the candidate test algorithm, a first test algorithm from the plurality of candidate test algorithms, wherein the key boosting amplitude corresponding to the candidate test algorithm refers to a boosting amplitude in the plurality of boosting amplitudes at which the candidate test algorithm detects a type I error that is controllable; and determine, if there are a plurality of first test algorithms, a target test algorithm from the first test algorithms based on sizes of type I error rates of the plurality of first test algorithms at the plurality of boosting amplitudes.

[0018] Optionally, the evaluation module is also used to acquire multiple second-test algorithms; different second-test algorithms are obtained by configuring the initial test algorithm with different hyperparameters; for each improvement, the power evaluation result of the reference second-test algorithm under the improvement is determined by referring to the adjusted index data of multiple experimental groups and the control index data of multiple experimental groups based on the improvement of the reference second-test algorithm; the reference second-test algorithm is any one of the multiple second-test algorithms; based on the power evaluation results of the multiple second-test algorithms under multiple improvement, a third-test algorithm is determined from the multiple second-test algorithms; if there are multiple third-test algorithms, one is determined from the multiple third-test algorithms as a candidate test algorithm based on the size of the hyperparameters configured for the multiple third-test algorithms.

[0019] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory; the memory stores computer-readable instructions, which, when executed by the processor, implement the above-described method.

[0020] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-readable instructions that, when executed by a processor, implement the above-described method.

[0021] Fifthly, embodiments of this application provide a computer program product including computer-readable instructions that, when executed by a processor, implement the method described above.

[0022] This application provides a strategy evaluation method, apparatus, electronic device, and storage medium. First, based on the first run-through index data of at least three experimental groups within a reference time period, an objective function for the target experimental group is determined. The objective function incorporates a variable intercept used to adjust the weighted result of the first run-through index data of multiple control experimental groups. Simultaneously, the constraints of the objective function include real weights for each control experimental group and a non-zero intercept. The constraints no longer limit the weights of each control experimental group to be greater than zero, the sum of the weights of multiple control experimental groups to 1, or the intercept to zero. This reduces the restrictions on the intercept and the weights of each control experimental group. Therefore, when solving the objective function, the differences between the target experimental group and the control experimental groups can be better captured, making the weights determined for each control experimental group more accurate. This results in more accurate control index data for the target experimental group within the strategy time period based on the target weights, leading to a higher accuracy rate of the strategy evaluation result obtained based on the control index data, thus improving the evaluation accuracy of the target strategy. Attached Figure Description

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description only constitute some embodiments of the present application. For those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0024] Figure 1 A schematic diagram suitable for the application scenario to which the embodiments of the present application are applied is shown. Figure 2 A flowchart of a strategy evaluation method according to one embodiment of the present application is shown. Figure 3 A flowchart of a strategy evaluation method according to another embodiment of the present application is shown. Figure 4 A flowchart of a strategy evaluation method according to another embodiment of the present application is shown. Figure 2 A flowchart of step S140 in the corresponding embodiment is shown. Figure 5 A schematic diagram of a target weight determination process in the embodiments of the present application is shown. Figure 6 A schematic diagram of a target verification algorithm determination process in the embodiments of the present application is shown. Figure 7 A block diagram of a strategy evaluation device according to one embodiment of the present application is shown. Figure 8 A structural block diagram of an electronic device for executing a strategy evaluation method according to the embodiments of the present application is shown. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some embodiments of the present application, rather than all the embodiments. According to the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0026] In the following description, the terms "first\second" are only to distinguish similar objects, and do not represent a specific order of the objects. Understandably, "first\second" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application. It is to be understood that the use of "a", "an", "the" and "at least one" herein are not intended to exclude the presence of zero or more than one of the referenced elements in the context. Further, the use of "and / or" is intended to represent an association between the items named. For example, A and / or B means that there can be A alone, B alone, or A and B together.

[0028] Abbreviations and key terms used in this application are defined as follows: Synthetic Control (SC): A causal inference method that constructs a "synthetic control" that is highly similar to the target unit affected by the policy, by combining multiple units that are not affected by the policy. In the context of internet city experiments, it is often used to simulate the performance of the experimental city in the absence of the policy, by combining other cities that do not implement the policy.

[0029] Online Experiment: A controlled experiment conducted in the real online environment of an internet product, by applying different product strategies to different user groups to evaluate the effect of the strategy. Compared with offline experiments, online experiments can obtain real user behavior feedback.

[0030] City-level Experiment: A large-scale online experiment that uses cities as the granularity, with different cities as experimental groups, and the entire product strategy applied in a specific city, while other cities serve as controls. This experimental method is often used for strategies that need to be implemented by region due to regulatory requirements. For example, a new commercial pricing strategy for an application in a certain province, a new dispatch algorithm for a food delivery application in a certain city, and a new subsidy strategy for a shopping software in a certain city.

[0031] Unit: In the field of experiments, a unit is a group of objects participating in the experiment. In the context of city experiments, a unit refers to a province or a city.

[0032] Treatment Unit: A unit that receives the treatment of the policy. In the context of city experiments, a treatment unit refers to a province or a city that receives the treatment of the policy.

[0033] Control Unit: A unit that does not receive the treatment of the policy, used to construct the control of the treatment unit. In the context of city experiments, it refers to other provinces or cities that do not implement the policy.

[0034] RMSE (Root Mean Square Error): a measure of the difference between predicted and actual values, with smaller values indicating better fit. In synthetic control, used to assess the quality of the pre-period fit.

[0035] BIAS: the bias of the fit, the difference between the expected value of the estimator and the true parameter value. An unbiased estimator is one whose expected value equals the true value. In synthetic control, the bias reflects the systematic prediction error of the model.

[0036] Type I Error: the probability of falsely rejecting the null hypothesis when it is true. In experiments, this manifests as a strategy that is actually ineffective but is statistically tested as effective. Industry standards typically control this to be less than 5%.

[0037] Statistical Power: the probability that a test algorithm correctly rejects the null hypothesis (or called the zero hypothesis) when it is false. The higher the statistical power of a test algorithm, the easier it is for the test algorithm to detect a real strategy effect.

[0038] Robustness of the model: the ability of a model to maintain its performance in the face of data perturbations, outliers, and distribution changes. A robust model will not produce large prediction errors due to small data changes.

[0039] Sample: In statistics, a sample is a subset of individuals or observations taken from a larger set, known as the population. Samples are used for statistical analysis and inference to estimate or predict the characteristics of the population without surveying the entire population. Typically, a sample refers to an object participating in the experiment in the unit (i.e., the experimental group).

[0040] P-value: a measure of the probability of observing the difference between the experimental and control groups under the assumption that the null hypothesis is true. If the P-value is very low (less than a pre-specified significance threshold), it means that the difference between the experimental and control groups is small, and the null hypothesis can be rejected.

[0041] Please refer to Figure 1 , which shows a schematic diagram applicable to the application scenario to which the embodiments of the present application are applicable. The application scenario includes a terminal 110 and a server 120.

[0042] The terminal 110 is a device with a policy evaluation request transceiving function, such as a smartphone, a tablet computer, an e-book reader, a music player, a wearable device, a smart home device, a vehicle-mounted terminal, and the like. The terminal 110 is installed with a client capable of sending a policy evaluation request. For example, the client can be an instant messaging client, a content interaction client, a mailbox client, a short message client, and the like.

[0043] The server 120 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDNs (Content Delivery Networks), and big data and artificial intelligence platforms.

[0044] In some embodiments, the terminal 110 can send a policy evaluation request to the server 120, and the server 120 determines a target function of a target experiment group based on the first empty running index data of the at least three experiment groups in the reference period in response to the policy evaluation request, then solves the target function under a preset constraint condition to obtain a target weight corresponding to the target experiment group, and fuses the second empty running index data of the multiple control experiment groups in the policy period through the target weight to obtain control index data of the target experiment group in the policy period, finally, the server 120 evaluates the target policy based on the control index data to obtain a policy evaluation result of the target policy, and returns the policy evaluation result to the terminal 110, so that the user can view the policy evaluation result through the terminal 110.

[0045] In yet some embodiments, the terminal 110 directly obtains the first empty running index data of the at least three experiment groups in the reference period and the second empty running index data of the multiple control experiment groups in the policy period from the server 120 in response to a local policy evaluation request, and then evaluates the target policy based on the obtained data to obtain a policy evaluation result of the target policy.

[0046] Of course, the server 120 can also directly determine the policy evaluation result in response to a local policy evaluation request.

[0047] For ease of understanding, the following explains the electronic device as the execution subject of the policy evaluation method of the present application.

[0048] Please refer to Figure 2 , Figure 2 Fig. 1 shows a flowchart of a policy evaluation method according to an embodiment of the present application, which is applied to an electronic device, which can be a terminal 110 or a server 120. Figure 1The method can comprise: In S110, based on the first empty running index data of the at least three experimental groups in the reference period, a target function of the target experimental group is determined.

[0049] In the reference period, no target strategy is applied to each experimental group. The at least three experimental groups include the target experimental group. The target function includes a weighted result of the first empty running index data of a plurality of control experimental groups and an intercept for adjusting the weighted result. The plurality of control experimental groups are experimental groups other than the target experimental group in the at least three experimental groups. The solution target of the target function is to minimize the target function by adjusting the weight of the plurality of control experimental groups and the intercept.

[0050] In this embodiment, the experimental group refers to an experimental group participating in a target test experiment (i.e., the unit described above). The target test experiment refers to an experiment for evaluating a target strategy. The target test experiment is used to detect the effect of the target strategy on the test subject, which can be a positive effect or a negative effect. The test subject can refer to a biological element in the biological field, a chemical element in the chemical field, a physical element in the physical field, and an application program in the electronic information field, etc. The biological element can be, for example, a virus, a cell, a bacterium, and an animal body, etc. The chemical element can be, for example, a compound, a single element, and a mixture, etc. The physical element can be, for example, a real physical object or a virtual virtual object. The application program can be, for example, a standalone application program, a script, a web program, and a small program, etc.

[0051] The target strategy refers to a strategy participating in the evaluation. The target strategy can be to add a new method or means, change an existing method or means, or delete an existing method or means. For example, the test subject is a certain virus, and the target strategy can be to add a virus inhibitor, replace a virus catalyst, and configure a specific virus survival environment. For another example, the test subject is a certain compound, and the target strategy can be to add a catalyst or delete an inhibitor. For another example, the test subject is a shopping application program, and the target strategy is to add a new commodity recommendation algorithm or delete an existing commodity recommendation algorithm in the shopping application program, etc.

[0052] The sample refers to the test subject or the sample acted on by the test subject. For biological elements, chemical elements, and physical elements, etc., the sample can refer to the test subject itself. For application programs, etc., the sample can also be a user, a device, or other application programs acted on by the test subject. For example, the sample can be a user faced by a conventional application program. The sample can also be a device faced by a test-type application program for testing the device. The sample can also be an application program faced by a test-type application program for testing the application program.

[0053] In the present application, the experimental group refers to an experimental group participating in a target test experiment, and each experimental group includes multiple samples. The reference period is a period set based on demand, and no target strategy is applied to each experimental group in this period, and the index data of each sample in the experimental group in the empty running state (state without applying the target strategy) is observed as the empty running index data. The first empty running index data of the experimental group is obtained by summarizing the empty running index data of each sample in the experimental group.

[0054] The preset index quantity refers to the quantity or parameter affected by the target strategy. For example, the experimental subject is a biological element, the sample is the biological element itself, the target strategy affects the survival time of the biological element, and the target strategy affects the survival time of the biological element. The index quantity can be the survival period and survival half period of the sample, and the like. For example, the experimental subject is a chemical element, the sample is the chemical element itself, the target strategy affects the time of the chemical reaction, and the target strategy affects the time of the chemical reaction. The index quantity can be the total time of the chemical reaction, and the like. For example, the experimental subject is an application program, the sample is a user using the application program, and the target strategy affects the time or frequency of the user using the application program. The index quantity can be the use time or use frequency of the user using the application program, and the like.

[0055] The index data is the specific value of the preset index quantity. For example, the experimental subject is a biological element, the sample is the biological element itself, the index quantity is the survival time, and the index data is the specific value of the survival time. For example, the experimental subject is a virus, the sample is an animal affected by the virus, the index quantity is the time of the virus disappearing in the animal body, and the index data is the specific value of the time of the virus disappearing in the animal body. For example, the experimental subject is a conference application program, the sample is a user, and the index quantity is the use time per day. The index data refers to the time of the user using the conference application program for remote conference per day. For example, the experimental subject is a shopping application program, the sample is a user, and the index quantity is the use frequency per day. The index data refers to the number of times of the user using the shopping application program per day.

[0056] In the present application, different experimental groups can refer to experimental groups located in different regions and different environments. For example, the experimental subject is a biological element, the sample is the biological element itself, and different experimental groups can refer to biological elements in different cities. For example, the sample is an animal affected by a virus, and different experimental groups can refer to the same kind of animal in different countries. For example, the experimental subject is a conference application program, and different experimental groups can refer to the same kind of animal in different countries. For example, the experimental subject is a shopping application program, the sample is a user, the index quantity is the use frequency per day, and the index data refers to the number of times of the user using the shopping application program per day.

[0057] Firstly, a reference period is configured based on the demand, no target strategy is applied to each experimental group in the reference period, and the index quantity of each sample in each experimental group is monitored to obtain the index data of each sample for the index quantity, and then the first empty running index data of the experimental group in the reference period is obtained.

[0058] After obtaining the first empty running index data of each experimental group in the reference period, a target function of a target experimental group is constructed based on the first empty running index data of each experimental group in the reference period. The target experimental group can be any one of the experimental groups.

[0059] Generally, the observation term of the target experimental group can be determined based on the first empty running index data of the target experimental group, the control term of the target experimental group can be determined based on the weighted results of the first empty running index data of the plurality of control experimental groups and the intercept for adjusting the weighted results, and finally a function indicating the difference (here, the difference can be a difference value, a square of the difference value, or an absolute value of the difference value, etc.) between the control term and the observation term of the target experimental group is constructed as the target function. The adjustable parameters in the target function are the weights of the plurality of control experimental groups and the intercept, and the solving target of the target function is to minimize the target function.

[0060] In some embodiments, the first empty running index data of the target experimental group can be obtained as the observation term of the target experimental group, or the first empty running index data of the target experimental group can be multiplied by a coefficient to obtain a product result as the observation term of the target experimental group, or the first empty running index data of the target experimental group can be operated by an operation function to obtain an operation result as the observation term of the target experimental group, wherein the operation function can be a trigonometric function, an exponential function, and a logarithmic function, etc.

[0061] In some embodiments, the weighted results of the first empty running index data of the plurality of control experimental groups and the intercept can be summed or subtracted to obtain a result as the control term of the target experimental group, or the weighted results of the first empty running index data of the plurality of control experimental groups can be operated by an operation function first, and the operation result and the intercept can be summed or subtracted to obtain a result as the control term of the target experimental group.

[0062] It is not difficult to understand that the reference period can include a plurality of time points, so that the first index data includes the index data of each time point, and thus for each time point, the observation term and the control term of the target experimental group are determined, and then the differences between the observation terms and the control terms of the target experimental group at the plurality of time points are summed to obtain the target function.

[0063] For example, the difference between the observation item and the control item of the target experiment group is the absolute value of the difference, and the adjustment mode of the weighted result of the first empty running index data of the plurality of control experiment groups by the intercept is to sum the intercept and the weighted result of the first empty running index data of the plurality of control experiment groups, and the objective function can be as formula one, formula two: (I) wherein, is the intercept, is the weight of the plurality of control experiment groups, is the reference period, is the index data of the first empty running index data of the target experiment group at time point t, is a set composed of the plurality of control experiment groups, is the index data of the first empty running index data of the i th control experiment group in the plurality of control experiment groups at time point t, is the index data of the first empty running index data of the i th control experiment group in the plurality of control experiment groups at time point t, is the weight of the i th control experiment group in the plurality of control experiment groups.

[0064] It is not difficult to understand that the objective function shown in the aforementioned formula one is only an example, and other types of objective functions that meet the requirements of the objective function of the present application can be constructed based on demand, for example, in the objective function, the adjustment mode of the weighted result of the first empty running index data of the plurality of control experiment groups by the intercept is to subtract the intercept and the weighted result of the first empty running index data of the plurality of control experiment groups.

[0065] In some embodiments, S110 can further include: SA1, obtaining a reference regularization item; the reference regularization item is associated with the weight of the plurality of control experiment groups; SA2, determining the objective function based on the first empty running index data of the at least three experiment groups and the reference regularization item.

[0066] In other words, the weight of the plurality of control experiment groups can also be used for regularization processing to obtain a reference regularization item, and then the objective function is determined based on the first empty running index data of the at least three experiment groups and the reference regularization item. The reference regularization item here can be a regularization item obtained by L1 regularization processing or L2 regularization processing based on the weight of the plurality of control experiment groups, or a regularization item obtained by preset budgeting based on the weight of the plurality of control experiment groups. The preset operation can be an operation method or means configured based on demand.

[0067] In the present application, the observation item and the control item of the target experiment group can be constructed in the aforementioned manner, then the difference between the observation item and the control item of the target experiment group is determined, and the result obtained by summing or subtracting the difference and the reference regularization item is used as the objective function.​

[0068] Optionally, the aforementioned process of determining the reference regularization term can comprise: determining the reference regularization term based on at least one of a sum of squares of the weights of the plurality of control experimental groups and an absolute sum of the weights of the plurality of control experimental groups.

[0069] The absolute sum of the weights of the plurality of control experimental groups is an L1 regularization term obtained by L1 regularization processing, and the sum of squares of the weights of the plurality of control experimental groups is an L2 regularization term obtained by L2 regularization processing.

[0070] In this application, the reference regularization term can be determined based on at least one of the L1 regularization term and the L2 regularization term, for example, multiplying the L1 regularization term by a specified coefficient λ to obtain the reference regularization term, for another example, multiplying the L2 regularization term by a specified coefficient λ to obtain the reference regularization term, for another example, performing weighted summation on the L1 regularization term and the L2 regularization term, and multiplying the summation result by a specified coefficient to obtain the reference regularization term. The specified coefficient λ here is the regularization parameter of the reference regularization term, and its specific value can be determined by 5-fold cross-validation method.

[0071] For example, the reference regularization term can include and and so on, wherein the weight set based on the demand, ∈[0,1]. Correspondingly, if the reference regularization term is , the difference between the observation term and the control term of the target experimental group and the operation of the reference regularization term obtain the target function in the form of summation, the target function can be as formula two, formula two is as follows: (Two) Similarly, if the reference regularization term is , the difference between the observation term and the control term of the target experimental group and the operation of the reference regularization term obtain the target function in the form of summation, the target function can be as formula three, formula three is as follows: (Three) S120, under the preset constraint condition, solving the target function to obtain the target weight corresponding to the target experimental group.

[0072] Wherein, the target weight includes the weight of each control experimental group; the constraint condition includes that the weight of each control experimental group is a real number and the intercept is not zero.

[0073] Constrained by the constraint condition, the target function is solved to obtain the weight of each control experimental group, and the weight of each control experimental group is used as the target weight of the target experimental group.

[0074] Optionally, as aforementioned, the reference regularization terms are added to the objective function, and there are multiple reference regularization terms, thus the objective function includes objective functions of the target experiment group for the multiple reference regularization terms, for example, the reference regularization terms are and , and the objective function is shown in Equation 2 and Equation 3; correspondingly, S120 includes: SB1, for each reference regularization term, solving the objective function of the target experiment group for the reference regularization term under the constraint condition to obtain the candidate weight of the target experiment group corresponding to the reference regularization term; the candidate weight includes the weight of each control experiment group; SB2, fusing the first empty running index data of the multiple control experiment groups through the candidate weight of the target experiment group corresponding to the reference regularization term to obtain the candidate index data of the target experiment group corresponding to the regularization term; SB3, determining the target weight of the target experiment group from the candidate weight of the target experiment group corresponding to the multiple reference regularization terms based on the candidate index data of the target experiment group corresponding to the multiple reference regularization terms.

[0075] That is, multiple reference regularization terms are configured, and the objective function determined for the target experiment group is also multiple, one reference regularization term corresponds to one objective function, thus the objective function corresponding to each reference regularization term can be solved to obtain the weight of each control experiment group under each reference regularization term as the candidate weight of the target experiment group under each reference regularization term.

[0076] Then, for each reference regularization term, the first empty running index data of the multiple control experiment groups are weighted and summed through the candidate weight of the target experiment group corresponding to the reference regularization term to realize the fusion of the first empty running index data of the multiple control experiment groups, and the candidate index data of the target experiment group corresponding to the reference regularization term is obtained, at this time, the candidate index data is the index data of the virtual synthetic group in the reference period without implementing the target strategy after synthesizing the multiple control experiment groups, and the synthetic group is used to be equivalent to the target experiment group.

[0077] In this way, the candidate index data of the target experiment group corresponding to the multiple reference regularization terms is obtained by traversing the multiple reference regularization terms (the candidate index data is multiple, and each reference regularization term corresponds to one candidate index data), and then the target weight of the target experiment group is determined from the candidate weight of the target experiment group corresponding to the multiple reference regularization terms based on the candidate index data of the target experiment group corresponding to the multiple reference regularization terms.

[0078] In some embodiments, SB3 can include: averaging or weighted summing the candidate indicator data corresponding to the target experiment group under the plurality of reference regularization terms, and taking the result as the basic indicator data of the target experiment group; then based on the candidate indicator data corresponding to the target experiment group under the plurality of reference regularization terms and the basic indicator data, selecting one of the candidate indicator data corresponding to the target experiment group under the plurality of reference regularization terms as the intermediate candidate indicator data, and determining the reference regularization term to which the selected intermediate candidate indicator data belongs as the basic regularization term, obtaining the candidate weight corresponding to the target experiment group under the basic regularization term as the target weight corresponding to the target experiment group.

[0079] The intermediate candidate indicator data can be selected as the one with the smallest difference (which can be a difference, an absolute value difference, a square of the difference, etc.) between the basic indicator data and the candidate indicator data corresponding to the target experiment group under the plurality of reference regularization terms, or the median or mode of the difference between the basic indicator data and the candidate indicator data corresponding to the target experiment group under the plurality of reference regularization terms.

[0080] In yet some embodiments, SB3 can include: determining an error indicator corresponding to the target experiment group under the reference regularization term based on the difference between the candidate indicator data corresponding to the target experiment group under the reference regularization term and the first empty run indicator data; determining the target regularization term from the plurality of reference regularization terms based on the error indicators corresponding to the target experiment group under the plurality of reference regularization terms; and obtaining the candidate weight corresponding to the target experiment group under the target regularization term as the target weight corresponding to the target experiment group.

[0081] In the present embodiment, the error indicator can include a mean square error RMSE and a fitting bias BIAS, etc., so as to select the target regularization term from the plurality of reference regularization terms by the mean square error RMSE and / or the fitting bias BIAS. Wherein the mean square error corresponding to the target experiment group under the reference regularization term is: , is the first empty run indicator data of the target experiment group, is the candidate indicator data of the target experiment group; and the fitting bias corresponding to the target experiment group under the reference regularization term is: .

[0082] For example, when the error indicator includes the mean square error RMSE, the one with the smallest mean square error can be selected as the target regularization term, for another example, when the error indicator includes the fitting bias BIAS, the one with the smallest fitting bias can be selected as the target regularization term, and for another example, when the error indicator is a fusion error indicator that fuses the mean square error RMSE and the fitting bias BIAS, the one with the smallest fusion error indicator can be selected as the target regularization term.

[0083] After the target regularization term is determined, the candidate weight corresponding to the target experiment group under the target regularization term can be obtained as the target weight corresponding to the target experiment group, so that a good one is selected from the multiple candidate weights corresponding to the target experiment group as the target weight.

[0084] And because the error index of the target regularization term to which the target weight belongs is smaller, it means that the difference between the first empty running index data corresponding to the target experiment group and the candidate index data corresponding to the target experiment group under the target regularization term (the candidate index data here is the index data of the virtual synthetic group corresponding to the target experiment group obtained by fusing multiple control experiment groups according to the candidate weight corresponding to the target experiment group under the target regularization term) is small, the difference between the target experiment group and the corresponding virtual synthetic group is small, the virtual synthetic group fits the target experiment group better, the virtual synthetic group is more accurate, and the synthesis effect is better.

[0085] S130, by the target weight, fusing the second empty running index data of the multiple control experiment groups in the strategy period to obtain the control index data of the target experiment group in the strategy period.

[0086] Among them, the target strategy is not applied to each control experiment group in the strategy period, and the strategy period is different from the reference period. Usually, the target experiment group is applied with the target strategy in the strategy period, and the index data of the target experiment group in the strategy period is the strategy index data.

[0087] The strategy period is a period different from the reference period based on the demand, usually, the strategy period and the reference period do not coincide, and the target strategy is not applied to each control experiment group in the strategy period. In the strategy period, the index data of each sample in the control experiment group in the empty running state (state without applying the target strategy) is taken as the empty running index data, so that the empty running index data of each sample in the control experiment group is summarized as the second empty running index data of the control experiment group. In the strategy period, the index data of each sample in the target experiment group in the strategy state (state with the target strategy applied) is taken, so that the index data of each sample in the target experiment group is summarized as the strategy index data of the target experiment group.

[0088] After the target weight is determined, the weight of each control experiment group in the target weight can be used to fuse the second empty running index data of the multiple control experiment groups in the strategy period, and the fused data is taken as the control index data of the target experiment group in the strategy period. The control index data is actually the index data of the virtual synthetic group corresponding to the target experiment group in the strategy period without applying the target strategy, so that the index data of the target experiment group in the two cases of applying the target strategy and not applying the target strategy is obtained: the strategy index data in the case of applying the target strategy, and the control index data in the case of not applying the target strategy.

[0089] Generally, the state of the sample in the experimental group can change over time, resulting in the first empty running indicator data of the control experimental group in the reference period being different from the second empty running indicator data of the control experimental group in the strategy period. Therefore, when constructing the control indicator data of the target experimental group in the strategy period, the second empty running indicator data of the control experimental group in the strategy period is used to avoid the situation that the control indicator data does not fit the strategy period and the control indicator data is inaccurate when the control indicator data is synthesized based on the first empty running indicator data of the control experimental group in the reference period.

[0090] S140, evaluate the target strategy based on the control indicator data to obtain a strategy evaluation result of the target strategy.

[0091] After obtaining the control indicator data of the target experimental group in the strategy period, the target strategy can be evaluated based on the control indicator data to obtain a strategy evaluation result of the target strategy.

[0092] As described above, the strategy indicator data of the target experimental group in the strategy period can also be collected. At this time, the strategy evaluation result of the target strategy can be determined based on the control indicator data and the strategy indicator data of the target experimental group in the strategy period.

[0093] In this application, a target test algorithm can also be selected to determine the strategy evaluation result of the target strategy based on the control indicator data and the strategy indicator data of the target experimental group in the strategy period. The target test algorithm can be a T-test algorithm, a placebo test algorithm, a permutation test algorithm, a Placebo algorithm, a Bootstrap algorithm, and a Jackknife algorithm.

[0094] Specifically, the P value and the confidence interval of the target strategy can be determined based on the control indicator data and the strategy indicator data of the target experimental group in the strategy period by using the target test algorithm, and then the strategy evaluation result of the target strategy can be determined based on the P value and the confidence interval. In this application, the null hypothesis for determining the P value and the confidence interval is that the target strategy applied to the target experimental group has no effect. Correspondingly, the P value refers to the possibility of observing the strategy indicator data of the target experimental group under the condition that the target strategy applied to the target experimental group has no effect. Here, the confidence interval can be a 90% confidence interval or a 95% confidence interval, etc.

[0095] When the target strategy is evaluated by P-value, a significant threshold a can be set. If P-value ≤ a, it means that the probability of observing the sample's strategy indicator data in the target experimental group is low when the target strategy has no effect on the target experimental group, which also means that the target strategy can be effective, thus the null hypothesis is rejected, and the strategy evaluation result that the target strategy is effective is obtained. If P-value > a, it means that the probability of observing the sample's strategy indicator data in the target experimental group is high when the target strategy has no effect on the target experimental group, which also means that the target strategy can be ineffective, thus the null hypothesis is not rejected, and the strategy evaluation result that the target strategy is ineffective is obtained.

[0096] When the target strategy is evaluated by the confidence interval, if the confidence interval does not include the control indicator data of the target experimental group, it means that the difference between the strategy indicator data and the control indicator data in the target experimental group is not 0, which also means that the target strategy can be effective, thus the null hypothesis is rejected, and the strategy evaluation result that the target strategy is effective is obtained. If the confidence interval includes the control indicator data of the target experimental group, it means that the difference between the strategy indicator data and the control indicator data in the target experimental group can be 0, which also means that the target strategy can be ineffective, thus the null hypothesis is not rejected, and the strategy evaluation result that the target strategy is ineffective is obtained.

[0097] Further, if the maximum value of the confidence interval is less than the control indicator data of the target experimental group, it means that the target strategy is effective, which can be a negative effect, thus the strategy evaluation result that the target strategy has a negative effect is obtained. Correspondingly, if the minimum value of the confidence interval is greater than the control indicator data of the target experimental group, it means that the target strategy is effective, which can be a positive effect, thus the strategy evaluation result that the target strategy has a positive effect is obtained.

[0098] When the target strategy is evaluated by P-value and the confidence interval, if P-value ≤ a and the confidence interval does not include the control indicator data of the target experimental group, the strategy evaluation result that the target strategy is effective is obtained. Conversely, if P-value > a or the confidence interval includes the control indicator data of the target experimental group, the strategy evaluation result that the target strategy is ineffective is obtained.

[0099] Of course, each experimental group can also be regarded as a target experimental group, and the strategy evaluation result of the target strategy for each experimental group can be determined by the foregoing S110-S140, and the strategy evaluation results of the experimental groups can be combined to obtain the total strategy evaluation result.

[0100] For example, the total evaluation result includes a plurality of strategy evaluation results obtained by taking a plurality of experimental groups as target experimental groups respectively, and the total evaluation result is determined to be effective for the target strategy if the proportion of the strategy evaluation results that are effective for the target strategy is higher than a proportion threshold (for example, 60%), and otherwise, the total evaluation result is determined to be ineffective for the target strategy.

[0101] In this embodiment, first, based on the first empty running index data of the at least three experimental groups in the reference period, the target function of the target experimental group is determined, and the target function introduces a variable intercept for adjusting the weighted results of the first empty running index data of the plurality of control experimental groups, and the weights of the control experimental groups in the constraint condition of the target function are all real numbers and the intercept is not zero, so that the constraint condition no longer limits the weights of the control experimental groups to be greater than zero, the sum of the weights of the plurality of control experimental groups to be 1, and the intercept to be zero, so that the constraint condition has less restrictions on the intercept and the weights of the control experimental groups. Thus, when the target function is solved, the differences between the target experimental group and the control experimental groups can be better captured, the weights determined for the control experimental groups are more accurate, the control index data of the target experimental group in the strategy period is more accurate based on the target weights, and the accuracy of the strategy evaluation result obtained by evaluating the target strategy based on the control index data is higher, thereby improving the evaluation accuracy of the target strategy.

[0102] Secondly, the reference regularization term is also introduced, which controls the model complexity and avoids overfitting in the small sample scenario, further improves the accuracy of the determined target weights, and thus improves the evaluation accuracy of the target strategy.

[0103] Furthermore, the introduced reference regularization term is multiple, which realizes the target of selecting an optimal target regularization term from multiple reference regularization terms, and thus the accuracy of the target weight is further improved when the target weight is determined by using the target function including the target regularization term, thereby further improving the evaluation accuracy of the target strategy.

[0104] In some embodiments, the reference period includes reference periods under different period categories; as Figure 3 As shown in FIG. 1, S110 includes: S111, for each period category, based on the first empty running index data of the at least three experimental groups in the reference period under the period category, a target function corresponding to the target experimental group under the period category is determined.

[0105] The period category can include a weekday category and a holiday category, and the holiday category can be divided into a holiday category and a regular holiday category. The period of the holiday category can refer to the period of the holiday category that belongs to the holiday, and the regular holiday category can refer to the period of the holiday category that belongs to the holiday.

[0106] However, the index difference of the samples in the same experiment group at different time period categories may be large, and therefore, in this application, the reference time periods of different time period categories are processed respectively to obtain the corresponding target functions.

[0107] For example, the time period categories are workday category and holiday category respectively. For the workday category, the target function corresponding to the target experiment group under the workday category is determined based on the first empty running index data of at least three experiment groups in the reference time period under the workday category. Similarly, for the holiday category, the target function corresponding to the target experiment group under the holiday category is determined based on the first empty running index data of at least three experiment groups in the reference time period under the holiday category. The form of the constructed target function is referred to the form in S110 in the foregoing embodiment, which will not be described here.

[0108] Correspondingly, S120 comprises: S121, solving the target function corresponding to the target experiment group under the time period category under the constraint condition to obtain the target weight corresponding to the target experiment group under the time period category.

[0109] Of course, the target weight determined for different time period categories is also different, and the specific solving method is referred to the description of the foregoing S120, which will not be described here.

[0110] It is worth mentioning that in solving the target weight, a plurality of reference regularization terms are also involved, and therefore, in fact, for each time period category, a target regularization term is also selected from a plurality of reference regularization terms, and the candidate weight under the target regularization term is obtained as the target weight.

[0111] Correspondingly, S130 comprises: S131, obtaining the time period category to which the strategy time period belongs as the target time period category; and fusing the second empty running index data of the plurality of control experiment groups based on the target weight corresponding to the target experiment group under the target time period category to obtain the control index data of the target experiment group in the strategy time period.

[0112] That is to say, for any strategy time period, the target weight under the time period category to which the strategy time period belongs is used to fuse the second empty running index data of the plurality of control experiment groups to obtain the control index data of the target experiment group in the strategy time period.

[0113] It is not difficult to understand that if the strategy time period is a strategy time period under a plurality of time period categories, for each strategy time period, the control index data of the target experiment group in the strategy time period is determined by the target weight of the strategy time period (the target weight may be different for different categories of strategy time period, and the target weight can be the same for the strategy time period of the same time period category), and at the same time, the strategy evaluation result of each strategy time period is determined according to the means of S140.

[0114] Then, the policy evaluation results of the multiple policy periods are combined, and the policy evaluation results of different policy periods are fused for the policy periods under different time period categories to obtain a final target evaluation result.

[0115] For example, the target evaluation result includes the policy evaluation results of the multiple policy periods, and the target evaluation result is determined to be valid for the target policy when the proportion of the policy evaluation results that are valid for the target policy is higher than a proportion threshold (for example, 60%), and otherwise, the target evaluation result is determined to be invalid for the target policy.

[0116] For another example, the target evaluation result includes the policy evaluation results of the multiple policy periods, and the total duration of the policy periods in which the target policy is valid and the total duration of all the policy periods are counted, and the target evaluation result is determined to be valid for the target policy when the ratio of the total duration of the policy periods in which the target policy is valid to the total duration of all the policy periods is higher than a proportion threshold (for example, 70%), and otherwise, the target evaluation result is determined to be invalid for the target policy.

[0117] Of course, each experimental group can also be taken as a target experimental group, and the target evaluation result of the target policy for each experimental group can be determined according to the foregoing S111-S131 and S140, and the target evaluation results of the multiple experimental groups can be combined to obtain a total target evaluation result.

[0118] For example, the total target evaluation result includes the multiple target evaluation results obtained when the multiple experimental groups are taken as target experimental groups, and the total target evaluation result is determined to be valid for the target policy when the proportion of the target evaluation results that are valid for the target policy is higher than a proportion threshold (for example, 60%), and otherwise, the total target evaluation result is determined to be invalid for the target policy.

[0119] In the embodiment, different target weights are used to fuse the second empty running index data of the multiple control experimental groups for the policy periods under different time period categories, so that the fusion effect for different time period categories is high, the control index data obtained for the policy periods under different time period categories is more accurate, and the target policy evaluation is more accurate, and the credibility of the policy evaluation result of the target policy is higher.

[0120] In some embodiments, as shown in FIG. 1A, Figure 4 S140 includes: S141, obtaining multiple improvement amplitudes and policy index data of the target experimental group in the policy period.

[0121] The improvement amplitude refers to the amplitude of the improvement of the index data, and the improvement amplitude is, for example, 1%, 2%, or 3%, etc.

[0122] S142, adjust the first empty running index data of each experimental group by each promotion amplitude respectively to obtain the adjusted index data of each experimental group under each promotion amplitude.

[0123] For each promotion amplitude, the first empty running index data of each experimental group is increased by the promotion amplitude to obtain the adjusted index data of each experimental group under the promotion amplitude.

[0124] For example, the experimental groups include a1, a2, a3 and a4, the promotion amplitude is 1%, the first empty running index data of the experimental group a1 is a11, the first empty running index data of the experimental group a2 is a21, the first empty running index data of the experimental group a3 is a31, and the first empty running index data of the experimental group a4 is a41. At this time, the adjusted index data obtained for the experimental group a1 under the promotion amplitude of 1% is 101% a11, the adjusted index data obtained for the experimental group a2 is 101% a21, the adjusted index data obtained for the experimental group a3 is 101% a31, and the adjusted index data obtained for the experimental group a4 is 101% a41.

[0125] S143, for each promotion amplitude, the reference test algorithm is used to determine the test efficiency evaluation result of the reference test algorithm under the promotion amplitude based on the adjusted index data of at least three experimental groups and the first empty running index data of the at least three experimental groups under the promotion amplitude.

[0126] The reference test algorithm is any one of a plurality of candidate test algorithms; the plurality of candidate test algorithms can include T-test algorithm, placebo test algorithm, permutation test algorithm, Placebo algorithm, Bootstrap algorithm and Jackknife algorithm, etc.

[0127] For any promotion amplitude, the test effect of the reference test algorithm is evaluated based on the difference between the adjusted index data of at least three experimental groups and the first empty running index data of the at least three experimental groups under the promotion amplitude to obtain the test efficiency evaluation result of the reference test algorithm under the promotion amplitude.

[0128] Generally, the index difference of the target experimental group is determined based on the difference (which can be the difference value, the absolute value difference or the square of the difference, etc.) between the adjusted index data of the target experimental group and the first empty running index data of the target experimental group under the promotion amplitude by using the reference test algorithm, and the index difference of the control experimental group is also determined in the same way, and then the index differences of at least three experimental groups are fused as the total index difference, and the test efficiency evaluation result of the reference test algorithm under the promotion amplitude is determined based on the total index difference.

[0129] For example, the smaller the total index difference is, the better the test efficiency of the reference test algorithm is, and the larger the total index difference is, the worse the test efficiency of the reference test algorithm is.

[0130] In some embodiments, S143 can include: SC1, for each lifting amplitude, determining, by the reference test algorithm, a probability that the reference test algorithm detects a type I error at the lifting amplitude based on a difference between the adjustment index data of the at least three experimental groups and the first empty running index data of the at least three experimental groups at the lifting amplitude, as a type I error rate of the reference test algorithm at the lifting amplitude; SC2, determining, based on the type I error rate of the reference test algorithm at the lifting amplitude, a test efficiency evaluation result of the reference test algorithm at the lifting amplitude.

[0131] The type I error refers to rejecting the original hypothesis when the original hypothesis is true. That is, first, based on the difference between the adjustment index data of the at least three experimental groups and the first empty running index data of the at least three experimental groups at the lifting amplitude, a probability that the candidate test algorithm detects a rejection of the original hypothesis is determined as a type I error rate, and then a test efficiency evaluation result of the reference test algorithm at the lifting amplitude is determined based on the type I error rate.

[0132] As described above, the adjustment index data is manually changed and is not caused by implementing the target strategy, and therefore, the type I error is detected based on the adjustment index data and the first empty running index data, and therefore, the higher the type I error rate of the reference test algorithm is, the higher the ability of the reference test algorithm to detect the type I error is, the lower the type I error rate of the reference test algorithm is, the better the test effect of the reference test algorithm is, and the lower the ability of the reference test algorithm to detect the type I error is, the worse the test effect of the reference test algorithm is. Therefore, if the type I error rate is higher than a probability threshold (the probability threshold is, for example, 0.5, etc.), it means that the reference test algorithm can detect the type I error with a high probability, and a test efficiency evaluation result of the reference test algorithm with a good detection effect is obtained, and correspondingly, if the type I error rate is not higher than the probability threshold, it means that the reference test algorithm cannot detect the type I error with a high probability, and a test efficiency evaluation result of the reference test algorithm with a poor detection effect is obtained.

[0133] In some embodiments, the type I error rate of the reference test algorithm at the lifting amplitude can be obtained as the test efficiency evaluation result of the reference test algorithm at the lifting amplitude In yet some embodiments, the adjustment indicator data comprises single-dimension adjustment indicator data of each of the plurality of dimensions; the first empty run indicator data comprises single-dimension empty run indicator data of each of the plurality of dimensions, whereby the aforementioned SC1 comprises: for each of the plurality of lift magnitudes, determining, by the reference test algorithm, a P-value between the single-dimension adjustment indicator data of the reference experimental group in the reference dimension and the single-dimension empty run indicator data of the reference experimental group in the reference dimension under the lift magnitude as the P-value of the reference experimental group in the reference dimension under the reference test algorithm under the lift magnitude; the reference experimental group is any one of the at least three experimental groups, and the reference dimension is any one of the plurality of dimensions; and determining the type I error rate of the reference test algorithm under the lift magnitude based on the P-values of the at least three experimental groups in the plurality of dimensions under the reference test algorithm under the lift magnitude.

[0134] In the present embodiment, the indicator data is multi-dimensional, for example, the test subject is a shopping application, the sample is a user, the indicator quantity is the number of uses and the use duration per day, and the involved dimensions are two. Therefore, the type I error rate of the reference test algorithm is determined in combination with the indicator data of the plurality of dimensions.

[0135] For example, when the lift magnitude is c1, the experimental groups are d1, d2 and d3, and the dimensions are b1 and b2, at this time, by the reference test algorithm, the P-value between the single-dimension adjustment indicator data of the reference experimental group d1 in the dimension b1 and the single-dimension empty run indicator data of the reference experimental group d1 in the dimension b1 under the lift magnitude c1 is determined as the P-value of the reference experimental group d1 in the dimension b1 under the reference test algorithm under the lift magnitude c1, and at the same time, by the reference test algorithm, the P-value between the single-dimension adjustment indicator data of the reference experimental group d1 in the dimension b2 and the single-dimension empty run indicator data of the reference experimental group d1 in the dimension b2 under the lift magnitude c1 is determined as the P-value of the reference experimental group d1 in the dimension b2 under the reference test algorithm under the lift magnitude c1, and in this way, for the lift magnitude c1, the P-value of the reference experimental group d1 in the dimension b1 under the reference test algorithm and the P-value of the reference experimental group d1 in the dimension b2 under the reference test algorithm are determined, and correspondingly, the P-value of the reference experimental group d2 in the dimension b1 under the reference test algorithm and the P-value of the reference experimental group d2 in the dimension b2 under the reference test algorithm are determined, and the P-value of the reference experimental group d3 in the dimension b1 under the reference test algorithm and the P-value of the reference experimental group d3 in the dimension b2 under the reference test algorithm are determined, at this time, the reference experimental group d1 under the lift magnitude c1 obtains 6 P-values.

[0136] Then, based on the determined plurality of P-values (usually the number of P-values is the product of the number of experimental groups and the number of dimensions), the type I error rate of the reference test algorithm under the lift magnitude c1 is determined in combination with the plurality of P-values.

[0137] For the reference test algorithm, select the P values higher than the significant threshold value from the P values of the reference test algorithm in multiple dimensions of at least three experimental groups under the promotion amplitude as the significant P values, determine the ratio of the number of significant P values to the total number of P values as the false positive rate, that is, the false positive rate P1= For example, in the foregoing example, the promotion amplitude c1 obtains a total of 6 P values, and the significant P values are 3, so the false positive rate is 1 / 2.

[0138] In some embodiments, the foregoing SC2 can further include: determining a false positive controllable test result of the reference test algorithm under the promotion amplitude based on the false positive rate of the reference test algorithm under the promotion amplitude, the number of dimensions of the multiple dimensions, and the number of experimental groups of the at least three experimental groups; the false positive controllable test result of the reference test algorithm is used to indicate whether the false positive detected by the reference test algorithm is controllable; determining the test efficacy evaluation result of the reference test algorithm under the promotion amplitude based on the false positive rate of the reference test algorithm under the promotion amplitude and the false positive controllable test result.

[0139] That is, for any promotion amplitude c1, in combination with the false positive rate of the reference test algorithm under the promotion amplitude c1, the number of dimensions of the multiple dimensions, and the number of experimental groups of the at least three experimental groups, it is evaluated whether the false positive of the reference test algorithm under the promotion amplitude c1 is controllable to obtain the false positive controllable test result of the reference test algorithm under the promotion amplitude c1, and then in combination with the false positive rate of the reference test algorithm under the promotion amplitude c1 and the false positive controllable test result, the test efficacy evaluation result of the reference test algorithm under the promotion amplitude is obtained.

[0140] In some embodiments, for any promotion amplitude c1, in combination with the number of dimensions of the multiple dimensions and the number of experimental groups of the at least three experimental groups, a controllable threshold value is determined, if the false positive rate of the reference test algorithm under the promotion amplitude c1 is higher than the controllable threshold value, it is determined that the false positive of the reference test algorithm is controllable, otherwise it is not controllable, here, the method for determining the controllable threshold value can be: multiplying the number of dimensions of the multiple dimensions and the number of experimental groups of the at least three experimental groups, and then multiplying the product result by a specified coefficient, the method for determining the controllable threshold value can also be: finding the ratio of the number of dimensions of the multiple dimensions and the number of experimental groups of the at least three experimental groups, and then multiplying the ratio by a specified coefficient.

[0141] In yet some embodiments, a controllable interval can also be determined based on the preset significant threshold value, the number of dimensions of the multiple dimensions, and the number of experimental groups of the at least three experimental groups; the false positive controllable test result of the reference test algorithm under the promotion amplitude is determined based on the comparison result of the false positive rate of the reference test algorithm under the promotion amplitude and the controllable interval.

[0142] The adjustment amount can be determined based on the dimension number of the plurality of dimensions and the experiment group number of the at least three experiment groups, and then the difference between the significant threshold and the adjustment amount is calculated as the lower limit value of the controllable interval, and the sum of the significant threshold and the adjustment amount is calculated as the upper limit value of the controllable interval. The ratio, product, etc. of the dimension number of the plurality of dimensions and the experiment group number of the at least three experiment groups can be determined as an initial operation result, and then the initial operation result is normalized, multiplied by a specified coefficient, etc. to obtain a result as the adjustment amount.

[0143] Alternatively, the adjustment coefficient s1 and the adjustment coefficient s2 can be determined based on the dimension number of the plurality of dimensions and the experiment group number of the at least three experiment groups, the adjustment coefficient s1 is smaller than the adjustment coefficient s2, the product of the significant threshold and s1 is the lower limit value of the controllable interval, and the product of the significant threshold and s2 is the upper limit value of the controllable interval. The ratio of the dimension number of the plurality of dimensions and the experiment group number of the at least three experiment groups and the inverse of the ratio can be determined, and the smaller one of the ratio and the inverse of the ratio is obtained as s1, and the other one is obtained as s2.

[0144] Further alternatively, a reference amount can be determined based on the dimension number of the plurality of dimensions and the experiment group number of the at least three experiment groups, and then the difference between the reference amount and the significant threshold is calculated as the lower limit value of the controllable interval, and the sum of the reference amount and the significant threshold is calculated as the upper limit value of the controllable interval. The way of determining the reference amount can include determining the ratio of the dimension number of the plurality of dimensions and the experiment group number of the at least three experiment groups and the inverse of the ratio, and obtaining the smaller one of the ratio and the inverse of the ratio as the reference amount.

[0145] For example, the controllable interval can be determined based on the preset significant threshold, the dimension number of the plurality of dimensions and the experiment group number of the at least three experiment groups in the manner of Formula Four, and Formula Four is as follows: (Four) wherein a is the significant threshold, is the 97.5% lower quantile of the normal distribution, and n is the product of the dimension number of the plurality of dimensions and the experiment group number of the at least three experiment groups. In Formula Four, the adjustment amount is That is, the adjustment amount integrates the preset significant threshold, the dimension number of the plurality of dimensions and the experiment group number of the at least three experiment groups.

[0146] If the type I error rate of the reference test algorithm under the improvement amplitude is located in the corresponding controllable interval, it is determined that the type I error controllable test result of the reference test algorithm under the improvement amplitude is type I error controllable, and if the type I error rate of the reference test algorithm under the improvement amplitude is not located in the corresponding controllable interval, it is determined that the type I error controllable test result of the reference test algorithm under the improvement amplitude is type I error uncontrollable.

[0147] As described above, for each lifting amplitude, the evaluation result of the test efficiency of the reference test algorithm under the lifting amplitude is obtained, thereby for each lifting amplitude, the evaluation results of the test efficiency of the plurality of reference test algorithms under the lifting amplitude are obtained by traversing the plurality of candidate test algorithms, and the evaluation results of the test efficiency of the plurality of reference test algorithms under the plurality of lifting amplitudes are obtained by continuing to traverse the plurality of lifting amplitudes. At this time, the evaluation result of the test efficiency is a plurality, and the number is the product of the number of the plurality of lifting amplitudes and the number of the plurality of reference test algorithms.

[0148] In S144, the target test algorithm is determined from the plurality of candidate test algorithms based on the evaluation result of the test efficiency of each of the plurality of candidate test algorithms under the plurality of lifting amplitudes.

[0149] In some embodiments, S144 can include: for each candidate test algorithm, determining a first test algorithm from the plurality of candidate test algorithms based on the number of key lifting amplitudes corresponding to the candidate test algorithm; the key lifting amplitude corresponding to the candidate test algorithm refers to the lifting amplitude in the plurality of lifting amplitudes in which the candidate test algorithm detects a controllable type I error; if the first test algorithm is a plurality, determining the target test algorithm from the first test algorithm based on the size of the type I error rate of the plurality of first test algorithms under the plurality of lifting amplitudes.

[0150] In some embodiments, S144 can include: for each candidate test algorithm, determining a first test algorithm from the plurality of candidate test algorithms based on the number of key lifting amplitudes corresponding to the candidate test algorithm; the key lifting amplitude corresponding to the candidate test algorithm refers to the lifting amplitude in the plurality of lifting amplitudes in which the candidate test algorithm detects a controllable type I error; if the first test algorithm is a plurality, determining the target test algorithm from the first test algorithm based on the size of the type I error rate of the plurality of first test algorithms under the plurality of lifting amplitudes. For each candidate test algorithm, traverse the type I error controllable test result of the candidate test algorithm under the plurality of lifting amplitudes. If the type I error controllable test result is a type I error controllable, the lifting amplitude is the key lifting amplitude of the candidate test algorithm.

[0151] The plurality of first test algorithms can be selected from the candidate test algorithm in the order from high to low according to the number of key lifting amplitudes, or the candidate test algorithm with the number of key lifting amplitudes greater than the number threshold (the number threshold is usually determined based on the number of lifting amplitudes, for example, the number threshold is 0.8 times the number of lifting amplitudes) is selected as the first test algorithm.

[0152] If the first test algorithm is a plurality, one of the plurality of first test algorithms can be selected as the target test algorithm, wherein the selection means is based on the size of the type I error rate of the plurality of first test algorithms under the plurality of lifting amplitudes.

[0153] As known from the foregoing, the higher the one-class error rate of the candidate test algorithm, the higher the ability of the candidate test algorithm to detect one-class errors, the lower the one-class error rate of the candidate test algorithm, the better the test effect of the candidate test algorithm, the lower the ability of the candidate test algorithm to detect one-class errors, the worse the test effect of the candidate test algorithm, and thus, the first test algorithm with the highest one-class error rate can be selected as the target test algorithm.

[0154] S145, determining a strategy evaluation result of the target strategy based on the control indicator data and the strategy indicator data of the target experiment group in the strategy period by the target test algorithm.

[0155] The evaluation process of S145 refers to the description in S140, and will not be repeated here.

[0156] In some embodiments, if the candidate test algorithm is configured by one of the plurality of hyperparameters, the method further comprises: obtaining a plurality of second test algorithms; different second test algorithms are obtained by configuring the initial test algorithm with different hyperparameters; for each boost amplitude, determining the test efficacy evaluation result of the reference second test algorithm under the boost amplitude based on the adjustment indicator data of the plurality of experiment groups under the boost amplitude and the control indicator data of the plurality of experiment groups by referring to the second test algorithm; the reference second test algorithm is any one of the plurality of second test algorithms; based on the test efficacy evaluation results of the plurality of second test algorithms under the plurality of boost amplitudes, determining a third test algorithm from the plurality of second test algorithms; if the number of third test algorithms is more than one, based on the size of the hyperparameters configured by the plurality of third test algorithms, selecting one from the plurality of third test algorithms as the candidate test algorithm. The initial test algorithm may, for example, be a T-Test algorithm.

[0157] It is not difficult to understand that each second test algorithm can be regarded as a candidate test algorithm, and the test efficacy evaluation result of the second test algorithm can be determined according to the foregoing processes of S141-S143. Accordingly, the test efficacy evaluation result of the second test algorithm also includes the one-class error rate and the one-class error controllable test result.

[0158] The plurality of second test algorithms can be selected as third test algorithms in descending order of the number of key boost amplitudes, or the second test algorithm with a number of key boost amplitudes greater than a number threshold (the number threshold is usually determined based on the number of boost amplitudes, for example, the number threshold is 0.8 times the number of boost amplitudes) can be selected as the third test algorithm.

[0159] If the third test algorithm is multiple, one of the multiple first test algorithms can be selected as a candidate test algorithm, wherein the selection means is based on the hyperparameters of the multiple third test algorithms under multiple lift magnitudes. For example, for the initial test algorithm being a T-Test algorithm, the third test algorithm with the maximum hyperparameter is determined as the candidate test algorithm.

[0160] It can be understood that for each initial test algorithm, multiple second test algorithms are obtained after the corresponding multiple hyperparameters are configured, and a candidate test algorithm is determined for the initial test algorithm according to the foregoing means, so that a candidate test algorithm can be determined for each initial test algorithm.

[0161] However, some test algorithms do not have hyperparameters, and these test algorithms themselves can be a candidate test algorithm.

[0162] It is again pointed out that the reference period can include reference periods of multiple period categories, and any one period category is a reference category, at this time, the first empty run index data of each experimental group in the reference period under the reference category can be adjusted by each lift magnitude respectively to obtain the adjusted index data of each experimental group in the reference period under the reference category under each lift magnitude; for each lift magnitude, the reference test algorithm is used to determine the test efficacy evaluation result of the reference test algorithm for the reference period category under the lift magnitude based on the adjusted index data of at least three experimental groups in the reference period under the reference category and the first empty run index data in the reference period under the reference category under the lift magnitude; based on the test efficacy evaluation results of the multiple candidate test algorithms for the reference period category respectively under the multiple lift magnitudes, the target test algorithm for the reference period category is determined from the multiple candidate test algorithms, so that a target test algorithm is obtained for each period category. Correspondingly, the period category to which the strategy period belongs is the target period category, and thus the target strategy evaluation result of the target strategy is determined based on the control index data and the strategy index data of the target experimental group in the strategy period by the target test algorithm corresponding to the target period category.

[0163] Of course, the strategy period can include multiple strategy periods of different period categories, for each strategy period, the target strategy evaluation result of the target strategy is determined based on the control index data and the strategy index data of the target experimental group in the strategy period by the target test algorithm corresponding to the target period category to which each strategy period belongs, and finally, the strategy evaluation results of the multiple strategy periods are summarized to obtain the target evaluation result.

[0164] In this embodiment, a target test algorithm is selected from multiple candidate test algorithms, which realizes the optimal use of the test algorithm, so that the evaluation effect is better when the target strategy is evaluated according to the target test algorithm, and the accuracy of the obtained strategy evaluation result is higher.

[0165] In addition, in the case that the test algorithm is configured by multiple hyperparameters, the best hyperparameter can also be selected for the test algorithm, so that the selection accuracy of the test algorithm is high, and it is easier to select a test algorithm with high detection effect to evaluate the target strategy.

[0166] Furthermore, different target test algorithms can also be selected for different time period categories, which further improves the pertinence of the target test algorithm, so that the evaluation effect of the target test algorithm for the target strategy evaluation is better, and the obtained strategy evaluation result is more accurate.

[0167] In order to more conveniently understand the scheme of the present application, the strategy evaluation method of the present application is explained below in combination with an example.

[0168] Firstly, the architecture of the evaluation system for evaluating the strategy in this example is as follows: ├── Data layer │ ├── Data access module │ ├── Data preprocessing module │ └── Data quality monitoring module ├── Algorithm layer │ ├── Weight fitting module │ │ ├── Workday effect │ │ ├── Weekend effect │ ├── Statistical inference module │ │ ├── T-Test │ │ ├── Placebo Test │ │ ├── Bootstrap Test │ │ ├── Permutation Test │ │ ├── Jackknife Test │ └── Model selection module └── Application layer ├── Visualization board └── Report generation Among them, the data access module is used to monitor the index data of the test group, the data preprocessing module is used to preprocess the index data (remove outliers, filter processing and format adjustment processing, etc.); the data quality monitoring module is used to manage the quality of the index data, for example, discarding the index data with poor quality.

[0169] The weight fitting module is used to determine the target weight of the target experiment group. When fitting the target weight, the weekday effect and the weekend effect are referenced to fit a target weight for weekdays and a target weight for weekends.

[0170] The statistical inference module is used to determine the test power evaluation results of T-Test, Placebo Test, Bootstrap Test, Permutation Test, and Jackknife Test, and the model selection module is used to select a target test algorithm based on the test power evaluation results of the given test algorithm.

[0171] The application layer is used to evaluate the target strategy based on the control indicator data and the strategy indicator data of the target experiment group in the strategy period through the target test algorithm, obtain the strategy evaluation result of the target strategy, and generate a report based on the strategy evaluation result. The strategy evaluation result and the report can also be directly output in the visual board. The report can include specific P-value, centroid interval, target test algorithm used, and target weight, etc.

[0172] In this example, the test subject is a conference application, the target strategy can be a member fee rule added to the conference application, and the test group is 34 different cities. The evaluation process of the target strategy is as follows: ├── 1. Experiment configuration phase │ ├── Select experiment group (34 cities) │ ├── Set time window (involve reference period and strategy period) │ ├── Select index quantity (such as number of participants, participation duration, purchase amount, and member order quantity, etc.) │ └── Configure algorithm hyperparameters (default hyperparameters / manual hyperparameters) ├── 2. Data preparation phase │ ├── Data upload (acquire index data) │ ├── Data quality check and cleaning (select non-missing data cities, and index data of weekdays and weekends) │ └── Experiment group screening (such as excluding cities controlled by special strategies) ├── 3. Algorithm running phase │ ├── Weight fitting calculation (determine target weight) │ ├── Test algorithm optimization (determine target test algorithm) │ └── Result verification └── 4. Result analysis phase ├── View the policy evaluation results of the target policy (including the test power evaluation results, P values, confidence intervals, etc. of different test algorithms) └── Generate an evaluation report (including whether the target policy is significant and the test power evaluation results of different test algorithms, etc.) First, as shown in Figure 5 , separate modeling for weekends and weekdays is performed to determine the target weight for weekdays and the target weight for weekends, and in the process of determining the target weight, the restriction is relaxed when creating the target function: the weight can be negative, the weight sum can not be 1, and the intercept can not be zero.

[0173] When creating the target function, the empty run index data (first empty run index data) of the 34 experimental groups in the reference period (i.e. the period of empty running state) is used to determine the error index, which includes the mean square error and the fitting bias, to determine the target weight based on the error index.

[0174] Then, as shown in Figure 6 , method optimization is performed based on the empty run index data (first empty run index data) of the 34 experimental groups in the reference period to select one from the T-test algorithm, placebo test algorithm, permutation test algorithm, Placebo algorithm, Bootstrap algorithm and Jackknife algorithm in the test algorithm pool as the target test algorithm (where selecting the target test algorithm involves hyperparameter optimization and test algorithm optimization, such as selecting hyperparameters for the T-test algorithm), to evaluate the target policy through the target test algorithm.

[0175] For example, for the T-test algorithm, the process of selecting hyperparameters can be as follows: Data segmentation: divide the reference period T into K consecutive intervals, , define: Cross-validation estimation: for each interval, , calculate the treatment effect estimator; Where: is the weight obtained based on the index data of the reference period .

[0176] Construct the T statistic: Statistical inference: calculate the P value and confidence interval The detection efficiency of the T-test algorithm configured by the hyperparameter K depends on the selection of the hyperparameter K-Fold, so a K value that meets the conditions (for example, the K value is configured, and the false positive rate is controllable) can be selected from multiple K values. The T-test algorithm configured by the selected K value is used as the third test algorithm, and the detection efficiency of the third test algorithm is relatively good. If the selected hyperparameter K that meets the conditions is 2, 3, and 4, the maximum hyperparameter value “4” can be selected to configure the T-test algorithm. The configured T-test algorithm is used as a candidate test algorithm, and the candidate test algorithm is the one with the best detection efficiency among the T-test algorithms configured by multiple hyperparameters K.

[0177] For example, if the determined target test algorithm is Placebo, the process of determining the P value and the confidence interval when evaluating the target strategy is as follows: First, calculate the effect size : is the second empty running index data of the control experiment group, is the weight of the control experiment group i, and in this example, the control experiment group has 33.

[0178] In turn, assume that the cities in the experiment group are the target experiment group, and repeat the determination process described above to obtain , There are 34 cities, and the variance of the effect size is calculated as follows: Then, calculate the Placebo test statistic Calculate the corresponding P value and confidence interval where a is a preset significant threshold.

[0179] In this example, compared with existing strategy evaluation methods, the empirical results of the conference application program are: • Mean indicators (such as average meeting duration): RMSE is reduced from 3.89% to 3.39%, a decrease of 12.8% • Count indicators (such as the number of participants): RMSE is reduced from 25385.4 to 16656.8, a decrease of 34.4% • Proportional metrics (e.g., paid conversion rate): RMSE is reduced from 0.46% to 0.30%, a 34.8% reduction • Retention metrics (e.g., 7-day retention rate): RMSE is reduced from 0.86% to 0.72%, a 16.3% reduction In this example, the introduction of the intercept term allows the model to better capture the systematic differences between the experimental cities and the control cities. The separation of weekday and weekend modeling fully considers the differences in usage patterns of the application in different time period categories, especially for office products such as the conference application.

[0180] Furthermore, the present scheme greatly improves the robustness and applicability of the evaluation of the target strategy by innovatively applying the K-fold T-Test method and introducing the Placebo test.

[0181] See Figure 7 , Figure 7 A block diagram of a strategy evaluation device according to an embodiment of the present application is shown. The device 800 includes: The determination module 810 is configured to determine a target function of a target experimental group based on first empty running index data of at least three experimental groups in a reference period. No target strategy is applied to each experimental group in the reference period. The at least three experimental groups include the target experimental group. The target function includes a weighted result of first empty running index data of a plurality of control experimental groups and an intercept for adjusting the weighted result. The plurality of control experimental groups are experimental groups other than the target experimental group in the at least three experimental groups. The solution target of the target function is to minimize the target function by adjusting the weights of the plurality of control experimental groups and the intercept; The solution module 820 is configured to solve the target function under a preset constraint condition to obtain a target weight corresponding to the target experimental group. The target weight includes the weights of the plurality of control experimental groups. The constraint condition includes that the weights of the control experimental groups are real numbers and the intercept is not zero. The fusion module 830 is configured to fuse second empty running index data of the plurality of control experimental groups in a strategy period by using the target weight to obtain control index data of the target experimental group in the strategy period. No target strategy is applied to each control experimental group in the strategy period. The strategy period is different from the reference period. The evaluation module 840 is configured to evaluate the target strategy based on the control index data to obtain a strategy evaluation result of the target strategy.

[0182] Optionally, the reference period includes reference periods under different period categories; the determination module 810 is further configured to determine, for each period category, a target function corresponding to the target experiment group under the period category based on the first empty running index data of the at least three experiment groups in the reference period under the period category; the solving module 820 is further configured to solve the target function corresponding to the target experiment group under the period category under the constraint condition to obtain a target weight corresponding to the target experiment group under the period category; and the fusion module 830 is further configured to obtain a period category to which the strategy period belongs as a target period category, and fuse the second empty running index data of the plurality of control experiment groups based on the target weight corresponding to the target experiment group under the target period category to obtain control index data of the target experiment group in the strategy period.

[0183] Optionally, the determination module 810 is further configured to obtain a reference regularization term; the reference regularization term is associated with the weights of the plurality of control experiment groups; and the target function is determined based on the first empty running index data of the at least three experiment groups and the reference regularization term.

[0184] Optionally, the reference regularization term is a plurality; the target function includes a target function of the target experiment group for the plurality of reference regularization terms; the solving module 820 is further configured to, for each reference regularization term, solve the target function of the target experiment group for the reference regularization term under the constraint condition to obtain a candidate weight corresponding to the target experiment group under the reference regularization term; the candidate weight includes the respective weights of the plurality of control experiment groups; the first empty running index data of the plurality of control experiment groups is fused by referring to the candidate weight corresponding to the target experiment group under the reference regularization term to obtain candidate index data corresponding to the target experiment group under the regularization term; and the target weight corresponding to the target experiment group is determined from the candidate weight corresponding to the target experiment group under the plurality of reference regularization terms based on the candidate index data corresponding to the target experiment group under the plurality of reference regularization terms.

[0185] Optionally, the solving module 820 is further configured to determine an error index corresponding to the target experiment group under the reference regularization term based on a difference between the candidate index data corresponding to the target experiment group under the reference regularization term and the first empty running index data; determine a target regularization term from the plurality of reference regularization terms based on the error index corresponding to the target experiment group under the plurality of reference regularization terms; and obtain the candidate weight corresponding to the target experiment group under the target regularization term as the target weight corresponding to the target experiment group.

[0186] Optionally, the solving module 820 is further configured to determine the reference regularization term based on at least one of a square sum of the weights of the plurality of control experiment groups and an absolute value sum of the weights of the plurality of control experiment groups.

[0187] Optionally, the evaluation module 840 is further configured to: obtain a plurality of lift magnitudes and strategy indicator data of the target experiment group in a strategy period; the target strategy is applied to the target experiment group in the strategy period; adjust the first empty running indicator data of the target experiment group by each of the plurality of lift magnitudes, to obtain adjusted indicator data of the target experiment group under each of the plurality of lift magnitudes; for each of the plurality of lift magnitudes, determine a test power evaluation result of the reference test algorithm under the lift magnitude based on the adjusted indicator data of the at least three experiment groups under the lift magnitude and the first empty running indicator data of the at least three experiment groups by the reference test algorithm; the reference test algorithm is any one of a plurality of candidate test algorithms; determine the target test algorithm from the plurality of candidate test algorithms based on the test power evaluation result of each of the plurality of candidate test algorithms under the plurality of lift magnitudes; and determine a strategy evaluation result of the target strategy based on the control indicator data and the strategy indicator data of the target experiment group in the strategy period by the target test algorithm.

[0188] Optionally, the evaluation module 840 is further configured to, for each of the plurality of lift magnitudes, determine a probability of the reference test algorithm detecting a type I error under the lift magnitude as a type I error rate of the reference test algorithm under the lift magnitude based on a difference between the adjusted indicator data of the at least three experiment groups under the lift magnitude and the first empty running indicator data of the at least three experiment groups by the reference test algorithm; and determine the test power evaluation result of the reference test algorithm under the lift magnitude based on the type I error rate of the reference test algorithm under the lift magnitude.

[0189] Optionally, the adjusted indicator data comprises single-dimension adjusted indicator data of each of a plurality of dimensions; the first empty running indicator data comprises single-dimension empty running indicator data of each of the plurality of dimensions; and the evaluation module 840 is further configured to, for each of the plurality of lift magnitudes, determine, by the reference test algorithm, a P value between the single-dimension adjusted indicator data of the reference experiment group in the reference dimension and the single-dimension empty running indicator data of the reference experiment group in the reference dimension under the lift magnitude as a P value of the reference experiment group in the reference dimension for the reference test algorithm under the lift magnitude; the reference experiment group is any one of the at least three experiment groups, and the reference dimension is any one of the plurality of dimensions; and determine the type I error rate of the reference test algorithm under the lift magnitude based on the P values of the at least three experiment groups in the plurality of dimensions for the reference test algorithm under the lift magnitude.

[0190] Optionally, the first empty-run indicator data includes single-dimensional empty-run indicator data for each of multiple dimensions; the evaluation module 840 is further used to determine the controllability test result of the reference test algorithm for the type I error under the improvement magnitude based on the type I error rate of the reference test algorithm under the improvement magnitude, the number of dimensions of multiple dimensions, and the number of experimental groups of at least three experimental groups; the controllability test result of the reference test algorithm for the type I error is used to indicate whether the type I error detected by the reference test algorithm is controllable; based on the type I error rate of the reference test algorithm under the improvement magnitude and the controllability test result of the type I error, the evaluation result of the test efficacy of the reference test algorithm under the improvement magnitude is determined.

[0191] Optionally, the evaluation module 840 is also used to determine the controllable interval based on a preset significance threshold, the number of dimensions in multiple dimensions, and the number of experimental groups of at least three experimental groups; and to determine the controllable test result of the first type error of the reference test algorithm under the improvement magnitude based on the comparison result of the first type error rate of the reference test algorithm with the controllable interval.

[0192] Optionally, the evaluation module 840 is further configured to, for each candidate testing algorithm, determine a first testing algorithm from multiple candidate testing algorithms based on the number of key improvement margins corresponding to the candidate testing algorithm; the key improvement margin corresponding to the candidate testing algorithm refers to the improvement margin of a type of error that the candidate testing algorithm detects that is controllable among multiple improvement margins; if there are multiple first testing algorithms, the target testing algorithm is determined from the first testing algorithms based on the magnitude of the type of error rate of the multiple first testing algorithms under multiple improvement margins.

[0193] Optionally, the evaluation module 840 is also used to acquire multiple second-test algorithms; different second-test algorithms are obtained by configuring the initial test algorithm with different hyperparameters; for each improvement, the test efficacy evaluation result of the reference second-test algorithm under the improvement is determined by referring to the adjusted index data of multiple experimental groups and the control index data of multiple experimental groups based on the improvement of the reference second-test algorithm; the reference second-test algorithm is any one of the multiple second-test algorithms; based on the test efficacy evaluation results of the multiple second-test algorithms under multiple improvement, a third-test algorithm is determined from the multiple second-test algorithms; if there are multiple third-test algorithms, one is determined from the multiple third-test algorithms as a candidate test algorithm based on the size of the hyperparameters configured for the multiple third-test algorithms.

[0194] It should be noted that the device embodiments in this application correspond to the aforementioned method embodiments. The specific principles in the device embodiments can be found in the content of the aforementioned method embodiments, and will not be repeated here.

[0195] Figure 8 A structural block diagram of an electronic device for performing a strategy evaluation method according to an embodiment of this application is shown. The electronic device may be...Figure 1 The terminal 110 or the server 120 in the system 100, and the like, need to be explained, Figure 8 The computer system 1200 of the electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present application.

[0196] As shown in Figure 8 The computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1202 or programs loaded from a storage portion 1208 into a random access memory (RAM) 1203, such as performing the methods in the above embodiments. In the RAM 1203, various programs and data required for system operation are also stored. The CPU 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0197] The following components are connected to the I / O interface 1205: an input portion 1206 including a keyboard, a mouse, and the like; an output portion 1207 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage portion 1208 including a hard disk, and the like; and a communication portion 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication portion 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as necessary. A removable media 1211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is mounted on the drive 1210 as necessary, so that a computer program read therefrom is installed into the storage portion 1208 as necessary.

[0198] In particular, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product including a computer program carried on a computer readable medium, the computer program containing program codes for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication portion 1209, and / or installed from the removable media 1211. When the computer program is executed by the central processing unit (CPU) 1201, various functions defined in the system of the present application are performed.

[0199] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present application, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, or the like, or any suitable combination of the above.

[0200] The flow and block diagrams in the drawings represent possible architectural, functional, and operational architectures of systems, methods, and computer program products according to various embodiments of the present application. Each block in the flow or block diagrams can represent a module, a segment, or a portion of code, which includes one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and computer-readable instructions.

[0201] The units described in the embodiments of the present application can be implemented by software, or by hardware, or by a combination of software and hardware. The units described may

[0202] As another aspect, the present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device. The computer readable storage medium carries computer readable instructions, which, when executed by a processor, implement the method in any of the above embodiments.

[0203] According to an aspect of the embodiments of the present application, a computer program product is provided, which includes computer readable instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer readable instructions from the computer readable storage medium, and the processor executes the computer readable instructions, so that the electronic device executes the method in any of the above embodiments.

[0204] In the embodiments of the present application, the term "module" or "unit" refers to a part of a computer program with a predetermined function, and works with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory), or a combination thereof, and similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a whole module or unit of the function of the module or unit, or a part of the module or unit.

[0205] It should be noted that, although several modules or units of the devices for action execution are mentioned in the foregoing detailed description, such division is not mandatory. Indeed, according to an embodiment of the application, the features and functionalities of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functionalities of one module or unit described above can be further divided into several modules or units embodied.

[0206] From the above description of the embodiments, those skilled in the art will readily appreciate that the example embodiments described herein can be implemented by software and / or by hardware. Accordingly, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash disk, a mobile hard disk, etc.) or on a network, and includes a number of instructions for making an electronic device (such as a personal computer, a server, a touch terminal, or a network device, etc.) execute the methods according to the embodiments of the present application.

[0207] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the application following, in general, the principles of the application and including such

[0208] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art will understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to part of the technical features; and these modifications or replacements do not drive the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method of policy evaluation, characterized by, The method comprises: determining a target function of a target experiment group based on first empty running index data of at least three experiment groups in a reference period; no target strategy is applied to each of the experiment groups in the reference period, and the at least three experiment groups include the target experiment group; the target function includes a weighted result of first empty running index data of a plurality of control experiment groups and an intercept for adjusting the weighted result; the plurality of control experiment groups are experiment groups other than the target experiment group in the at least three experiment groups; and a solution target of the target function is to minimize the target function by adjusting the weights of the plurality of control experiment groups and the intercept; solving the target function under a preset constraint condition to obtain a target weight corresponding to the target experiment group; the target weight includes the weights of the plurality of control experiment groups; and the constraint condition includes that the weights of each of the control experiment groups are real numbers and the intercept is not zero; fusing second empty running index data of the plurality of control experiment groups in a strategy period through the target weight to obtain control index data of the target experiment group in the strategy period; no target strategy is applied to each of the control experiment groups in the strategy period, and the strategy period is different from the reference period; evaluating the target strategy based on the control index data to obtain a strategy evaluation result of the target strategy.

2. The method of claim 1, wherein, The reference period includes reference periods under different period categories; The method for determining a target function of a target experiment group based on first empty running index data of at least three experiment groups in a reference period, solving the target function under a preset constraint condition to obtain a target weight corresponding to the target experiment group, and fusing second empty running index data of a plurality of control experiment groups in a strategy period through the target weight to obtain control index data of the target experiment group in the strategy period comprises: for each period category, determining a target function corresponding to the target experiment group under the period category based on first empty running index data of the at least three experiment groups in a reference period under the period category; solving the target function corresponding to the target experiment group under the period category under the constraint condition to obtain a target weight corresponding to the target experiment group under the period category; obtaining a period category to which the strategy period belongs as a target period category; fusing second empty running index data of the plurality of control experiment groups based on the target weight corresponding to the target experiment group under the target period category to obtain control index data of the target experiment group in the strategy period.

3. The method of claim 1, wherein, The method for determining a target function of a target experiment group based on first empty running index data of at least three experiment groups in a reference period comprises: obtaining a reference regularization term; the reference regularization term is associated with the weights of the plurality of control experiment groups; determining the target function based on the first empty running index data of the at least three experiment groups and the reference regularization term.

4. The method of claim 3, wherein, The reference regularization term is multiple; the target function includes a target function of the target experiment group for the multiple reference regularization terms; The target function is solved under the preset constraint condition to obtain the target weight corresponding to the target experiment group, comprising: For each reference regularization term, the target function of the target experiment group for the reference regularization term is solved under the constraint condition to obtain the candidate weight corresponding to the target experiment group under the reference regularization term; the candidate weight includes the weight of each of the multiple control experiment groups; The first empty running index data of the multiple control experiment groups is fused through the candidate weight corresponding to the target experiment group under the reference regularization term to obtain the candidate index data corresponding to the target experiment group under the regularization term; Based on the candidate index data corresponding to the target experiment group under the multiple reference regularization terms, the target weight corresponding to the target experiment group is determined from the candidate weight corresponding to the target experiment group under the multiple reference regularization terms.

5. The method of claim 4, wherein, The target weight corresponding to the target experiment group is determined from the candidate weight corresponding to the target experiment group under the multiple reference regularization terms based on the candidate index data corresponding to the target experiment group under the multiple reference regularization terms, comprising: Based on the difference between the candidate index data corresponding to the target experiment group under the reference regularization term and the first empty running index data, the error index corresponding to the target experiment group under the reference regularization term is determined; Based on the error index corresponding to the target experiment group under the multiple reference regularization terms, the target regularization term is determined from the multiple reference regularization terms. The candidate weight corresponding to the target experiment group under the target regularization term is obtained as the target weight corresponding to the target experiment group.

6. The method according to any one of claims 3-5, characterized in that, The reference regularization term is obtained, comprising: Based on at least one of the sum of squares of the weights of the multiple control experiment groups and the sum of absolute values of the weights of the multiple control experiment groups, the reference regularization term is determined.

7. The method of claim 1, wherein, The target strategy is evaluated based on the control index data to obtain a strategy evaluation result of the target strategy, comprising: Obtain multiple promotion amplitudes and strategy index data of the target experiment group in the strategy period; the target experiment group is applied with the target strategy in the strategy period; Adjust the first empty running index data of each experiment group under each promotion amplitude to obtain the adjusted index data of each experiment group under each promotion amplitude; For each promotion amplitude, the test efficiency evaluation result of the reference test algorithm under the promotion amplitude is determined based on the adjusted index data of the at least three experiment groups under the promotion amplitude and the first empty running index data of the at least three experiment groups by referring to the test algorithm; the reference test algorithm is any one of the multiple candidate test algorithms; Based on the test efficiency evaluation result of each of the multiple candidate test algorithms under the multiple promotion amplitudes, the target test algorithm is determined from the multiple candidate test algorithms; The target test algorithm is used to determine a strategy evaluation result of the target strategy based on the control index data and the strategy index data of the target experiment group in the strategy period.

8. The method of claim 7, wherein, The reference test algorithm is used to determine, for each of the promotion amplitudes, a test power evaluation result of the reference test algorithm at the promotion amplitude based on the difference between the adjustment index data and the first empty running index data of the at least three experiment groups at the promotion amplitude. The reference test algorithm is used to determine, for each of the promotion amplitudes, a probability of detecting a type I error of the reference test algorithm at the promotion amplitude based on the difference between the adjustment index data and the first empty running index data of the at least three experiment groups at the promotion amplitude, as a type I error rate of the reference test algorithm at the promotion amplitude. The reference test algorithm is used to determine a test power evaluation result of the reference test algorithm at the promotion amplitude based on the type I error rate of the reference test algorithm at the promotion amplitude.

9. The method of claim 8, wherein, The adjustment index data includes single-dimension adjustment index data of each of a plurality of dimensions; and the first empty running index data includes single-dimension empty running index data of each of the plurality of dimensions. The reference test algorithm is used to determine, for each of the promotion amplitudes, a probability of detecting a type I error of the reference test algorithm at the promotion amplitude based on the difference between the adjustment index data and the first empty running index data of the at least three experiment groups at the promotion amplitude, as a type I error rate of the reference test algorithm at the promotion amplitude. The reference test algorithm is used to determine, for each of the promotion amplitudes, a P value between the single-dimension adjustment index data of a reference dimension of a reference experiment group and the single-dimension empty running index data of the reference dimension of the reference experiment group at the promotion amplitude, as a P value of the reference experiment group for the reference test algorithm at the reference dimension at the promotion amplitude; the reference experiment group is any one of the at least three experiment groups, and the reference dimension is any one of the plurality of dimensions. The reference test algorithm is used to determine a type I error rate of the reference test algorithm at the promotion amplitude based on the P values of the at least three experiment groups for the reference test algorithm at the plurality of dimensions at the promotion amplitude.

10. The method of claim 8, wherein, The first empty running index data includes single-dimension empty running index data of each of the plurality of dimensions. The reference test algorithm is used to determine, for each of the promotion amplitudes, a test power evaluation result of the reference test algorithm at the promotion amplitude based on the type I error rate of the reference test algorithm at the promotion amplitude, a number of dimensions of the plurality of dimensions, and a number of experiment groups of the at least three experiment groups. The reference test algorithm is used to determine a type I error controllable test result of the reference test algorithm at the promotion amplitude based on the type I error rate of the reference test algorithm at the promotion amplitude, the number of dimensions of the plurality of dimensions, and the number of experiment groups of the at least three experiment groups; the type I error controllable test result of the reference test algorithm is used to indicate whether a type I error detected by the reference test algorithm is controllable. determine, based on the false positive rate of the reference test algorithm under the promotion amplitude and the false positive controllable test result, a test efficacy evaluation result of the reference test algorithm under the promotion amplitude.

11. The method of claim 10, wherein, The false positive controllable test result of the reference test algorithm under the promotion amplitude is determined based on the false positive rate of the reference test algorithm under the promotion amplitude, the number of dimensions of the plurality of dimensions, and the number of experimental groups of the at least three experimental groups, including: determine a controllable interval based on a preset significant threshold, the number of dimensions of the plurality of dimensions, and the number of experimental groups of the at least three experimental groups; determine the false positive controllable test result of the reference test algorithm under the promotion amplitude based on the comparison result of the false positive rate of the reference test algorithm under the promotion amplitude and the controllable interval.

12. The method of claim 7, wherein, The target test algorithm is determined from the plurality of candidate test algorithms based on the test efficacy evaluation results of the plurality of candidate test algorithms under the plurality of promotion amplitudes, including: For each candidate test algorithm, a first test algorithm is determined from the plurality of candidate test algorithms based on the number of key promotion amplitudes corresponding to the candidate test algorithm; the key promotion amplitude corresponding to the candidate test algorithm refers to the promotion amplitude for which the candidate test algorithm detects a false positive rate that is controllable. If the first test algorithm is multiple, a target test algorithm is determined from the first test algorithm based on the size of the false positive rate of the plurality of first test algorithms under the plurality of promotion amplitudes.

13. The method of claim 7, wherein, Before determining, for each promotion amplitude, the test efficacy evaluation result of the reference test algorithm under the promotion amplitude by the reference test algorithm based on the adjustment index data of the plurality of experimental groups and the first empty run index data of the plurality of experimental groups under the promotion amplitude, the method further includes: obtain a plurality of second test algorithms; different second test algorithms are obtained by configuring an initial test algorithm with different hyperparameters; For each promotion amplitude, determine the test efficacy evaluation result of the reference second test algorithm under the promotion amplitude by the reference second test algorithm based on the adjustment index data of the plurality of experimental groups and the control index data of the plurality of experimental groups under the promotion amplitude; the reference second test algorithm is any one of the plurality of second test algorithms; determine a third test algorithm from the plurality of second test algorithms based on the test efficacy evaluation results of the plurality of second test algorithms under the plurality of promotion amplitudes; If the number of third test algorithms is multiple, one of the plurality of third tests is determined as a candidate test algorithm based on the size of the hyperparameters configured by the plurality of third test algorithms.

14. A policy evaluation apparatus characterized by comprising: The device includes: The determining module is configured to determine a target function of a target experiment group based on first empty running index data of at least three experiment groups in a reference period; no target strategy is applied to each of the experiment groups in the reference period, and the at least three experiment groups include the target experiment group; the target function includes a weighted result of first empty running index data of a plurality of control experiment groups and an intercept for adjusting the weighted result; the plurality of control experiment groups are experiment groups other than the target experiment group in the at least three experiment groups; and a solution target of the target function is to minimize the target function by adjusting weights of the plurality of control experiment groups and the intercept. The solving module is configured to solve the target function under a preset constraint condition to obtain a target weight corresponding to the target experiment group; the target weight includes weights of the plurality of control experiment groups; and the constraint condition includes that the weights of each of the control experiment groups are real numbers and the intercept is not zero. The fusing module is configured to fuse second empty running index data of the plurality of control experiment groups in a strategy period by using the target weight to obtain control index data of the target experiment group in the strategy period; no target strategy is applied to each of the control experiment groups in the strategy period, and the strategy period is different from the reference period. The evaluation module is configured to evaluate the target strategy based on the control index data to obtain a strategy evaluation result of the target strategy.

15. An electronic device, comprising: The method comprises the following steps: A processor; A memory having computer readable instructions stored thereon, wherein the computer readable instructions are executed by the processor to implement the method according to any one of claims 1-13.

16. A computer readable storage medium, characterized in that, A computer readable instruction is stored thereon, when the computer readable instruction is executed by the processor, the method according to any one of claims 1-13 is implemented.

17. A computer program product, characterised in that, A computer readable instruction is stored thereon, when the computer readable instruction is executed by the processor, the method according to any one of claims 1-13 is implemented.