Test information determination method and apparatus, device, medium, and program product
By obtaining the observation values of the control group and the experimental group, and using the reference observation values in the preset test environment to predict the indicator data of the control group, the problem of the A/B test scheme being unable to be tested after full configuration of the business was solved, and the effective testing of the control group and cost reduction were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-09-07
- Publication Date
- 2026-07-24
AI Technical Summary
Existing A/B testing solutions cannot effectively test the control group after the business is fully configured, resulting in the test subjects not being able to perceive the business usage experience, increasing R&D and experimentation costs, and failing to obtain the indicator data of the control group after full configuration.
By obtaining the first, second, and third observations of the control group, and using the reference observations in the preset test environment, the indicator data of the control group after full configuration of the business are predicted. The configuration of the target business to the control group is prohibited within the effective period. The test is carried out in combination with the observation data of the experimental group.
It enables effective testing of the control group after full configuration of the business, protects the usage rights of the control group, reduces R&D and experimental costs, and can predict the real control data of the control group to complete long-term testing.
Smart Images

Figure CN117707913B_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to the field of computer technology, specifically to the field of software testing technology, and in particular to a method, apparatus, device, medium, and program product for determining test information. Background Technology
[0002] With the rapid development and progress of computer technology, all kinds of applications have emerged in the Internet environment. If these applications are configured with new versions of services, long-term testing is required to obtain the long-term impact of the new versions on the applications.
[0003] The A / B testing scheme used in related technologies involves randomly selecting two groups of test subjects as the experimental and control groups for the new version of the service. The new version of the service is released to the experimental group, while the control group is not. The test results of the new version of the service are obtained by comparing and analyzing the indicator data sampled from the experimental group and the indicator data sampled from the control group. However, after the full release of the service, the service will be configured for both the experimental and control groups. Therefore, the above testing scheme is not suitable for service testing after the service has been fully configured. Summary of the Invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a method, apparatus, device, medium and program product for determining test information, which enables the test object in the control group to perform test on the business after the full volume of business.
[0005] In a first aspect, this application provides a method for determining test information, the method comprising: acquiring a first observation value of a control group, a second observation value of the control group, and a third observation value that is time-matched with the first observation value; the first observation value is an observation value collected from the control group before configuring the target service to the control group, the second observation value is an observation value collected from the control group after configuring the target service to the control group, and the third observation value is an observation value collected from the experimental group after configuring the target service to the experimental group; the target service includes a multimedia push service, and the observation value is the value of the recommendation effect index of the multimedia push service; determining a first difference between the first observation value and the third observation value, and determining a reference observation value of the control group under a preset test environment based on the first difference and the second observation value; the effective time of the preset test environment matches the time when the target service is configured in the control group, and configuring the target service to the control group is prohibited during the effective time; the reference observation value is the value of the recommendation effect index of the multimedia push service in the control group under the preset test environment; and determining the test information of the target service based on the reference observation value.
[0006] Secondly, this application provides a test information determining device, the test information determining device comprising:
[0007] The acquisition module is used to acquire the first observation value of the control group, the second observation value of the control group, and the third observation value that is time-matched with the first observation value; the first observation value is the observation value collected from the control group before configuring the target service to the control group, the second observation value is the observation value collected from the control group after configuring the target service to the control group, and the third observation value is the observation value collected from the experimental group after configuring the target service to the experimental group; the target service includes multimedia push service, and the observation value is the value of the recommendation effect index of multimedia push service.
[0008] The processing module is used to determine the first difference between the first observation and the third observation, and to determine the reference observation of the control group under the preset test environment based on the first difference and the second observation. The effective time of the preset test environment matches the time when the target service is configured in the control group, and the configuration of the target service to the control group is prohibited during the effective time. The reference observation is the value of the recommendation effect index of the multimedia push service in the control group under the preset test environment.
[0009] The testing module is used to determine the test information for the target service based on reference observations.
[0010] In one embodiment, the target service is used to recommend multimedia push services to target objects, and the observations are used to test the recommendation effect of the multimedia push services. The target objects include at least one of the experimental group and the control group.
[0011] In one embodiment, the processing module is specifically used for,
[0012] The first difference and the first time are fitted to obtain the target correlation. The target correlation is used to characterize the relationship between the observation difference and time. The observation difference is used to characterize the degree of fitting difference between the observations of the experimental group and the observations of the control group at the same time. The first time is the same time corresponding to the first observation and the third observation.
[0013] The generation time of the second observation is mapped according to the target correlation to obtain the difference in observations corresponding to the control group.
[0014] The reference observation is determined based on the difference between the observed values and the second observation.
[0015] In one embodiment, the processing module is specifically used to perform fitting processing on the first difference and the first time based on the observability parameter to obtain the target probability density function; the target probability density function is used to characterize the target correlation; the observability parameter is related to the distribution of the first difference.
[0016] In one embodiment, the target probability density function is a probability density function of a sub-exponential distribution type, and the processing module is specifically used to...
[0017] By fitting the first difference and the first time based on the parameters, at least two sub-exponential distribution functions are obtained.
[0018] The target probability density function is the probability density function that best matches the distribution of the first difference among at least two exponential distribution functions.
[0019] In one embodiment, the processing module is specifically used for,
[0020] The initial probability density function is determined immediately based on the preset parameters.
[0021] Based on the first time and the initial probability density function, determine the prediction difference output by the initial probability density function.
[0022] Based on the loss between the first difference and the predicted difference, the preset obedience parameters are iteratively trained to obtain the obedience parameters and the target probability density function.
[0023] In one embodiment, the processing module is specifically used for,
[0024] The difference between the observed values is corrected based on the preset values to obtain the first value.
[0025] A reference observation is determined based on the first and second observations; the magnitude of the reference observation is positively correlated with the magnitude of the second observation.
[0026] In one embodiment, the target service includes at least two services, and the processing module is specifically used for,
[0027] For each service, the difference between the observed values corresponding to the service is corrected according to the preset values to obtain the first value corresponding to the service.
[0028] The first value of all business transactions is multiplied to obtain the intermediate value.
[0029] A reference observation is determined based on the median and the second observation; the magnitude of the reference observation is positively correlated with the magnitude of the second observation.
[0030] In one embodiment, the test module is specifically used for,
[0031] Obtain a fourth observation that matches the time of the second observation; the fourth observation is the observation collected from the experimental group after the target service is configured to the experimental group.
[0032] Test information is determined based on the reference observation and the fourth observation.
[0033] In one embodiment, the same time corresponding to the second and fourth observations is designated as the second time, and at least one reference time precedes the second time. The test module is specifically used for...
[0034] Obtain observations for the control group and the experimental group at at least one reference time.
[0035] Based on the weighting coefficients of at least one reference time and the weighting coefficients of the second time, the observations of at least one reference time, the second observation, and the fourth observation are smoothed to obtain test information.
[0036] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in embodiments of this application.
[0037] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in embodiments of this application.
[0038] Fifthly, embodiments of this application provide a computer program product including instructions that, when executed, cause the method described in embodiments of this application to be performed.
[0039] The method, apparatus, equipment, medium, and program products for determining test information proposed in this application address two main problems: First, if existing A / B testing schemes are used to conduct long-term testing of the business, the test subjects in the control group will be unable to perceive the user experience brought by the business for a considerable period, thus harming the right of some test subjects to use the business. Furthermore, the inability of some test subjects to use the business directly increases the development and experimental costs of the business. Second, if a full configuration of the business is chosen, it is impossible to obtain the indicator data of the control group that is unaffected by the business release after the full configuration, thus making it impossible to complete long-term testing of the business.
[0040] Therefore, this application uses the time point of full service configuration as the dividing line. By utilizing the indicator data (i.e., observed values) observed in the experimental group before full service configuration and the indicator data observed in the control group, combined with the indicator data observed in the control group after full service configuration, it predicts the indicator data of the control group when it does not configure the service after the aforementioned time point, thereby achieving the testing of the service. Specifically, firstly, the first observation value and the second observation value of the control group are obtained; where the first observation value is the observation value collected from the control group before configuring the target service, and the second observation value is the observation value collected from the control group after configuring the target service; and a third observation value matching the time of the first observation value is obtained; the third observation value is the observation value collected from the experimental group after configuring the target service, and the target service includes multimedia push service, and the observation value is the value of the recommendation effect indicator of multimedia push service. Then, the first difference between the first observation value and the third observation value is determined. Through the first difference, the actual difference between the observation values of the two groups before service configuration can be obtained. Based on the actual difference and the data observed in the control group after configuring the target service (i.e., the second observation value), a reference observation value for the control group in a preset test environment is determined. The effective time of the preset test environment matches the time the control group configures the target service, and configuring the target service to the control group is prohibited during the effective time. The reference observation value is the recommendation performance index of the multimedia push service in the preset test environment. The test information for the target service is then determined based on this reference observation value. This approach ensures full configuration of the service, allowing the control group to use it and protecting their right to use the service, while also reducing the development and experimental costs of the service. Furthermore, after full configuration of the service, the actual control data of the control group under the assumption that the service is not configured can be predicted, enabling testing of the service based on the experimental data of the experimental group and the actual control data, thus achieving the goal of long-term testing of the service.
[0041] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0042] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0043] Figure 1 A schematic diagram of the structure of the test information determination system provided in the embodiments of this application;
[0044] Figure 2 This is an experimental deployment architecture diagram provided for an embodiment of this application;
[0045] Figure 3 A flowchart illustrating the method for determining test information provided in an embodiment of this application;
[0046] Figure 4 This application provides a diagram of a reverse experiment deployment architecture.
[0047] Figure 5 Another experimental deployment architecture diagram provided for embodiments of this application;
[0048] Figure 6 This is another experimental deployment architecture diagram provided in the embodiments of this application;
[0049] Figure 7 Gamma distribution diagram provided for embodiments of this application;
[0050] Figure 8 The log-normal distribution diagram provided in the embodiments of this application;
[0051] Figure 9 This is a schematic diagram illustrating the effect of determining the test information provided in the embodiments of this application;
[0052] Figure 10 This is a schematic diagram illustrating the effect of determining multiple test information provided in the embodiments of this application;
[0053] Figure 11 A schematic diagram of the structure of the device for determining test information provided in the embodiments of this application;
[0054] Figure 12 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0055] The present application will now be described in optional detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0056] It is understood that the term "multiple" as used in this application refers to "two" and "more than two".
[0057] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0058] The following explains the terms used in the embodiments of this application:
[0059] 1. Layered domain design (experiment infrastructure design)
[0060] Layered design is essentially the division of experimental traffic. By properly configuring the subordinate structure of traffic domains and traffic layers, the orthogonality, scalability, reusability, and high availability of experimental traffic can be ensured.
[0061] 2. Crossover experiment
[0062] The reverse experimental design is essentially a repeated measurements design, in which the experimental and control groups are swapped after the experimental period ends. The original experimental group is used as the new control group and the original control group is used as the new experimental group to repeat the experiment. This reduces the impact of confounding covariates on the experimental effect and improves the efficiency of testing and the accuracy of conclusions.
[0063] 3. Holdout experiment
[0064] Holdout testing, also known as long-term holdout testing, is a type of A / B testing strategy. It involves partially implementing the new service while keeping a portion of the business objects offline for an extended period. These objects serve as a control group. The metrics generated by this control group are then compared with the metrics of other groups that have implemented the new service to assess its long-term impact. For example, during full rollout, 99% of the business objects are made available with the new service, while 1% are kept offline. The 1% group acts as the control group, and the 99% group as the experimental group. The resulting metrics are then compared to determine the overall impact of the new service on the business.
[0065] 4. Long-term effects
[0066] Long-term effects refer to the impact of a business on another business over a long period of time.
[0067] 5. Sub-exponential distribution
[0068] The sub-exponential distribution can be seen as an extension of the sub-Gaussian distribution, representing a distribution pattern of random variables that satisfies certain specific properties. Common sub-exponential distributions include the Gamma distribution, the log-normal distribution, and the Pareto distribution.
[0069] 6. Optimization methods
[0070] Optimization methods refer to methods for finding the extrema of a specific function under given constraints. Common optimization methods include linear search, gradient descent, Newton's method, and genetic algorithms.
[0071] 7. Decay function
[0072] A decay function is a mathematical expression that describes the decay of a dependent variable over time. It is generally required that the function value converges to 0 as time increases.
[0073] 8. Smoothing
[0074] Smoothing is a data processing technique that removes noise from data, thereby helping to capture the main characteristics of data trends. It's also used in parameter estimation to address the problem of data sparsity. The main idea is to allocate a portion of the probability density in the entire probability space according to certain business rules to low-frequency sparse events, making the estimated probability distribution more reliable under sparse conditions.
[0075] The A / B testing scheme used in related technologies involves randomly selecting two groups of test objects as the experimental and control groups for the new service. The test results are obtained by comparing and analyzing the observed metrics data from the experimental and control groups. However, if this testing scheme is used, the new service will be configured in both the experimental and control groups after full configuration. Therefore, the above testing scheme is not suitable for testing the service after full configuration.
[0076] Based on this, embodiments of this application provide a method, apparatus, device, medium, and program product for determining test information. Taking the time node of full service configuration as the dividing line, by using the indicator data (i.e., observed values) observed in the experimental group before the full service configuration and the indicator data observed in the control group, combined with the indicator data observed in the control group after the full service configuration, the indicator data of the control group not configuring the service after the aforementioned time node is predicted, thereby realizing the testing work of the service.
[0077] Figure 1 This is a schematic diagram of a test information determination system provided in an embodiment of this application. The test information determination method provided in this embodiment can be applied to the test information determination system 100. (Reference) Figure 1 The system 100 includes at least two clients 101 (e.g., Figure 1 The diagram shows clients 101a, 101b, and 101c, and server 102. It should be noted that, although... Figure 1Only three clients 101a, 101b, and 101c are depicted, but those skilled in the art will understand that this application can support any number of clients, at least two or more.
[0078] Understandably, at least two clients 101 can provide at least two test objects, which can be divided into experimental and control groups. The server 102 obtains the service's metric data (i.e., observations) by acquiring the usage of the service by the test objects in at least two clients 101.
[0079] Optionally, client 101 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, etc. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. The computer devices are capable of executing various applications, such as various Internet-related applications, communication applications (such as email applications), short message service applications, and can use various communication protocols.
[0080] Optionally, refer to Figure 2 This paper presents an experimental deployment architecture diagram. This architecture employs a layered domain design, involving business experimental domains for conducting experiments. These business experimental domains include long-term benefit verification domains and multi-layered experimental domains. The multi-layered experimental domains include an experimental release layer and multiple experimental layers (such as...). Figure 2 The experiment layer is shown as Experiment Layer 1, Experiment Layer 2, ..., Experiment Layer n. The experiment release layer mainly allocates and configures services to each experiment layer, allocating and configuring one or more services that are the same and / or similar to each experiment layer. Each experiment layer can include one or more experiments, and each experiment includes a control group and an experimental group. The control group and experimental group included in each experiment are used to test a certain service. The long-term benefit verification domain includes a long-term experimental group and a long-term control group. Both A / B testing and reversal experiments are carried out in the experiment layer of the multi-layered experiment domain. Service configuration can be performed in the experiment release layer, including full configuration and partial configuration. Full-scale experiments after full configuration can be carried out in the long-term experimental group or in the experiment layer. Long-term holdout experiments are carried out in the long-term benefit verification domain.
[0081] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0082] The following will combine Figure 1 and Figure 2 The technical solutions of this application and how they solve the aforementioned technical problems will be described in detail with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0083] like Figure 3 As shown in the illustration, this application provides a method for pushing media content, specifically applied in the field of content recommendation. This method can be applied to... Figure 1 The method for server 102 shown specifically includes the following steps:
[0084] S11. Obtain the first observation of the control group, the second observation of the control group, and the third observation that is time-matched with the first observation.
[0085] The first observation value is the observation value collected from the control group before configuring the target service to the control group; the second observation value is the observation value collected from the control group after configuring the target service to the control group; and the third observation value is the observation value collected from the experimental group after configuring the target service to the experimental group. The target service includes multimedia push service, and the observation value is the value of the recommendation effect index of multimedia push service.
[0086] Optionally, the first observation value is the true value of the recommendation performance index of the multimedia push service in the control group before the target service is configured; the second observation value is the true value of the recommendation performance index of the multimedia push service in the control group after the target service is configured; and the third observation value is the true value of the recommendation performance index of the multimedia push service in the experimental group after the target service is configured, which is time-matched with the first observation value. Here, the true value can be understood as the value actually collected.
[0087] It should be noted that time matching refers to the same time; for example, the same time period or the same point in time.
[0088] Specifically, the first, second, and third observations are all observations of the same observation indicator, which can be used to test (or evaluate) the performance of the target service and can be a key performance indicator (KPI) for testing the target service. The first and second observations are the observations of this observation indicator collected or detected by the control group at different time periods. The first and third observations are the observations of this observation indicator collected or detected from the control group and the experimental group, respectively, within the same time period. Specifically, the configuration time of the target service is used as the time node. The first observation is the observation collected or detected by the control group before this time node, the second observation is the observation collected or detected by the control group after this time node, and the third observation is the observation collected or detected by the experimental group before this time node.
[0089] Optionally, the control group is set as the first time period before configuring the target service and as the second time period after configuring the target service; both the first time period and the second time period contain one or more preset periods, wherein the number of preset periods contained in the first time period and the second time period may be the same or different, and this application embodiment does not limit this.
[0090] Specifically, the first observation value is the observation value collected or detected by the control group in each observation period within one or more observation cycles (i.e., the first time period) included in the aforementioned first time period; the second observation value is the observation value collected or detected by the control group in each observation period within one or more observation cycles included in the aforementioned second time period; and the third observation value is the observation value collected or detected in each observation period within one or more observation cycles included in the aforementioned first time period. The first observation value can be the average, maximum, or minimum value of multiple values of the observed indicator within an observation period, or it can be the value of the observed indicator at a preset time within that observation period; the second and third observation values are similar and will not be elaborated upon here.
[0091] Optionally, the control group includes at least one test subject (i.e., a control group). The control subjects included in the control group during the first time period can be the same as those included in the second time period, or they can be different. Similarly, the experimental group includes at least one test subject (i.e., an experimental subject). The experimental subjects included in the experimental group during the first time period can be the same as those included in the second time period, or they can be different.
[0092] In one exemplary implementation, the target service can be the specific content to be tested, such as the layout of the interface, the shape or color of interactive controls, the object operation method (click or swipe), the process, or the push of advertisements. In some instances, the target service may also correspond to its service information, which may be an identifier of the service. It should be understood that the service information may also include a description of the service function, application scenarios, and other information.
[0093] For example, taking the target business as changing the color of a target control from red to blue, assume there are 100 test subjects in the first time period, including 50 experimental subjects and 50 control subjects. For the target control being blue, this blue control is pushed to the 50 experimental subjects in the first time period for their use, obtaining the first observation value. For the target control being red, the red control is pushed to the other 50 control subjects in the first time period for their use, obtaining the third observation value. If the same 100 test subjects are used in the second time period, and the target control is blue, the blue control is pushed to the 50 control subjects in the second time period for their use, obtaining the second observation value.
[0094] Optionally, assuming that the 100 test subjects from the first time period are not used in the second time period, but another 200 test subjects are used, including 100 experimental subjects and 100 control subjects, and the target control is blue, this blue control is pushed to the other 100 test subjects in the second time period for the test subjects to use, so as to obtain the second observation value.
[0095] In this embodiment, when the target service is not configured for the test object, it can be understood that the test object uses a service other than the target service (assuming it is the first service). The first service and the target service are different services for the same project within the service. The target service can be a new service within the service compared to the first service, while the first service can be an existing service within the service. As an example, the experimental group's test objects use the target service in the first time period. The experimental group's clients can generate indicator data (such as the second observation value) under the target service based on the experimental object's operation data for the target service, and then the experimental group's clients can send the generated indicator data to the server. Similarly, the control group's clients do not configure the target service in the first time period (i.e., use the first service). The control group's clients can generate indicator data (such as the first observation value) under the first service based on the control object's operation data for the first service, and then the control group's clients can send the generated indicator data to the server.
[0096] To test the target service, multiple clients are selected by the server as experimental and control groups. In other words, multiple test objects are selected by the server as experimental and control objects. It is worth noting that there are multiple ways for the server to select the experimental and control groups in this embodiment of the application. It can be a random selection method or a selection method according to a selection rule. An exemplary selection rule can be described as follows: test objects with even-numbered client identifiers are selected as experimental objects in the experimental group, and test objects with odd-numbered client identifiers are selected as test objects in the control group. Any of the above selection methods can be used in the process of selecting the experimental and control groups, and this embodiment of the application does not limit this.
[0097] For example, the test subjects include object A1, object A2, object A3, and object A4. The server can assign object A1 and object A3 to the control group set of control objects in the first time period, and assign object A2 and object A4 to the experimental group set of experimental objects. That is, object A2 and object A4 are control objects in the first time period and the second time period, and object A2 and object A3 are experimental objects in the first time period and the second time period.
[0098] For example, the test subjects include objects A1, A2, A3, A4, A5, A6, A7, and A8. The server can assign objects A1 and A4 to the control group (set of control subjects in the first time period), and objects A2 and A3 to the experimental group (set of experimental subjects in the first time period). That is, objects A1 and A4 are control subjects in the first time period, and objects A2 and A3 are experimental subjects in the first time period. The server can also assign objects A5 and A6 to the control group (set of control subjects in the second time period), and objects A7 and A8 to the experimental group (set of experimental subjects in the first time period). That is, objects A5 and A6 are control subjects in the second time period, and objects A7 and A8 are experimental subjects in the second time period. This is just an example; in actual implementation, the number of test subjects in each time period is generally much greater than four.
[0099] As an example, the first and third observations can be obtained using A / B testing. When conducting A / B testing, one or more experimental layers can be established. Each experimental layer can contain two or more experimental sets, each set can conduct experiments on a specific business function, and each set can include a control group and an experimental group. For the experiment on the target business function, a control group and an experimental group corresponding to the original experiment of the target business function can be established in a pre-defined experimental layer. Optionally, based on the experiment on the target business function conducted in the control group and experimental group, a reversal experiment can also be established to verify the effect of the original experiment. The reversal experiment also corresponds to a reversal control group and a reversal experimental group.
[0100] For example, refer to Figure 4 This application provides an A / B testing architecture diagram. The original control group experiment is performed in bucket 1 of experimental layer 1, and the original experimental group experiment is performed in bucket 2 of experimental layer 1. To verify the experimental effects of the original experimental group and the original control group within the first time period, a reversal experiment can also be established. Specifically, the reversal experiment can be established based on the following three scenarios:
[0101] Scenario 1: When the flow rate in experimental layer 1 is sufficient, a new inversion experiment can be created while retaining the original experimental group and the original control group (i.e., retaining the original experiment). For example, Figure 4 The experiment shown is a reverse control group experiment conducted in bucket 3 of experimental layer 1, and an experimental group experiment conducted in bucket 4 of experimental layer 1.
[0102] Scenario 2: When the throughput in Experiment Layer 1 is insufficient, and the original experiment is not coupled with other experiments in Experiment Layer 1, a new inversion experiment layer can be created in the multi-domain experiment for inversion experiments. For example, Figure 5 The experiment shown is a reverse control group experiment conducted in bucket 1 of the reverse experimental layer, and a reverse experimental group experiment conducted in bucket 2 of the reverse experimental layer.
[0103] Scenario 3: When there is insufficient bandwidth in Experimental Layer 1, and the original experiment is coupled with other experiments in Experimental Layer 1, the original experimental group and control group are not retained. Instead, the original experimental group and control group are taken offline (i.e., the original experiment is taken offline), bandwidth is released, and the original experimental group is used as the reverse control group for the experiment, while the original control group is used as the reverse experimental group. For example, refer to... Figure 6 As shown, the reverse experimental group experiment was conducted in bucket 1 of experimental layer 1, and the reverse control group experiment was conducted in bucket 2 of experimental layer 1.
[0104] If both the original experiment and the reverse experiment meet expectations, the observed value of the original control group can be used as the first observed value, and the observed value of the original experimental group can be used as the third observed value; or, the observed value of the reverse control group can be used as the first observed value, and the observed value of the reverse experimental group can be used as the third observed value. It should be noted that the specific selection of the observed value of the original control group or the reverse control group as the first observed value, or a combination of both (e.g., a weighted sum of the two) to determine the first observed value, is not limited in this embodiment of the application. The same applies to the third observed value, which will not be elaborated here.
[0105] According to some embodiments of this application, an application programming interface (API) can be pre-configured. When conducting A / B testing, the API can be used to obtain data related to the target business for each experimental object and each test object.
[0106] It should be noted that the test subjects mentioned in the embodiments of this application may be, for example, users, and the data related to the test subjects involved in the embodiments of this application (such as the observed values of test indicators) are all data collected with the consent and authorization of the test subjects.
[0107] S12. Determine the first difference between the first and third observations, and determine the reference observation of the control group under the preset test environment based on the first difference and the second observation.
[0108] The preset test environment's effective time matches the time the control group configures the target service, and configuring the target service in the control group is prohibited during the effective time. The reference observation value is the recommendation performance index of the multimedia push service in the control group under the preset test environment. When the target service is a multimedia push service, the preset test environment can be understood as a test environment that prohibits configuring the multimedia push service in the control group, and the effective time of this test environment matches the time the control group configures the multimedia push service.
[0109] Specifically, the reference observation value is the fitted value of the recommendation performance index of the multimedia push service in the control group under a preset test environment. At the same time point, this fitted value is infinitely close to the true value of the recommendation performance index of the multimedia push service under the preset test environment.
[0110] Optionally, the multimedia push service can be a project within a client application, such as a browser, chat software, or novel application.
[0111] Optionally, the recommendation performance metrics of multimedia push services can characterize the target audience's interest in multimedia push services, or the revenue brought to the services applied by the multimedia push services by the conversion behavior of the recommended multimedia push services.
[0112] Optionally, conversion behavior is used to characterize the operation behavior of the target object on the multimedia push service; for example, conversion behavior can be one or more of the following behaviors: exposure behavior, click behavior, shallow conversion behavior, and deep conversion behavior.
[0113] Specifically, exposure behavior refers to the target audience's viewing behavior of media content; click behavior refers to the target audience's clicking action on the exposed multimedia push service; shallow conversion behavior refers to the target audience's action of clicking on items or links contained in the multimedia push service, being redirected to the corresponding page, and performing corresponding actions, such as application download, application activation, and form registration; deep conversion behavior refers to the target audience's action of clicking on items or links contained in the media content, being redirected to the corresponding page, and performing corresponding actions, such as paying, downloading the application, and achieving application retention the next day. Among these, the deep conversion operation of the media content is a follow-up operation to the shallow conversion operation of the media content.
[0114] Optionally, the first difference is used to characterize the true relative difference or the degree of true difference between the observed indicators of the experimental group and the observed indicators of the control group at the same time before the target service was configured (i.e., the first indicator difference). In the content recommendation scenario, the first difference is used to characterize the degree of difference between the true value of the recommendation effect indicator of the multimedia push service of the control group and the true value of the recommendation effect indicator of the multimedia push service of the experimental group at a certain period of time before the target service was configured for the control group (i.e., the degree of true difference).
[0115] The first difference can be a value 'a' obtained by subtracting the first and third observations, or a value 'b' obtained by quotienting the first observation and 'a', with 'b' serving as the first difference. In the embodiment of step S11 above, the first difference is used to characterize the true relative difference between the observed indicators in the experimental group and the observed indicators in the control group at a certain time within the first time period. Specifically, the first difference can be determined according to the following formula:
[0116]
[0117] Among them, a k This represents the first difference of the k-th service (i.e., the target service) at time t. This represents the observation value of the k-th service in the control group at time t (e.g., the t-th time period), which can be understood here as the second observation value. This represents the observation value of the k-th service in the experimental group at time t, which can be understood as the first observation value.
[0118] In practical applications, the preset test environment is actually a hypothetical test environment. Specifically, it is a test environment in which the target service is not configured for the observation group during the time period of the target service configuration.
[0119] In one feasible implementation, reference observations are used to characterize observations that are unrelated to the deployment of the target service after configuring the target service to the control group; or, to match the timing of configuring the target service to the control group, they are observations that assume the target service is not configured to the control group; or, they are observations of the control group that are unrelated to the deployment of the target service. It is understood that reference observations are not the actual observations collected or detected by the control group, but rather predicted observations of the control group that are unaffected by the deployment of the target service, or, assuming the service is not configured to the control group, possible observations of the control group.
[0120] In addition, matching the effective time of the preset test environment with the time when the control group configures the target service can be understood in practical applications as the same as the effective time of the preset test environment, which is how long the control group has configured the target service.
[0121] It should be noted that the reference observation value is actually the observed value of the observation index in the control group during the second time period, assuming that the target service is still not configured for the control group. In conjunction with the embodiment in step S11 above, the reference observation value can be understood as the observed value of the observation index generated by the control group's control object under the first service during the second time period.
[0122] Optionally, the observed values can be used to reflect the operational attributes of test subjects in the experimental or control groups for different or the same business at different time periods. In other words, the second observed value of the experimental group under the observed indicator can be used to reflect the operational attributes of the experimental subjects in the experimental group for the target business in the first time period, the first observed value of the control group under the observed indicator can be used to reflect the operational attributes of the control subjects in the control group for the first business in the first time period, and the third observed value of the control group under the observed indicator can be used to reflect the operational attributes of the control subjects in the control group for the target business in the second time period. Similarly, the reference observed value of the control group under the observed indicator can reflect the operational attributes of the control subjects in the control group for the first business in the second time period.
[0123] It should be noted that relative difference, degree of difference, extent of difference, and difference in the embodiments of this application all have the same meaning.
[0124] S13. Determine the test information for the target service based on the reference observations.
[0125] Among them, the test information of the target service is the test information after the observation group configures the target service.
[0126] In one implementation, the test information can be the target observation value of the control group without the target service configured. This target observation value corresponds to the same time as the reference observation value. The target observation value can be understood as the result of correcting the reference observation value; specifically, it is the corrected value of the recommendation performance index of the multimedia push service under the preset test conditions. Therefore, determining the test information for the target service based on the reference observation value includes: correcting the reference observation value to obtain the target observation value. Specifically, this can be achieved by correcting the reference observation value using one or more observation values before and after the time corresponding to the reference observation value, thereby obtaining the target observation value.
[0127] In another implementation, the test information can be the second difference between the experimental group and the control group when the control group is not configured with the target service. The second difference is used to characterize the relative difference between the target observation and the fourth observation (i.e., the second index difference) after the target service is configured in the control group. The second difference can be calculated using the formula for the first difference mentioned above, which will not be elaborated here. The fourth observation is the observation in the experimental group that matches the second observation in time.
[0128] In another implementation, the test information can be the test results used to characterize the target service when the control group is not configured with the target service; for example, it can be the decay rate of the observed indicators over a certain period of time after the target service is configured in the control group.
[0129] For example, suppose the multimedia push service is an advertisement in a browser, and its recommendation performance metric is the revenue generated by the target audience's click behavior on the advertisement. The first observation is the revenue value obtained by the browser of the control group before the advertisement is configured (i.e., within the first time period); the second observation is the revenue value obtained by the browser of the control group after the advertisement is configured (i.e., within the second time period) and the control group clicks on the advertisement (i.e., within the second time period); the third observation is the revenue value obtained by the browser of the experimental group after the advertisement is configured (i.e., within the second time period) and the experimental group clicks on the advertisement (i.e., within the second time period). Then, the first difference between the first revenue value and the second revenue value is determined, and based on the first difference and the second revenue value, the reference revenue value obtained by the browser of the control group in the second time period without the advertisement is determined (i.e., the reference observation value). Then, based on the reference revenue value, the target revenue value of the browser corresponding to the advertisement in the second time period is determined (i.e., the target observation value). This target revenue value is only one way of representing the test information of the target service.
[0130] The method for determining test information proposed in this application addresses two main problems with existing A / B testing schemes. First, using existing A / B testing schemes for long-term testing of a business results in a prolonged period where test subjects in the control group cannot perceive the user experience provided by the business, thus infringing on their right to use the business. Furthermore, the inability of some test subjects to use the business directly increases the development and experimental costs. Second, if a full configuration of the business is chosen, it is impossible to obtain the indicator data of the control group unaffected by the business release after full configuration, making it impossible to complete long-term testing of the business.
[0131] Therefore, this application uses the time point of full service configuration as the dividing line. By utilizing the indicator data (i.e., observed values) observed in the experimental group before full service configuration and the indicator data observed in the control group, combined with the indicator data observed in the control group after full service configuration, it predicts the indicator data of the control group when it does not configure the service after the aforementioned time point, thereby achieving the testing of the service. Specifically, firstly, the first observation value and the second observation value of the control group are obtained; where the first observation value is the observation value collected from the control group before configuring the target service, and the second observation value is the observation value collected from the control group after configuring the target service; and a third observation value matching the time of the first observation value is obtained; the third observation value is the observation value collected from the experimental group after configuring the target service, and the target service includes multimedia push service, and the observation value is the value of the recommendation effect indicator of multimedia push service. Then, the first difference between the first observation value and the third observation value is determined. Through the first difference, the actual difference between the observation values of the two groups before service configuration can be obtained. Based on the actual difference and the data observed in the control group after configuring the target service (i.e., the second observation value), a reference observation value for the control group in a preset test environment is determined. The effective time of the preset test environment matches the time the control group configures the target service, and configuring the target service to the control group is prohibited during the effective time. The reference observation value is the recommendation performance index of the multimedia push service in the preset test environment. The test information for the target service is then determined based on this reference observation value. This approach ensures full configuration of the service, allowing the control group to use it and protecting their right to use the service, while also reducing the development and experimental costs of the service. Furthermore, after full configuration of the service, the actual control data of the control group under the assumption that the service is not configured can be predicted, enabling testing of the service based on the experimental data of the experimental group and the actual control data, thus achieving the goal of long-term testing of the service.
[0132] In one embodiment of this application, a correlation can be determined based on a first difference and a first time. This correlation characterizes the relationship between the difference between the observed values of the control group and the experimental group and time, assuming the target service is not configured for the control group, after the target service is configured in the control group. A reference observed value is then determined based on this difference. Therefore, in one implementation, determining the reference observed value of the control group after configuring the target service based on the first difference and the second observed value includes: performing fitting processing on the first difference and the first time to obtain a target correlation; the target correlation characterizes the relationship between the observed value difference and time, and the observed value difference characterizes the degree of fitting difference between the observed values of the experimental group and the observed values of the control group at the same time; the first time is the same time corresponding to the first observed value and the third observed value; mapping the generation time of the second observed value according to the target correlation to obtain the observed value difference corresponding to the control group; and determining the reference observed value based on the observed value difference and the second observed value.
[0133] Optionally, the first time can be the time when the observation is generated. For example, in a sampling scenario, the first time can be the time when the observation is collected; similarly, in a detection scenario, the first time can be the time when the observation is detected. Furthermore, the first time can be a preset period or a specific point in time.
[0134] Specifically, the observation difference is used to characterize the relative difference between the observed values of the experimental group and the observed values of the control group at the same time. Optionally, in conjunction with the embodiment in step S11 above, when the observation difference corresponds to the first time period, the observation difference can specifically characterize the degree of fitting difference between the first observed value and the third observed value at the same time; when the observation difference corresponds to the second time period, the observation difference can specifically characterize the degree of fitting difference between the fourth observed value and the reference observed value.
[0135] Specifically, since the target correlation represents the relationship between the observation difference and time, an observation difference can be determined based on a time. In other words, there is actually a mapping relationship between the observation difference and time. By mapping time through this target correlation, the corresponding observation difference can be obtained.
[0136] In one implementation, before configuring the target service in the control group, there is at least one first time point, each corresponding to a first difference. A curve of a preset distribution type, similar to the distribution of the at least one first difference, can be fitted based on the first difference corresponding to each of the at least one first time point, according to a preset distribution type. The horizontal axis of this curve represents time, and the vertical axis represents the observation difference. Then, based on the generation time of the second observation, the observation difference corresponding to that generation time in the curve can be determined, and a reference observation value can be determined based on the observation difference and the second observation. It is understood that this curve conforms to the preset distribution type; therefore, this curve can not only characterize the target correlation but also the correlation between the observation difference and the generation time of the fourth observation.
[0137] Optionally, in the content recommendation scenario, the first difference is used to characterize the true difference in the recommendation performance index of the multimedia push service before configuring the target service to the control group. That is, the degree of difference between the true value of the recommendation performance index of the multimedia push service of the control group and the true value of the recommendation performance index of the multimedia push service of the experimental group during a period of time before configuring the target service to the control group. Therefore, the first difference is the true difference in the recommendation performance index of the multimedia push service before configuring the target service to the control group.
[0138] The observation difference is used to characterize the degree of fit difference of the recommendation effect index of the multimedia push service. That is, after configuring the target service to the control group, the degree of difference between the fitted value (i.e., the reference observed value) of the recommendation effect index of the multimedia push service in the control group and the true value (i.e., the degree of fit difference) of the recommendation effect index of the multimedia push service in the experimental group. Therefore, the observation difference is the fitting difference value of the recommendation effect index of the multimedia push service.
[0139] Optionally, the target association is used to characterize the relationship between the degree of fit difference of the recommendation effect index of the multimedia push service and time under a preset test environment.
[0140] For example, by fitting multiple first differences within a first time period, the correlation between the observed value differences and the first time (i.e., the time points within the first time period) can be obtained, which is the target correlation mentioned above. This correlation can be represented by a fitted curve, where the observed value difference is the fitted value in the fitted curve corresponding to each time point within the first time period.
[0141] It is understandable that, since the observation difference is used to characterize the degree of fit difference between the reference observation and the second observation, the reference observation can be calculated in reverse using the observation difference and the second observation. The specific calculation formula can be found in the calculation formula of the first difference mentioned above, and will not be repeated here.
[0142] For example, assuming the first difference for each day in the 7 days before the target service is configured in the control group, there are 7 corresponding first differences: day 1, day 2, day 3, day 4, day 5, day 6, and day 7. Specifically, based on the first differences of these 7 days according to a preset distribution type, a curve of that distribution type can be fitted. Based on this curve, the difference of the observed value for any day after day 8 can be obtained.
[0143] In this embodiment, by fitting the first difference and the same time (i.e., the first time) corresponding to the first and third observations, the relationship (i.e., the target correlation relationship) between the difference between the observations of the experimental group and the observations of the control group at the same time (i.e., the observation difference) and the first time is determined. Furthermore, the generation time of the second observation can be mapped based on the target correlation relationship to determine the observation difference corresponding to the control group. This allows for the prediction of the actual control data (i.e., the reference observation value) of the control group under the assumption of no service configuration after full service configuration, based on the observation difference and the second observation. This enables the testing of the service based on the experimental data of the experimental group and the actual control data.
[0144] In one embodiment of this application, fitting processing is performed on the first difference and the first time to obtain the target association relationship, including: fitting processing of the first difference and the first time based on the observability parameter to obtain the target probability density function; the target probability density function is used to characterize the target association relationship; the observability parameter is related to the distribution of the first difference.
[0145] Since the target correlation characterizes the relationship between the difference in observed values and time, it can be understood that the value of the target probability density function at a certain time point is the difference in observed values corresponding to that time point.
[0146] In content recommendation scenarios, the target probability density function is used to characterize the relationship between the fitted difference (i.e., the difference in observed values) of the recommendation performance index of multimedia push services and time under a preset test environment. The value of the target probability density function at a certain time point is used to characterize the degree of difference (i.e., the degree of fitted difference) between the fitted value (i.e., the reference observed value) of the recommendation performance index of the control group's multimedia push services and the true value (i.e., the fourth observed value) of the experimental group's multimedia push services at that time point.
[0147] Specifically, based on the first difference, the first time, and the parameters, a curve can be fitted, where the vertical axis of the curve represents the target density function and the horizontal axis represents the first time.
[0148] Optionally, to improve the accuracy of the curve's vertical axis, making it closer to the first difference at the same time interval (i.e., to better fit the difference curve formed by the first difference in the first time interval), a correction coefficient can be used to modify the curve, resulting in a modified curve. The vertical axis of this modified curve represents the target density function, and the horizontal axis represents the first time interval.
[0149] In this case, determining the target probability density function based on the first difference, the first time, and the following parameters specifically involves determining the initial probability density function based on the first difference, the first time, and the following parameters, and obtaining the target probability density function based on the correction coefficient and the initial probability density function.
[0150] Optionally, the average of multiple first differences can be calculated, as well as the average of multiple initial probability density functions at the same time as the multiple first differences, and the correction coefficient can be obtained based on the relative difference between the average of the first differences and the average of the initial probability density functions.
[0151] In this embodiment, the first difference and the first time can be fitted by the parameters related to the distribution of the first difference, so as to determine the target probability density function used to characterize the target association relationship.
[0152] In one embodiment of this application, considering that in practical applications, the observed indicators of most services tend to gradually decay over time during use, the target probability density function can specifically be a probability density function of a sub-exponential distribution type to ensure its practicality. However, the distribution of the first difference may differ among different services, making different sub-exponential distribution types applicable. This results in the inability to accurately describe all services using the same sub-exponential distribution type when using the probability density function to characterize the relationship between the observed difference and the first time. Therefore, different sub-exponential distribution types can be used to obtain different sub-exponential distribution functions. That is, based on the first difference, the first time, and the following parameters, multiple probability density functions of sub-exponential distribution types (i.e., sub-exponential distribution functions) are determined, and the optimal probability density function is selected as the target probability density function. Specifically, the target probability density function is obtained by fitting the first difference and the first time based on the following parameters, including: fitting the first difference and the first time based on the following parameters to obtain at least two sub-exponential distribution functions; and determining the probability density function of the sub-exponential distribution type that is closest to the distribution of the first difference among the at least two sub-exponential distribution functions as the target probability density function.
[0153] It is understandable that different sub-exponential distribution types have different parameters.
[0154] In the content recommendation scenario, the target probability density function is obtained by fitting the first difference and the first time based on the obedience parameter. Specifically, this includes fitting the true difference value (i.e., the first difference) and the first time of the recommendation effect index of the multimedia push service before configuring the target service to the control group based on the obedience parameter, and obtaining at least two exponential distribution functions. Each exponential distribution function is used to characterize the relationship between the fitted difference value (i.e., the difference of observed value) of the multimedia push effect index and time.
[0155] Optionally, the sub-exponential distribution type may include one or more of the following distribution types: Gamma distribution, log-normal distribution, Pareto distribution, etc. The target probability density function is one of the aforementioned sub-exponential distribution types. It should be noted that the parameters are determined by the sub-exponential distribution type. For example, if the sub-exponential distribution type is Gamma distribution, the parameters include shape parameters and scaling parameters; if the sub-exponential distribution type is log-normal distribution, the parameters include mean and variance.
[0156] As an example, since observed metrics tend to decay over time, for experiments with returns within a reasonable range, we can assume that the changes in observed values over time have a sub-exponential distribution shape, for example, referring to... Figure 7 The shape shown follows a Gamma distribution, or as... Figure 8 The shape shown follows a log-normal distribution.
[0157] Alternatively, for the Gamma distribution, the following formula applies:
[0158]
[0159] Where x represents a random variable, Γ(α) represents the Gamma function, α represents the shape parameter, and β represents the scaling parameter. Both α and β are conformational parameters.
[0160] For the log-normal distribution, the corresponding formula is as follows:
[0161]
[0162] Where x represents a random variable, Γ(α) represents the Gamma function, α represents the mean of the logarithm of the random variable x, and β represents the variance of the logarithm of the random variable x. Both α and β are parameters. Figure 7 The σ = β shown in the figure 2 .
[0163] It should be noted that, since the area under the curve of the sub-exponential distribution is required to be 1, in order to remove this requirement from the limitations of this application's embodiments and thus achieve the purpose of long-term testing of the business, the formula for the sub-exponential distribution can be modified accordingly to obtain a core function that can describe the distribution shape. Taking the Gamma distribution as an example, by removing the denominator from the formula corresponding to the Gamma distribution above and replacing the arbitrary variable x with time t, the following formula is obtained:
[0164] g s (t)=β α t α-1 e -βt ,t>0
[0165]
[0166] Among them, g s (t) represents the core function of the sub-exponential distribution s, which can characterize any one of the prediction difference, the initial probability density function, and the target probability density function. Let represent the probability density function obtained from the prediction under the sub-exponential distribution s, where s represents the sub-exponential distribution type.
[0167] In one possible implementation, for each sub-exponential distribution, the average of multiple first differences and the average of multiple initial probability density functions occurring at the same time as the multiple first differences can be calculated. The correction coefficient for that sub-exponential distribution is then obtained based on the relative difference between the average of the first differences and the average of the initial probability density functions. It is understood that the closer the correction coefficient is to 1, the closer the sub-exponential distribution function under that sub-exponential distribution is to the distribution of the first differences.
[0168]
[0169]
[0170] Where, δ k,s Let represent the initial probability density function of the k-th business (i.e. the target business) at time t under the sub-exponential distribution type s.
[0171] In this embodiment, using the same sub-exponential distribution type to describe the relationship between the observation difference and the first time for different services may result in a significant discrepancy between the described relationship and the actual situation. Therefore, by fitting the first difference and the first time based on the conformance parameter, at least two sub-exponential distribution functions are determined. Then, among these at least two sub-exponential distribution functions, the sub-exponential distribution type that best matches the distribution of the first difference is identified, and its corresponding sub-exponential distribution function is determined as the target probability density function. This ensures that the relationship between the observation difference and the first time, represented by the obtained target probability density function, more closely matches the relationship between the first difference and the first time, thereby improving the accuracy of the test information.
[0172] In one embodiment of this application, considering that the conformance parameter is the key to determining whether the probability density function of a sub-exponential distribution type accurately describes the target association relationship, the conformance parameter can be optimized by the loss between the first difference and the predicted difference to obtain the optimal conformance parameter, thereby determining the probability density function that best describes the relationship between the observed difference and the first time under this sub-exponential distribution type. For example, in a content recommendation scenario, the conformance parameter can be optimized in the following way so that the correlation between the fitted difference value and time of the recommendation effect index represented by the target probability density function is more closely related to the correlation between the actual difference value and time of the recommendation effect index. Therefore, fitting the first difference and the first time based on the conformance parameter to obtain the target probability density function includes: determining the initial probability density function according to the preset conformance parameter and the first time; determining the predicted difference output by the initial probability density function according to the first time and the initial probability density function; and iteratively training the preset conformance parameter based on the loss between the first difference and the predicted difference to obtain the conformance parameter and the target probability density function.
[0173] Specifically, with the goal of minimizing the loss between the first difference and the predicted difference, the preset obedience parameters are iteratively trained to obtain the obedience parameters and the target probability density function.
[0174] In one implementation, the initial probability density function can be calculated based on the formula for the corresponding sub-exponential distribution type.
[0175] Optionally, based on the loss between the first difference and the predicted difference, the preset obedience parameters are iteratively trained to obtain the obedience parameters and the target probability density function, including:
[0176] For each type of exponential distribution, find information about g. s The optimal parameters α* and β* that (t) follow are an optimization problem, specifically expressed as follows:
[0177]
[0178] Where α and β represent preset compliance parameters, g s (t;α,β) represents the prediction difference, L(a k (t),g s (t;α,β)) represents the loss between the first difference and the predicted difference, λ represents the regularization parameter, Φ(α,β) represents the regularization function, and T represents the time or the number of preset periods.
[0179] Optionally, the above formula means that during multiple iterations, a set of optimal obedience parameters α* and β* are determined from the preset obedience parameters obtained in each iteration, so that the loss between the first difference after regularization and the predicted difference is minimized, and the predicted curve fits the true curve best.
[0180] Optionally, the loss function used to calculate the loss between the first difference and the predicted difference can be, but is not limited to, any one of the absolute value loss function, squared loss function, Huber loss function, and other custom loss functions. The regularization function can be, but is not limited to, L1 regularization function, L2 regularization function, and other custom functions. In practical applications, appropriate loss functions and regularization functions can be selected to optimize the conformance parameters according to specific needs. This application embodiment does not impose any limitations on the specific forms of the loss function and regularization function.
[0181] In this embodiment, the preset compliance parameters are iteratively trained by the loss between the first difference and the predicted difference. The preset compliance parameters are continuously optimized through iterative training to obtain the optimal preset compliance parameters. The optimal preset compliance parameters are then determined as the final compliance parameters. Based on the final compliance parameters, the probability density function (i.e., the target probability density function) that best describes the relationship between the observed difference and the first time under this exponential distribution type is determined.
[0182] In one embodiment of this application, determining a reference observation value based on the difference between observation values and a second observation value includes: correcting the difference between observation values according to a preset value to obtain a first value; determining a reference observation value based on the first value and the second observation value; the magnitude of the reference observation value is positively correlated with the magnitude of the second observation value.
[0183] For example, in a content recommendation scenario, the fitting difference value of the recommendation effect index is corrected according to a preset value to obtain the corrected fitting difference value (i.e., the first value). Then, based on the true value of the recommendation effect index of the multimedia push service of the control group after configuring the target service to the control group (i.e., the second observation value) and the corrected fitting difference value, the fitting value of the recommendation effect index (i.e., the reference observation value) is determined.
[0184] It should be noted that the first value, the second observation, and the reference observation all correspond to the same time.
[0185] Optionally, the difference between the preset value and the observed value is summed to obtain the first value. Considering that in practical applications, the reference observed value corresponds to the observed value of the indicator obtained when the target service is not configured in the control group, while the second observed value corresponds to the observed value of the indicator obtained when the target service is configured in the control group, generally, the observed value obtained by configuring the target service is greater than the observed value obtained by not configuring the target service. In reality, the reference observed value is less than the second observed value. Therefore, to make the reference observed value more closely reflect the actual situation, the purpose of the preset value is to ensure that the reference observed value at the same time is less than the second observed value. Therefore, it can be understood that the preset value is a constant greater than 0. Since the reference observed value is generally less than 1, the constant 1 is considered a preferred value for the preset value.
[0186] In one possible implementation, the difference between the observations and the second observation are used to calculate the reference observation based on the following formula:
[0187]
[0188] in, This represents the reference observation value of the k-th service (i.e., the target service) at time t. This represents the observation value of the k-th service in the control group at time t (e.g., the t-th time period), which can be understood here as the second observation value. This represents the difference in observations at time t for the k-th service, where n represents a preset value, and n≥0.
[0189] In this embodiment, the difference between the observed values is corrected using a preset value, so that the reference observed value determined based on the corrected difference between the observed values (i.e., the first value) and the second observed value is more in line with the actual situation, thereby improving the accuracy of the test results of the target service.
[0190] In one embodiment of this application, considering the existence of multiple service configurations, the target service includes at least two services. Determining a reference observation value based on the difference between observation values and a second observation value includes: for each service, correcting the difference between observation values corresponding to the service according to a preset value to obtain a first value corresponding to the service; performing product processing on the first values of all services to obtain an intermediate value; determining a reference observation value based on the intermediate value and the second observation value; the magnitude of the reference observation value is positively correlated with the magnitude of the second observation value.
[0191] For example, in a content recommendation scenario, for each multimedia push service, the fitted difference value of the recommendation effect index corresponding to that multimedia push service is corrected according to a preset value to obtain the corrected fitted difference value (i.e., the first value) for that multimedia push service. Then, the fitted difference values corresponding to all multimedia push services are multiplied to obtain the multiplied fitted difference value (i.e., the median value). After that, based on the actual value of the recommendation effect index of the multimedia push service in the control group (i.e., the second observation value) and the multiplied fitted difference value after configuring the target service in the control group, the fitted value of the recommendation effect index (i.e., the reference observation value) is determined.
[0192] Optionally, the magnitude of the reference observation is proportional to the magnitude of the second observation.
[0193] In practical applications, the services included in the target service are not required to be fully configured at the same time. They can be configured at the same or different times according to the actual needs of each service.
[0194] It is understandable that after configuring the target service to the control group, the reference observations are obtained in a preset test environment, which is mainly for an environment that assumes the target service is not configured to the control group. Therefore, the reference observations can be understood as observations that are counterfactual representations of the control group.
[0195] In one possible implementation, the difference between the observations and the second observation are used to calculate the reference observation based on the following formula:
[0196]
[0197] in, y represents the reference observation value at time t for all services included in the target service. t,0 This represents the observation value at time t (e.g., the t-th time period) of all services included in the target service within the control group. This can be understood as the second observation value. This represents the difference in observations at time t for the k-th service, where n represents a preset value (n≥0), K represents the total number of services included in the target service, and ∏ is the product operator.
[0198] In this embodiment, when the target service includes at least two services, the difference between the observed values corresponding to each service can be corrected according to a preset value to obtain a first value for each service. The first values of all services are then multiplied to obtain an intermediate value, thereby achieving the correction and fusion of the difference between the observed values for all services. Furthermore, based on the intermediate value and the second observed value, a reference observed value that integrates multiple services can be determined, thus achieving comprehensive prediction for multiple services.
[0199] In one embodiment of this application, determining test information for a target service based on a reference observation includes: acquiring a fourth observation that matches the time of a second observation; the fourth observation is an observation collected from the experimental group after configuring the target service to the experimental group; and determining test information based on the reference observation and the fourth observation.
[0200] For example, in a content recommendation scenario, the fourth observation is the actual value of the recommendation performance index of the multimedia push service in the experimental group after the multimedia push service is configured in the experimental group, matching the time of configuration with the control group. In other words, it is the value of the recommendation performance index collected from the experimental group after the multimedia push service is configured.
[0201] Specifically, determining test information based on the reference observation and the fourth observation includes obtaining test information for the target service based on the relative difference between the reference observation and the fourth observation.
[0202] Optionally, the relative difference between the reference observation and the fourth observation can be determined according to the following formula:
[0203]
[0204] in, This represents the relative difference between the reference observation and the fourth observation for the k-th service at time t. This represents the reference observation value of the k-th service (i.e., the target service) at time t. This represents the observation value of the k-th service in the experimental group at time t (e.g., the t-th time period), which can be understood as the fourth observation value.
[0205] In another testing scenario, during long-term testing of the target service, if the target service is never configured for the control group, the test information of the target service can be determined directly by comparing the relative differences between the actual observations collected or detected by the experimental group and the actual observations collected or detected by the control group. That is, the relative difference between the actual observations of the experimental group and the actual observations of the control group can be calculated using the following formula:
[0206]
[0207] in, This indicates the relative difference between the true observed values of the experimental group and the true observed values of the control group. This represents the actual observed value of the k-th service (i.e., the target service) at time t. This represents the actual observed value of the k-th service in the experimental group at time t (e.g., the t-th time period). The actual observed values in the experimental group include the third and fourth observed values.
[0208] Optionally, considering that observed data is easily affected by periodicity and randomness in actual application, smoothing is required when evaluating the relative difference between the actual observed values of the experimental group and the control group at time t. A smoothing period can be set, containing multiple consecutive time intervals, points in time, or moments. The smoothing period includes time t and one or more times within a preset time interval preceding time t. The actual observed values of the experimental group and the control group at each time within the preset time interval preceding time t are combined with the actual observed values of the experimental group and the control group at time t to achieve smoothing. Using time t to represent periodic time data, for example, with a period of days, assuming time t is the 14th day of the entire test period, then one or more times within the preset time interval could be from the 7th to the 13th day.
[0209] The specific formula is expressed as follows:
[0210]
[0211] Where, r t (d) represents the relative difference between the true observed values of the experimental group and the control group at time t, where d represents the smoothing period, and y t-i,1 y represents the actual observed value of the experimental group at time ti. t-i,0 ω represents the actual observed value of the control group at time ti. t-i This represents the weighting coefficient of the actual observation at time *ti*, which characterizes the degree of influence of the actual observation at time *ti* on the actual observation at time *t*. Generally, the closer *ti* is to *t*, the greater the influence and the larger the weighting coefficient. It should be noted that the above formula can be used to calculate test information for a single service or to obtain comprehensive test information for multiple services.
[0212] In this embodiment, a fourth observation value that matches the time of the second observation value is obtained and combined with a reference observation value to determine the test information so as to achieve the test of the target service after configuring the target service to the control group.
[0213] In one embodiment of this application, considering that the observed values may be affected by periodicity, randomness, etc. during actual application, smoothing processing is required when evaluating the relative difference between the second observed value of the experimental group and the fourth observed value of the control group. Specifically, the same time corresponding to the second and fourth observed values is the second time. At least one reference time is included before the second time. Determining test information based on the reference observed value and the fourth observed value includes: obtaining the observed values of the control group and the experimental group at at least one reference time; and smoothing the observed values of the at least one reference time, the second observed value, and the fourth observed value according to the weighting coefficient of the at least one reference time and the weighting coefficient of the second time to obtain the test information.
[0214] For example, in a content recommendation scenario, the recommendation performance index of the multimedia push service for the control group and the experimental group at at least one reference time is obtained, based on the weighting coefficient of at least one reference time and the weighting coefficient of the second time.
[0215] The test information is obtained by smoothing the values of the recommendation performance index of the multimedia push service at at least one reference time, the actual values of the recommendation performance index of the multimedia push service in the control group after configuring the target service, and the actual values of the recommendation performance index of the multimedia push service in the experimental group. Here, the sampling time of the actual values of the recommendation performance index of the multimedia push service in the experimental group matches the time when the multimedia push service was configured in the control group; that is, the recommendation performance index values collected from the experimental group after configuring the multimedia push service.
[0216] It should be noted that if the reference time is before the target service configuration of the control group, the observation value of the control group at at least one reference time is the actual collected indicator value (i.e., the first observation value), and the observation value of the experimental group at at least one reference time is the third observation value; if it is after the target service configuration of the control group, the observation value of the control group at at least one reference time is the reference observation value determined by the method provided in the embodiments of this application, and the observation value of the experimental group at at least one reference time is the fourth observation value.
[0217] In a preferred embodiment, at least one reference time precedes the second time, and the at least one reference time and the second time form a smoothing period. Test information is obtained by smoothing the observations within this smoothing period. Optionally, the at least one reference time and the second time are consecutive time periods, points in time, or moments.
[0218] Specifically, the smoothing process for the first, second, and fourth observations is performed using the following formula:
[0219]
[0220] in, Let d represent the relative difference between the reference observation and the fourth observation of the k-th service at time t (i.e., the second time point), where d represents the smoothing period, and y represents the relative difference between the reference observation and the fourth observation of the k-th service. t-i,1 This represents the observation value of the experimental group at time ti (the specific value can be determined from the third and fourth observation values based on time ti). ω represents the reference observation at time ti. t-i ω represents the weighting coefficient of the observation at time ti. t-i Specifically, this includes a weighting coefficient for at least one reference time and a weighting coefficient for a second time. This weighting coefficient characterizes the degree of influence of an observation at either the reference time or the second time on the observation at the second time. Generally, the closer *ti* is to *t*, the greater the influence and the larger the weighting coefficient. It should be noted that the above formula can be used to calculate test information for a single service or comprehensive test information for multiple services.
[0221] In this embodiment, considering that observed values are easily affected by periodicity and randomness in actual application, the observed values of the control group and the experimental group at at least one reference time are obtained. Based on the weighting coefficients of the at least one reference time and the weighting coefficient of the second time, the observed values, the second observed value, and the fourth observed value are smoothed to obtain test information. In this way, smoothing the data can reduce the impact of external factors and data fluctuations on the test results of the target business, thereby improving the accuracy of the test results.
[0222] In summary, the method provided in this application embodiment can achieve the following beneficial effects:
[0223] 1. Able to conduct long-term testing of the target business using short-term metric data.
[0224] Given that experimental benefits tend to gradually diminish over time, this embodiment utilizes the decay characteristics of a sub-exponential distribution to model data from short-term reversal experiments. By optimizing the loss function and selecting the model, the distribution shape with the best fit is found, thereby evaluating metrics for businesses that cannot achieve long-term business performance without a control group. Therefore, this embodiment can make better predictions of long-term business benefits using short-term experimental data, and also solves the problem of benefit evaluation when long-term experiments are not possible. Specific experimental results are detailed below. Figure 9As can be seen, the method provided in this application embodiment can predict test information for the next year based on 14-day short-term experimental data. Its output results are also basically consistent with the actual effects of long-term experiments, fully verifying the high accuracy of the prediction effect of this application embodiment. Curve 1 represents the true relative difference, curve 2 represents the relative difference (i.e., test information) under the Gamma distribution, and curve 3 represents the relative difference (i.e., test information) under the log-normal distribution.
[0225] 2. The method output is flexible and can be adapted to various distribution trends.
[0226] Since the area under the curve of a sub-exponential distribution is required to be 1, to remove this requirement and address the limitations of the method provided in this application, the probability density functions of various sub-exponential distributions are modified to obtain a core function describing the distribution shape, controlled by parameters α and β. Because different distributions have different shapes, and the same distribution also has significantly different shapes under different parameters, the curves constructed from the probability density functions of various distributions are highly flexible and can effectively fit various index data.
[0227] 3. Low complexity and high training efficiency
[0228] Since the relationship between the observed value difference and the first time point is determined only from data from short-term experiments, which typically last no more than 14 days, the sample size available for training is limited. Using conventional machine learning models might lead to underfitting. However, given the assumptions, this embodiment only requires optimizing two parameters, α and β. The algorithm has fewer dimensions, lower complexity, and faster convergence, resulting in higher training efficiency. Generally, the relationship between the observed value difference and the first time point can be obtained instantly, making it suitable for real-time prediction in business applications, helping businesses understand test information in a timely manner.
[0229] 4. It can obtain the decay rate and effective period of the observed index.
[0230] Since business operations often cannot guarantee long-term effectiveness, in addition to evaluating the cumulative long-term benefits of already launched operations, business stakeholders also focus on the rate of decay and the effective period of the operations. The test information obtained in this application embodiment can output a relative difference curve of the business with respect to observed metrics. Specifically, the rate of decay of the business can be evaluated in two ways: 1) the percentage of the relative difference on day N after the business's launch relative to the original relative difference; 2) the time required for the relative difference to decay to K% of the original relative difference. Therefore, the method for determining test information in this application can output the rate of decay and the effective period of the relative difference, thereby helping business stakeholders to schedule projects and manage experiments, optimize and adjust existing operations in a timely manner, and contribute to the continuous and efficient development of the business.
[0231] 5. High accuracy in long-term comprehensive testing of multiple services.
[0232] When evaluating the long-term comprehensive benefits of multiple services, the SED algorithm effectively combines the results of long-term holdout experiments with the algorithm's predictions, successfully covering all experimental services and providing accurate revenue estimates.
[0233] Reference Figure 10 This provides a schematic diagram illustrating the effect of determining multiple test information, such as... Figure 10 As shown, line 1 is the curve constructed using the third and fourth observations of the experimental group, line 2 is the curve constructed using the first and second observations, line 3 is the curve constructed using reference observations under the Gamma distribution, and line 4 is the curve constructed using reference observations under the log-normal distribution. Therefore, the method provided in this application embodiment can still guarantee a high accuracy rate when testing multiple services.
[0234] Furthermore, the testing process not only introduces a correction coefficient to correct the difference between the result and the actual return during the reversal period, but also introduces a smoothing period to reduce the impact of external factors and data fluctuations on the test results, thereby further improving the accuracy and robustness of the test information determination method provided in this application embodiment for predicting the observations of the target business.
[0235] It should be noted that although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed in order to achieve the desired result.
[0236] Figure 11 This is a block diagram of a test information determination device according to an embodiment of this application.
[0237] like Figure 11 As shown, the device for determining test information includes: an acquisition module 1101, a processing module 1102, and a test module 1103. Among them,
[0238] The acquisition module 1101 is used to acquire the first observation value of the control group, the second observation value of the control group, and the third observation value that is time-matched with the first observation value; the first observation value is the observation value collected from the control group before configuring the target service to the control group, the second observation value is the observation value collected from the control group after configuring the target service to the control group, and the third observation value is the observation value collected from the experimental group after configuring the target service to the experimental group; the target service includes multimedia push service, and the observation value is the value of the recommendation effect index of multimedia push service.
[0239] The processing module 1102 is used to determine the first difference between the first observation and the third observation, and to determine the reference observation of the control group under the preset test environment based on the first difference and the second observation. The effective time of the preset test environment matches the time when the target service is configured in the control group, and the configuration of the target service to the control group is prohibited during the effective time. The reference observation is the value of the recommendation effect index of the multimedia push service in the control group under the preset test environment.
[0240] Test module 1103 is used to determine test information for the target service based on reference observations.
[0241] In one embodiment, the processing module 1102 is specifically used for,
[0242] The first difference and the first time are fitted to obtain the target correlation. The target correlation is used to characterize the relationship between the observation difference and time. The observation difference is used to characterize the degree of fitting difference between the observations of the experimental group and the observations of the control group at the same time. The first time is the same time corresponding to the first observation and the third observation.
[0243] The generation time of the second observation is mapped according to the target correlation to obtain the difference in observations corresponding to the control group.
[0244] The reference observation is determined based on the difference between the observed values and the second observation.
[0245] In one embodiment, the processing module 1102 is specifically used to perform fitting processing on the first difference and the first time based on the obedience parameter to obtain the target probability density function; the target probability density function is used to characterize the target correlation; the obedience parameter is related to the distribution of the first difference.
[0246] In one embodiment, the target probability density function is a probability density function of the sub-exponential distribution type, and the processing module 1102 is specifically used for,
[0247] By fitting the first difference and the first time based on the parameters, at least two sub-exponential distribution functions are obtained.
[0248] The target probability density function is the probability density function that best matches the distribution of the first difference among at least two exponential distribution functions.
[0249] In one embodiment, the processing module 1102 is specifically used for,
[0250] The initial probability density function is determined immediately based on the preset parameters.
[0251] Based on the first time and the initial probability density function, determine the prediction difference output by the initial probability density function.
[0252] Based on the loss between the first difference and the predicted difference, the preset obedience parameters are iteratively trained to obtain the obedience parameters and the target probability density function.
[0253] In one embodiment, the processing module 1102 is specifically used for,
[0254] The difference between the observed values is corrected based on the preset values to obtain the first value.
[0255] A reference observation is determined based on the first and second observations; the magnitude of the reference observation is positively correlated with the magnitude of the second observation.
[0256] In one embodiment, the target service includes at least two services, and the processing module 1102 is specifically used for,
[0257] For each service, the difference between the observed values corresponding to the service is corrected according to the preset values to obtain the first value corresponding to the service.
[0258] The first value of all business transactions is multiplied to obtain the intermediate value.
[0259] A reference observation is determined based on the median and the second observation; the magnitude of the reference observation is positively correlated with the magnitude of the second observation.
[0260] In one embodiment, the test module 1103 is specifically used for,
[0261] Obtain a fourth observation that matches the time of the second observation; the fourth observation is the observation collected from the experimental group after the target service is configured to the experimental group.
[0262] Test information is determined based on the reference observation and the fourth observation.
[0263] In one embodiment, the same time corresponding to the second and fourth observations is designated as the second time, and at least one reference time precedes the second time. The test module 1103 is specifically used for...
[0264] Obtain observations for the control group and the experimental group at at least one reference time.
[0265] Based on the weighting coefficients of at least one reference time and the weighting coefficients of the second time, the observations of at least one reference time, the second observation, and the fourth observation are smoothed to obtain test information.
[0266] The test information determination device proposed in this application addresses the issue that existing A / B testing schemes rely on comparing and analyzing indicator data observed in the experimental group and the control group to test the service. This presents two main problems: First, using existing A / B testing schemes for long-term service testing results in the control group being unable to perceive the user experience of the service for an extended period, thus infringing on the right of some test subjects to use the service. Furthermore, the inability of some test subjects to use the service directly increases the development and experimental costs of the service. Second, if a full configuration of the service is chosen, it is impossible to obtain indicator data for the control group unaffected by the service release after full configuration, making it impossible to complete long-term testing of the service.
[0267] Therefore, this device uses the time point of full service configuration as the dividing line. By combining the indicator data (i.e., observed values) observed in the experimental group before full service configuration and the indicator data observed in the control group, along with the indicator data observed in the control group after full service configuration, it predicts the indicator data for the control group not configuring the service after the aforementioned time point, thereby achieving the testing of the service. Specifically, it first obtains the first observation value and the second observation value of the control group; where the first observation value is the observation value collected from the control group before configuring the target service, and the second observation value is the observation value collected from the control group after configuring the target service; and it obtains the third observation value that matches the time of the first observation value; the third observation value is the observation value collected from the experimental group after configuring the target service, and the target service includes multimedia push service, and the observation value is the value of the recommendation effect indicator of multimedia push service. Then, it determines the first difference between the first observation value and the third observation value. Through the first difference, it can know the actual difference between the observation values of the two groups before service configuration. Based on the actual difference and the data observed in the control group after configuring the target service (i.e., the second observation value), a reference observation value for the control group in a preset test environment is determined. The effective time of the preset test environment matches the time the control group configures the target service, and configuring the target service to the control group is prohibited during the effective time. The reference observation value is the recommendation performance index of the multimedia push service in the preset test environment. The test information for the target service is then determined based on this reference observation value. This approach ensures full configuration of the service, allowing the control group to use it and protecting their right to use the service, while also reducing the development and experimental costs of the service. Furthermore, after full configuration of the service, the actual control data of the control group under the assumption that the service is not configured can be predicted, enabling testing of the service based on the experimental data of the experimental group and the actual control data, thus achieving the goal of long-term testing of the service.
[0268] It should be understood that the units recorded in the test information determining device are related to the reference. Figure 3 The steps in the described method correspond to each other. Therefore, the operations and features described above for the method also apply to the test information determination device and its constituent units, and will not be repeated here. The test information determination device can be pre-implemented in a computer device's browser or other security applications, or it can be loaded into the computer device's browser or its security applications through download or other means. The corresponding units in the test information determination device can cooperate with the units in the computer device to implement the solutions of the embodiments of this application.
[0269] The division of modules or units mentioned in the detailed description above is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0270] It should be noted that for details not disclosed in the test information determination device of this application embodiment, please refer to the details disclosed in the above embodiments of this application, which will not be repeated here.
[0271] The following is for reference. Figure 12 , Figure 12 A schematic diagram of a computer device suitable for implementing embodiments of this application is shown, such as... Figure 12 As shown, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1202 or programs loaded from storage section 1208 into random access memory (RAM) 1203. The RAM 1203 also stores various programs and data required for the system's operating instructions. The CPU 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0272] The following components are connected to I / O interface 1205: an input section 1206 including a keyboard, mouse, etc.; an output section 1207 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN card, modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to I / O interface 1205 as needed. Removable media 1211, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1210 as needed so that computer programs read from them can be installed into storage section 1208 as needed.
[0273] Specifically, according to embodiments of this application, the flowchart above refers to... Figure 3 The described process can be implemented as a computer software program. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program contains program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 1209, and / or installed from removable medium 1211. When the computer program is executed by central processing unit (CPU) 1201, it performs the functions defined in the system of this application.
[0274] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0275] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operational instructions of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two connected blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified functions or operational instructions, or using a combination of dedicated hardware and computer instructions.
[0276] The units or modules described in the embodiments of this application can be implemented in software or hardware. The described units or modules can also be configured in a processor; for example, a processor can be described as including a violator detection unit, a multimodal detection unit, and a recognition unit. The names of these units or modules do not necessarily constitute a limitation on the unit or module itself.
[0277] On the other hand, this application also provides a computer-readable storage medium, which may be included in the computer device described in the above embodiments, or may exist independently and not assembled into the computer device. The aforementioned computer-readable storage medium stores one or more programs that, when used by one or more processors, execute the method for determining test information described in this application. For example, it may execute... Figure 3 The steps of the method for determining test information are shown.
[0278] This application provides a computer program product including instructions that, when executed, cause the method described in this application to be performed. For example, it can execute... Figure 3 The steps of the method for determining test information are shown.
[0279] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A method for determining test information, applied in the field of content recommendation, characterized in that, include: Obtain the first observation value of the control group, the second observation value of the control group, and the third observation value that is time-matched with the first observation value; The first observation value is an observation collected from the control group before configuring the target service to the control group; the second observation value is an observation collected from the control group after configuring the target service to the control group; and the third observation value is an observation collected from the experimental group after configuring the target service to the experimental group. The target service includes a multimedia push service, and the observation value is the value of the recommendation effect index of the multimedia push service. Determine the first difference between the first observation and the third observation, and perform fitting processing on the first difference and the first time to obtain the target correlation relationship; The target correlation is used to characterize the relationship between the observed value difference and time. The observed value difference is used to characterize the degree of fit difference between the observed values of the experimental group and the observed values of the control group at the same time. The first time is the same time corresponding to the first observed value and the third observed value. The generation time of the second observed value is mapped according to the target correlation to obtain the observed value difference corresponding to the control group. The reference observed value of the control group in a preset test environment is determined according to the observed value difference and the second observed value. The effective time of the preset test environment matches the time when the control group configures the target service, and the configuration of the target service to the control group is prohibited during the effective time. The reference observation value is the value of the recommendation effect index of the multimedia push service in the control group under the preset test environment. The preset test environment is a virtual data control environment constructed to determine the reference observation value data of the control group in the state without the target service configured. The test information for the target service is determined based on the reference observations.
2. The method for determining test information according to claim 1, characterized in that, The fitting process of the first difference and the first time to obtain the target correlation includes: The first difference and the first time are fitted based on the observability parameter to obtain the target probability density function; the target probability density function is used to characterize the target correlation; the observability parameter is related to the distribution of the first difference.
3. The method for determining test information according to claim 2, characterized in that, The target probability density function is a probability density function of the sub-exponential distribution type. The step of fitting the first difference and the first time based on the observable parameters to obtain the target probability density function includes: Based on the observability parameters, the first difference and the first time are fitted to obtain at least two exponential distribution functions. The probability density function of the sub-exponential distribution type that is closest to the distribution of the first difference among the at least two sub-exponential distribution functions is determined as the target probability density function.
4. The method for determining test information according to claim 2 or 3, characterized in that, The step of fitting the first difference and the first time based on the observable parameters to obtain the target probability density function includes: The initial probability density function is determined based on the preset observability parameters and the first time. Based on the first time and the initial probability density function, determine the prediction difference output by the initial probability density function; Based on the loss between the first difference and the predicted difference, the preset obedience parameters are iteratively trained to obtain the obedience parameters and the target probability density function.
5. The method for determining test information according to any one of claims 1-4, characterized in that, The step of determining the reference observation value of the control group under the preset test environment based on the difference between the observed values and the second observed value includes: The difference between the observed values is corrected according to a preset value to obtain a first value; The reference observation value is determined based on the first value and the second observation value; the magnitude of the reference observation value is positively correlated with the magnitude of the second observation value.
6. The method for determining test information according to any one of claims 1-5, characterized in that, The target service includes at least two services, and determining the reference observation value of the control group under the preset test environment based on the observation difference and the second observation value includes: For each service, the difference between the observed values corresponding to the service is corrected according to a preset value to obtain the first value corresponding to the service; The first values of all the aforementioned services are multiplied to obtain the intermediate value; The reference observation value is determined based on the intermediate value and the second observation value; the magnitude of the reference observation value is positively correlated with the magnitude of the second observation value.
7. The method for determining test information according to any one of claims 1-6, characterized in that, The step of determining the test information of the target service based on the reference observations includes: Obtain a fourth observation that matches the time of the second observation; the fourth observation is an observation collected from the experimental group after the target service is configured in the experimental group. The test information is determined based on the reference observation and the fourth observation.
8. The method for determining test information according to claim 7, characterized in that, The time corresponding to the second observation and the fourth observation is defined as the second time, and at least one reference time precedes the second time. Determining the test information based on the reference observation and the fourth observation includes: Obtain the observations of the control group and the experimental group at at least one reference time; Based on the weighting coefficients of the at least one reference time and the second time, the observed values of the at least one reference time, the second observed value, and the fourth observed value are smoothed to obtain the test information.
9. A device for determining test information, characterized in that, include: The acquisition module is used to acquire the first observation value of the control group, the second observation value of the control group, and the third observation value that matches the time of the first observation value; The first observation value is an observation collected from the control group before configuring the target service to the control group; the second observation value is an observation collected from the control group after configuring the target service to the control group; and the third observation value is an observation collected from the experimental group after configuring the target service to the experimental group. The target service includes a multimedia push service, and the observation value is the value of the recommendation effect index of the multimedia push service. The processing module is used to determine the first difference between the first observation value and the third observation value, and to perform fitting processing on the first difference and the first time to obtain the target correlation relationship; The target correlation is used to characterize the relationship between the observed value difference and time. The observed value difference is used to characterize the degree of fit difference between the observed values of the experimental group and the observed values of the control group at the same time. The first time is the same time corresponding to the first observed value and the third observed value. The generation time of the second observed value is mapped according to the target correlation to obtain the observed value difference corresponding to the control group. The reference observed value of the control group in a preset test environment is determined according to the observed value difference and the second observed value. The effective time of the preset test environment matches the time when the control group configures the target service, and the configuration of the target service to the control group is prohibited during the effective time. The reference observation value is the value of the recommendation effect index of the multimedia push service in the control group under the preset test environment. The preset test environment is a virtual data control environment constructed to determine the reference observation value data of the control group in the state without the target service configured. The testing module is used to determine the test information of the target service based on the reference observations.
10. The apparatus for determining test information according to claim 9, characterized in that, The processing module is specifically used to fit the first difference and the first time based on the observability parameter to obtain the target probability density function; the target probability density function is used to characterize the target correlation; the observability parameter is related to the distribution of the first difference.
11. The apparatus for determining test information according to claim 10, characterized in that, The target probability density function is a probability density function of the sub-exponential distribution type. The processing module is specifically used to perform fitting processing on the first difference and the first time based on the observability parameter to obtain at least two sub-exponential distribution functions. The probability density function of the sub-exponential distribution type that is closest to the distribution of the first difference among the at least two sub-exponential distribution functions is determined as the target probability density function.
12. The apparatus for determining test information according to claim 10 or 11, characterized in that, The processing module is specifically used to: determine an initial probability density function based on preset compliance parameters and the first time; determine the prediction difference output by the initial probability density function based on the first time and the initial probability density function; and iteratively train the preset compliance parameters based on the loss between the first difference and the prediction difference to obtain the compliance parameters and the target probability density function.
13. The apparatus for determining test information according to any one of claims 9-12, characterized in that, The processing module is specifically used to correct the difference between the observed values according to a preset value to obtain a first value; The reference observation value is determined based on the first value and the second observation value; the magnitude of the reference observation value is positively correlated with the magnitude of the second observation value.
14. The apparatus for determining test information according to any one of claims 9-13, characterized in that, The target service includes at least two services. The processing module is specifically used to correct the difference between the observed values corresponding to each service according to a preset value to obtain a first value corresponding to the service. The first values of all the aforementioned services are multiplied to obtain the intermediate value; The reference observation value is determined based on the intermediate value and the second observation value; The magnitude of the reference observation is positively correlated with the magnitude of the second observation.
15. The apparatus for determining test information according to any one of claims 9-14, characterized in that, The testing module is specifically used to acquire a fourth observation value that matches the time of the second observation value; the fourth observation value is the observation value collected from the experimental group after the target service is configured to the experimental group; The test information is determined based on the reference observation and the fourth observation.
16. The apparatus for determining test information according to claim 15, characterized in that, The time corresponding to the second observation and the fourth observation is the second time. The second time includes at least one reference time. The test module is specifically used to obtain the observation values of the control group and the experimental group at the at least one reference time. Based on the weighting coefficients of the at least one reference time and the second time, the observed values of the at least one reference time, the second observed value, and the fourth observed value are smoothed to obtain the test information.
17. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for determining test information as described in any one of claims 1 to 8.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the method for determining test information as described in any one of claims 1 to 8.
19. A computer program product, characterized in that, The computer program product includes instructions that, when executed, cause the method as described in any one of claims 1 to 8 to be performed.