AB experiment sampling method, apparatus, device, and medium

By generating a cumulative traffic contribution curve and performing inflection point detection, stratified sampling and locking buckets, the problem of large traffic differences in AB experiments in existing technologies is solved, the uniformity of user volume and activity is achieved, and the experimental effect is improved.

CN120429543BActive Publication Date: 2025-10-21ZHIZHESIHAIBEIJINGTECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510926460.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-21
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

In the existing technology, AB experiments have poor experimental results due to the large difference in traffic between the two groups AB, making it difficult to accurately judge the effectiveness of the strategy.

Method used

By determining the traffic contribution of the locked bucket, generating a cumulative traffic contribution curve and performing inflection point detection, stratification is performed based on the inflection point, and locked buckets are extracted for each experimental group to ensure uniformity in user volume and activity.

Benefits of technology

It reduces the false positive rate of AB experiments, improves experimental results, ensures the uniformity of user volume and activity, and reduces data analysis noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429543B_ABST
    Figure CN120429543B_ABST
Patent Text Reader

Abstract

The application discloses an AB experiment sampling method and device, equipment and medium, which are applied to the technical field of data processing, and are used to solve the problem that the AB experiment effect is poor due to large flow difference between AB two groups. Specifically, a locking bucket of a current AB experiment task is determined from each flow bucket, the flow contribution degree of each locking bucket is obtained, the cumulative flow contribution curve is generated based on the flow contribution degree of each locking bucket, the inflection point detection is performed on the cumulative flow contribution curve to obtain each flow contribution degree inflection point, each locking bucket is layered based on each flow contribution degree inflection point to obtain each locking bucket sampling layer, and the target user flow of each experiment group is obtained by extracting the locking bucket from each locking bucket sampling layer for each experiment group, so that the flow difference between each experiment group can be reduced by using the flow contribution degree inflection point for layering and then sampling, and then the false positive rate of the AB experiment can be reduced, and the experiment effect of the AB experiment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to an AB experiment sampling method, device, equipment and medium. Background Art

[0002] AB experiment is a controlled experimental method. Its core is to divide different strategies into two groups, A and B, and verify the differences in the effectiveness of different strategies through data collection and statistical analysis. It is widely used in Internet products, recommendation systems, advertising systems, digital operations and intelligent marketing, natural sciences, psychology, economics and biomedicine, and is an important means of data-driven and scientific research.

[0003] Normally, when diverting two groups A and B, if the flow difference between the two groups A and B is large, it will interfere with the subsequent AB experimental data analysis. For example, it is difficult to determine whether the difference in experimental effect is caused by the strategy itself or by the flow difference. Therefore, in the experimental diversion stage, the flow difference between the two groups A and B needs to be maintained at a small level to avoid noise in the subsequent AB experimental data analysis. However, the diversion method in the existing technology cannot ensure that the flow difference between the two groups A and B is maintained at a small level, which leads to poor AB experimental results. Summary of the Invention

[0004] This application provides an AB experiment sampling method, device, equipment, and medium to solve the problem of poor AB experiment results due to the large difference in flow rates between the two AB groups in the prior art. The technical solutions provided by this application are as follows:

[0005] In one aspect, the present application provides an AB experiment sampling method, comprising:

[0006] Determine the traffic bucket corresponding to the current AB experiment task from each traffic bucket of the traffic sampling layer as the locked bucket;

[0007] Obtain the traffic contribution of each locked bucket; wherein the traffic contribution of each locked bucket is determined based on the current number of users of the locked bucket and the current user activity;

[0008] Generate a cumulative traffic contribution curve based on the traffic contribution of each locked bucket;

[0009] Perform inflection point detection on the cumulative flow contribution curve to obtain the inflection points of each flow contribution;

[0010] Based on the inflection points of each traffic contribution, each locked bucket is stratified to obtain a sampling layer for each locked bucket;

[0011] A locked bucket is extracted for each experimental group from each locked bucket sampling layer to obtain the target user traffic of each experimental group.

[0012] Optionally, determining the traffic bucket corresponding to the current AB experiment task from each traffic bucket of the traffic sampling layer as the locked bucket includes:

[0013] Based on the number of experimental groups and the proportion of experimental group traffic in the current AB experiment task, randomly select traffic from each traffic bucket that is not locked by other AB experiment tasks in the traffic sampling layer as the locked bucket of the current AB experiment task.

[0014] Optionally, obtain the traffic contribution of each locked bucket, including:

[0015] Obtaining the traffic index values ​​of each locked bucket; wherein the traffic index values ​​of each locked bucket are predetermined based on the current number of users of the locked bucket and the current user activity;

[0016] Based on the traffic index values ​​of each locked bucket, the traffic contribution of each locked bucket is determined.

[0017] Optionally, based on the traffic contribution of each locked bucket, a cumulative traffic contribution curve is generated, including:

[0018] Sort the traffic contribution of each locked bucket in descending order to construct a traffic contribution sequence;

[0019] The first traffic contribution in the traffic contribution sequence is determined as the first data point, and the sum of the i-th traffic contribution in the traffic contribution sequence and each traffic contribution before the i-th traffic contribution is sequentially determined as a data point to construct a data point sequence; where i is a positive integer greater than 1 and less than or equal to n, and n is the total number of locked buckets;

[0020] Based on the data point sequence, a cumulative flow contribution curve is generated.

[0021] Optionally, perform inflection point detection on the cumulative traffic contribution curve to obtain inflection points of each traffic contribution, including:

[0022] The inflection point detection operation is cyclically performed on the cumulative traffic contribution curve until the termination condition is determined to be satisfied, and the candidate traffic contribution inflection points obtained by each inflection point detection operation are determined as the respective traffic contribution inflection points; wherein the inflection point detection operation is:

[0023] Perform inflection point detection on the cumulative flow contribution curve to be detected to obtain a candidate flow contribution inflection point; wherein, when the inflection point detection operation is performed for the first time, the cumulative flow contribution curve to be detected is the cumulative flow contribution curve; when the inflection point detection operation is not performed for the first time, the cumulative flow contribution curve to be detected is the two cumulative flow contribution curve segments segmented when the inflection point detection operation was performed last time;

[0024] Taking the candidate traffic contribution inflection point as the dividing point, the cumulative traffic contribution curve to be detected is divided into two cumulative traffic contribution curve segments.

[0025] Optionally, each locked bucket is stratified based on each traffic contribution inflection point to obtain each locked bucket sampling layer, including:

[0026] Determine each flow contribution interval based on each two adjacent flow contribution inflection points among each flow contribution inflection point;

[0027] Based on the traffic contribution of each locked bucket, each locked bucket is divided into a matching traffic contribution interval to obtain each locked bucket sampling layer.

[0028] Optionally, a locked bucket is extracted from each locked bucket sampling layer for each experimental group to obtain the target user traffic of each experimental group, including:

[0029] For each locked bucket sampling layer, randomly select a locked bucket for each experimental group from the locked bucket sampling layer according to the traffic proportion of the experimental group;

[0030] For each experimental group, the union of the locked buckets randomly selected for the experimental group in each locked bucket sampling layer is determined as the target user traffic of the experimental group.

[0031] On the other hand, the present application provides an AB experiment sampling device, comprising:

[0032] The traffic locking unit is used to determine the traffic bucket corresponding to the current AB experiment task from each traffic bucket of the traffic sampling layer as the locking bucket;

[0033] A contribution acquisition unit, configured to acquire the traffic contribution of each locked bucket; wherein the traffic contribution of each locked bucket is determined based on the current number of users of the locked bucket and the current user activity;

[0034] A curve generating unit, configured to generate a cumulative traffic contribution curve based on the traffic contribution of each locked bucket;

[0035] An inflection point determination unit is used to detect the inflection point of the cumulative flow contribution curve to obtain the inflection point of each flow contribution;

[0036] A traffic stratification unit is used to stratify each locked bucket based on each traffic contribution inflection point to obtain a sampling layer for each locked bucket;

[0037] The traffic extraction unit is used to extract a locked bucket for each experimental group from each locked bucket sampling layer to obtain the target user traffic of each experimental group.

[0038] On the other hand, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned AB experimental sampling method when executing the computer program.

[0039] On the other hand, the present application provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a processor, the above-mentioned AB experimental sampling method is implemented.

[0040] The beneficial effects of this application are as follows:

[0041] The present application determines the traffic contribution of each locked bucket based on the current user volume and current user activity, and generates a cumulative traffic contribution curve based on the traffic contribution of each locked bucket. After performing inflection point detection on the cumulative traffic contribution curve, the mutation point of the traffic contribution can be obtained. Therefore, after stratifying each locked bucket using the mutation point, locked bucket sampling layers with different traffic contribution intervals can be obtained. Subsequently, after extracting locked buckets from each locked bucket sampling layer for each experimental group, the user volume and user activity of the target user traffic of each experimental group can be made more uniform, thereby reducing the false positive rate of the AB experiment and improving the experimental effect of the AB experiment.

[0042] Other features and advantages of the present application will be described in the following description, and in part, will become apparent from the description or be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the written description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic diagrams and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0044] Figure 1 This is a diagram showing the distribution of total users using the traditional sampling method;

[0045] Figure 2 This is a diagram of the total user activity distribution using the traditional sampling method;

[0046] Figure 3 This is a schematic diagram of the overview of the AB experimental sampling method in the embodiments of the present application;

[0047] Figure 4 This is a schematic diagram of the overall distribution of the number of users in each traffic bucket in the embodiment of the present application;

[0048] Figure 5 This is a schematic diagram of the sorting and waiting when sampling multiple AB experimental tasks in the traditional sampling method;

[0049] Figure 6 This is a schematic diagram of the locking flow when sampling multiple AB experimental tasks in the embodiment of the present application;

[0050] Figure 7 This is a schematic diagram of the division method of the locked bucket sampling layer in an embodiment of the present application;

[0051] Figure 8 This is a schematic diagram of the operation of extracting a locked barrel for each experimental group in the embodiment of the present application;

[0052] Figure 9 This is a schematic diagram comparing the sampling scheme and the random sampling scheme in terms of user volume and overall active false positive rate in the embodiment of the present application;

[0053] Figure 10 This is a schematic diagram comparing the sampling scheme and the random sampling scheme in different dimensions of user volume and overall active false positive rate in the embodiment of the present application;

[0054] Figure 11 This is a schematic diagram of the composition structure of the AB experiment sampling device in the embodiment of the present application;

[0055] Figure 12 Schematic diagram of the hardware structure of the electronic device in the embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and beneficial effects of this application more clearly understood, the technical solutions of this application will be clearly and completely described below in conjunction with the embodiments and drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of this application.

[0057] In order to facilitate those skilled in the art to better understand this application, the technical terms involved in this application are briefly introduced below.

[0058] AB experiment: A statistical evaluation method used to assess whether a change to a variable has a significant impact on a target indicator. In an AB experiment, participants are divided into two or more groups. One group uses an existing method, function, or design that applies the variable as a control group, while the other group uses a new method, function, or design that applies the variable as an experimental group. By comparing the outcome indicators of the control and experimental groups, it is determined whether the new method, function, or design has produced a positive effect.

[0059] The sampling system is the core basic system that supports AB experiments, data analysis, and traffic management. Its core goal is to achieve randomization, representative distribution, and control of traffic through a scientific diversion mechanism.

[0060] Experimental layer domain structure: An experimental domain has multiple experimental layers. Within each experimental layer, users are aggregated into buckets according to certain rules. For example, according to a certain hash algorithm, users are assigned to a fixed number of buckets, and the number of users in each bucket conforms to a normal distribution with a mean of M / N.

[0061] Traffic sampling layer: This refers to the technical layer used to extract multiple buckets as traffic pools for implementing AB group policies. If an AB experiment is running within the traffic sampling layer, the running AB experiment will occupy this part of the traffic. New AB experiment sampling can only use the remaining traffic. For example, the corresponding number of buckets are extracted from the remaining user traffic in a certain experimental domain or experimental layer according to the traffic ratio.

[0062] Sampling refers to the process of selecting a subset of individuals or samples from a population for research or investigation. The characteristics and properties of the population are inferred by studying and analyzing the selected samples. Common sampling methods include random sampling, systematic sampling, stratified sampling, and cluster sampling. Random sampling involves randomly selecting buckets from a flow sampling stratum with equal probability, without replacement. Each bucket within the remaining flow has an equal chance of being selected into the experimental or control group. Systematic sampling involves selecting individuals at a certain interval as samples according to predetermined rules. For example, if the population of subjects is randomly numbered from 1 to 100, and only subjects with the last digit of the number being 7 are sampled, i.e., samples numbered 7, 17, 27, etc., there is essentially a fixed interval of 10 between adjacent numbers.

[0063] False positive rate: refers to the probability that the experimental group is actually invalid but the statistical test mistakenly judges it to be valid. It is a type I error in statistics.

[0064] After introducing the technical terms involved in this application, the application scenarios and design concepts of this application are briefly introduced.

[0065] At present, the traditional sampling method is: (1) All users are randomly divided into buckets, and the number of users in each bucket follows a normal distribution with a mean of M / N, where M is the total number of users and N is the total number of buckets. Under normal circumstances, M is more than a hundred times N, and some buckets have more users and some buckets have fewer users; (2) Sampling is performed based on the traffic proportion of each experimental group. For example, according to the traffic proportion of k% vs k%, k%*N buckets are randomly sampled from the two experimental groups as the user traffic of the two experimental groups. The above sampling method is simulated a lot, and the total user volume distribution obtained each time is as follows Figure 1 As shown in the figure, the total user activity distribution is as follows Figure 2 As shown in the figure, when the sampling results of the two experimental groups fall far apart in the distribution of the total number of users, the difference in the number of users and the difference in user activity in the two experimental groups are large, which brings a large noise impact to the subsequent AB experimental data analysis.

[0066] To this end, in this application, after determining the traffic bucket corresponding to the current AB experimental task as the locked bucket from each traffic bucket of the traffic sampling layer, the traffic contribution corresponding to each locked bucket determined based on the current user volume and current user activity is obtained, and based on the traffic contribution of each locked bucket, a cumulative traffic contribution curve is generated, and the inflection point detection is performed on the cumulative traffic contribution curve. After obtaining the inflection points of each traffic contribution, each locked bucket is stratified based on the inflection points of each traffic contribution to obtain each locked bucket sampling layer, and locked buckets are extracted for each experimental group from each locked bucket sampling layer to obtain the target user traffic of each experimental group. In this way, by determining the traffic contribution of each locked bucket based on the current user volume and current user activity, and generating a cumulative traffic contribution curve based on the traffic contribution of each locked bucket, the mutation point of the traffic contribution can be obtained after performing inflection point detection on the cumulative traffic contribution curve. Therefore, after stratifying each locked bucket using the mutation point, locked bucket sampling layers with different traffic contribution intervals can be obtained. Subsequently, after extracting locked buckets for each experimental group from each locked bucket sampling layer, the user volume and user activity of the target user traffic of each experimental group can be made more uniform, thereby reducing the false positive rate of the AB experiment and improving the experimental effect of the AB experiment.

[0067] After introducing the application scenarios and design ideas of this application, the technical solutions provided by this application are described in detail below.

[0068] The present application provides an AB experiment sampling method for use in electronic devices such as computers and servers. Figure 3 As shown, the general process of the AB experimental sampling method provided in the embodiment of the present application is as follows:

[0069] Step 301: Determine the traffic bucket corresponding to the current AB experiment task from each traffic bucket of the traffic sampling layer as a locked bucket.

[0070] In the implementation of this application, each traffic bucket of the traffic sampling layer is a measurement indicator of different dimensions such as experimental conflict intensity, experimental isolation requirements and user traffic characteristics. From the various experimental layers contained in the experimental domain corresponding to the current experimental type (such as UI, algorithm, etc.), the sampling system determines all the traffic buckets contained in the experimental layer where the traffic sampling layer acts. Among them, the various traffic buckets contained in the experimental domain corresponding to the current experimental type are divided in a random manner. For example, all users are randomly divided into multiple traffic buckets, and the number of users in each traffic bucket as a whole obeys the average value of M / N. Figure 4 In the normal distribution shown, M and N are positive integers, M is the total number of users, and N is the total number of traffic buckets.

[0071] When there are multiple AB experimental tasks during the sampling process, such as Figure 5 As shown, the traditional sampling method is to perform sequential sampling for each AB experimental task in turn, that is, after extracting traffic buckets as target user traffic for each experimental group of the current AB experimental task, sampling operations can be performed for each experimental group of the next AB experimental task. Therefore, during the sampling process, when sampling operations are performed for each experimental group of the current AB experimental task, the remaining traffic buckets in the traffic sampling layer need to be locked, resulting in the AB experimental tasks waiting in queue. For this reason, in the embodiment of the present application, Figure 6 As shown, for the first AB experimental task, first randomly sample and lock N*k%*K traffic buckets (N is the total number of traffic buckets, k% is the proportion of experimental group traffic, and K is the number of experimental groups) from the remaining traffic buckets in the traffic sampling layer as locked buckets, but do not perform sampling operations first, and continue to randomly extract locked buckets for each subsequent AB experimental task, and also do not perform sampling operations first. In this way, since the speed of locking N*k%*K traffic buckets is very fast, by locking the traffic buckets required by multiple experimental groups of each AB experimental task, even if there are multiple AB experimental tasks for sequential sampling, there will be no queuing problem. In other words, in the embodiment of the present application, when the sampling system determines the traffic bucket corresponding to the current AB experimental task from each traffic bucket of the traffic sampling layer as the locked bucket, it can adopt but is not limited to the following methods:

[0072] Based on the number of experimental groups and the percentage of experimental group traffic in the current AB experiment task, randomly select traffic from each traffic bucket within the traffic sampling layer that is not locked by other AB experiment tasks as the locked bucket for the current AB experiment task. For example, if all users are randomly divided into N traffic buckets, and the current AB experiment requires two experimental groups with k% and k% traffic, each experimental group requires N*k% buckets, for a total of two experimental groups. Therefore, N*k%*2 traffic buckets must be randomly selected from each traffic bucket within the traffic sampling layer that is not locked by other AB experiment tasks as the locked buckets for the current AB experiment task.

[0073] Step 302: Obtain the traffic contribution of each locked bucket; wherein the traffic contribution of each locked bucket is determined based on the current number of users of the locked bucket and the current user activity.

[0074] In the embodiment of the present application, the traffic contribution of each locked bucket can be obtained by, but is not limited to, the following methods:

[0075] First, each traffic index value of each locked bucket is obtained; wherein each traffic index value of the locked bucket is predetermined based on the current number of users and current user activity of the locked bucket.

[0076] In the specific implementation, in order to solve the traffic problems in different scenarios, based on the scenario characteristics and the current number of users and current user activity of each traffic bucket, the various traffic indicators of each traffic bucket are dynamically configured, and the traffic indicator values ​​corresponding to each traffic indicator are calculated.

[0077] For example, in an AB experiment sampling scenario, in order to balance user traffic, at least the following eight traffic indicators are included:

[0078] Traffic indicator 1: the number of users entering each experimental group: the number of deduplicated users entering each experimental group;

[0079] Traffic indicator 2: Total active days: the sum of the active days of users entering each experimental group;

[0080] Traffic indicator 3: Number of low-activity users: the number of deduplicated low-activity users in the scenario;

[0081] Traffic indicator 4: Active days of low-activity users: the sum of the active days of low-activity users in the scenario;

[0082] Traffic indicator 5: Number of medium active users: the number of users with medium activity in the scenario without duplicates;

[0083] Traffic indicator 6: Active days of moderately active users: the sum of the active days of moderately active users in the scenario;

[0084] Traffic indicator 7: Number of highly active users: The number of highly active users in the scenario without duplicates;

[0085] Traffic indicator 8: Active days of highly active users: the sum of the active days of highly active users in the scenario.

[0086] In an embodiment of the present application, the user classification method based on the activity threshold is as follows: users whose average daily usage time is less than a first time threshold (for example, 5 minutes) or whose weekly visits are less than or equal to a first number threshold (for example, 2 times) are low-activity users; users whose average daily usage time is within a first time range (for example, 5-30 minutes) and whose weekly visits are within a first number range (for example, 3-7 times) are medium-activity users; users whose average daily usage time is greater than a second time threshold (for example, 30 minutes) and whose daily visits are greater than or equal to a second number threshold (for example, 1 time) are high-activity users. The activity threshold can be recalibrated when the change in MAU / DAU is within a set range (for example, ±20%), and can also be configured differently according to the scenario. For example, during an e-commerce promotion, the classification standard for high-activity users is increased by a set step size (for example, by 50%), that is, the activity threshold is increased by a set step size (for example, by 50%). Among them, MAU / DAU = daily active users (deduplicated) / monthly active users (deduplicated); MAU ​​is the total number of unique users who are active at least once within the statistical period (for example, 30 days); DAU is the total number of unique users who are active at least once within a single day (24 hours).

[0087] In specific implementation, each flow index value of each flow bucket is pre-calculated and stored in the database in a specific format based on the experimental domain identifier, experimental layer identifier, and flow bucket identifier. Thus, based on the experimental domain identifier, experimental layer identifier, and flow bucket identifier corresponding to the current AB experiment task, each flow index value of each locked bucket can be directly obtained from the database, thereby saving calculation time and improving sampling efficiency. Specifically, each flow index value of each flow bucket in the database is stored in the form of a flow index table, where the flow index table is shown in Table 1:

[0088] Table 1.

[0089]

[0090] Then, based on the traffic index values ​​of each locked bucket, the traffic contribution of each locked bucket is determined.

[0091] In the embodiment of the present application, based on the traffic index values ​​of each locked bucket, the traffic contribution of each locked bucket can be determined by, but not limited to, the following methods:

[0092] First, for each flow index, based on the flow index value of each locked bucket under the flow index, the locked buckets are sorted to obtain the rank of each locked bucket under the flow index.

[0093] In a specific implementation, for each traffic index, based on the traffic index value of each locked bucket under the traffic index, when sorting each locked bucket to obtain the rank of each locked bucket under the traffic index, any of the following methods may be used, but not limited to:

[0094] The first method: For each traffic index, based on the traffic index value of each locked bucket under the traffic index, the competitive ranking method is used to sort the locked buckets to obtain the rank of each locked bucket under the traffic index, that is:

[0095] For each flow index, the lock buckets are sorted in descending order based on their flow index values ​​under the flow index. If multiple lock buckets have the same flow index value under the flow index, they are assigned the same rank. The rank of the next lock bucket skips the positions occupied by all parallel lock buckets, thereby obtaining the rank of each lock bucket under the flow index.

[0096] This scheme is robust when there are outliers in the flow metric values ​​of the locked bucket.

[0097] The second method: For each traffic indicator, based on the traffic indicator value of each locked bucket under the traffic indicator, a competitive ranking method is used to sort each locked bucket. Then, a weighted significance strategy and a traffic contribution correction strategy are introduced. That is:

[0098] First, for each flow index, based on the flow index value of each locking bucket under the flow index, a competitive ranking method is used to perform weighted significance sorting on each locking bucket to obtain the initial rank of each locking bucket under the flow index. Specifically, for each flow index, based on the product of the flow index value of each locked bucket under the flow index and the confidence coefficient of the flow index, a competitive ranking method is used to arrange the locked buckets in descending order. If the flow index values ​​of multiple locked buckets under the flow index and the product of the confidence coefficient of the flow index are the same, they are assigned the same initial rank, and the initial rank of the next locked bucket skips the positions occupied by all parallel locked buckets, thereby obtaining the initial rank of each locked bucket under the flow index; wherein, the confidence coefficient of the flow index is positively proportional to the importance of the flow index, for example, the confidence coefficient of a flow index with higher importance is a value greater than 1, such as any value between 1.05 and 1.1, the confidence coefficient of a flow index with medium importance is 1, such as any value between 0.95 and 0.99, and the confidence coefficient of a flow index with lower importance is a value less than 1, such as any value between 0.95 and 0.99.

[0099] Then, for each flow index, based on the bucket flow contribution factor of each locked bucket, the initial rank of each locked bucket under the flow index is corrected to obtain the rank of each locked bucket under the flow index; wherein, the bucket flow contribution factor of the locked bucket is calculated as the ratio of the bucket flow of the locked bucket to the total flow of all locked buckets. Specifically, for each flow index, based on the bucket flow contribution factor of each locked bucket, the following formula is used to correct the initial rank of each locked bucket under the flow index to obtain the rank of each locked bucket under the flow index:

[0100] Rank = initial rank × (1 + bucket flow contribution factor)

[0101] This scheme is robust when there are outliers in the traffic metric values ​​of the locked bucket, and at the same time, it reduces the ranking jitter caused by sampling fluctuations.

[0102] Then, for each locked bucket, the ranks of the locked bucket under each flow index are summed to obtain the sum of the ranks of the locked bucket under each flow index as the flow contribution of the locked bucket. In specific implementation, the flow contribution of each locked bucket can be calculated using the following formula:

[0103]

[0104] Wherein, j represents the number of the locked bucket, and j is a positive integer greater than 1 and less than or equal to n; i represents the number of the flow indicator, and i is a positive integer greater than 1 and less than or equal to n; Indicates the traffic contribution of the locked bucket numbered j; Indicates the flow index value of the locked bucket numbered j under the flow index numbered i.

[0105] Step 303: Generate a cumulative traffic contribution curve based on the traffic contribution of each locked bucket.

[0106] In the embodiment of the present application, when generating a cumulative traffic contribution curve based on the traffic contribution of each locked bucket, the following methods may be used, but are not limited to:

[0107] First, the traffic contribution of each locked bucket is sorted in descending order to construct a traffic contribution sequence.

[0108] Then, the first traffic contribution in the traffic contribution sequence is determined as the first data point, and the sum of the i-th traffic contribution in the traffic contribution sequence and each traffic contribution before the i-th traffic contribution is determined as a data point in turn to construct a data point sequence; wherein, i represents the number of the traffic indicator, and i is a positive integer greater than 1 and less than or equal to n, and n is the total number of locked buckets.

[0109] Finally, based on the data point sequence, a cumulative traffic contribution curve is generated.

[0110] Step 304: Perform inflection point detection on the cumulative flow contribution curve to obtain inflection points of each flow contribution.

[0111] In the embodiment of the present application, when performing inflection point detection on the cumulative traffic contribution curve to obtain the inflection points of each traffic contribution, the following methods may be used, but are not limited to:

[0112] Step 3041: cyclically perform an inflection point detection operation on the cumulative traffic contribution curve; wherein, the inflection point detection operation is: performing an inflection point detection on the cumulative traffic contribution curve to be detected to obtain a candidate traffic contribution inflection point; when the inflection point detection operation is performed for the first time, the cumulative traffic contribution curve to be detected is the cumulative traffic contribution curve; when the inflection point detection operation is not performed for the first time, the cumulative traffic contribution curve to be detected is the two cumulative traffic contribution curve segments segmented when the inflection point detection operation was performed last time; with the candidate traffic contribution inflection point as the segmentation point, the cumulative traffic contribution curve to be detected is divided into two cumulative traffic contribution curve segments.

[0113] Step 3042: Determine whether the termination condition is met; if so, execute step 3043; if not, return to step 3041; wherein, the termination condition may be but is not limited to that the total number of candidate traffic contribution inflection points is not less than a set threshold.

[0114] Step 3043: Determine the candidate flow contribution inflection points obtained by each execution of the inflection point detection operation as individual flow contribution inflection points.

[0115] Step 305: Based on the inflection points of the traffic contributions, the locked buckets are stratified to obtain sampling layers of the locked buckets.

[0116] In specific implementation, based on the inflection points of each traffic contribution, each locked bucket is stratified to obtain the sampling layer of each locked bucket, which can be adopted but not limited to the following methods:

[0117] First, each flow contribution interval is determined based on every two adjacent flow contribution inflection points among each flow contribution inflection point.

[0118] Then, based on the traffic contribution of each locked bucket, each locked bucket is divided into a matching traffic contribution interval to obtain each locked bucket sampling layer.

[0119] For example, see Figure 7As shown, assuming that there are flow contribution inflection point 1, flow contribution inflection point 2 and flow contribution inflection point 3, then from 0 to flow contribution inflection point 2 is used as a locked bucket sampling layer, from flow contribution inflection point 2 to flow contribution inflection point 1 is used as a locked bucket sampling layer, from flow contribution inflection point 1 to flow contribution inflection point 3 is used as a locked bucket sampling layer, and from flow contribution inflection point 3 to the last point is used as a locked bucket sampling layer.

[0120] Step 306: Extract locked buckets for each experimental group from each locked bucket sampling layer to obtain target user traffic for each experimental group.

[0121] In the embodiment of the present application, a locked bucket is extracted from each locked bucket sampling layer for each experimental group to obtain the target user traffic of each experimental group. Figure 8 As shown, the following methods may be used but are not limited to:

[0122] First, for each locked bucket sampling layer, according to the traffic proportion of the experimental group, a locked bucket is randomly selected for each experimental group from the locked bucket sampling layer.

[0123] Then, for each experimental group, the union of the locked buckets randomly selected for the experimental group in each locked bucket sampling layer is determined as the target user traffic of the experimental group.

[0124] The effect of the above scheme is verified. The verification scheme is: simulate 1000 sampling k% vs k% traffic AB experiments. At this time, the AB traffic strategies are consistent and there is no deviation in theory. The random sampling scheme and the sampling scheme provided by the embodiment of the present application are respectively used to analyze the false positive rate of the traffic indicator. The verification scheme can be verified from two aspects. On the one hand, it is verified from the user volume and activity of the two groups AB. On the other hand, it is seen whether it is better when decomposing in some important dimensions (avoiding overall improvement leading to local deterioration). The specific effects are as follows: Figure 9 and Figure 10 As shown, it can be seen that by implementing the sampling scheme provided in the embodiment of the present application, the false positive rate of the number of users and activity of the two groups AB diversion has been greatly reduced, and in different dimensions, such as Figure 10 In the three dimensions of L1, L2 and L3, the false positive rate of user volume and activity has also decreased significantly.

[0125] Based on the above embodiments, the present application provides an AB experiment sampling device, see Figure 11 As shown, the AB experiment sampling device 400 provided in the embodiment of the present application includes at least:

[0126] The traffic locking unit 401 is used to determine the traffic bucket corresponding to the current AB experiment task from each traffic bucket of the traffic sampling layer as a locked bucket;

[0127] Contribution acquisition unit 402, configured to acquire the traffic contribution of each locked bucket; wherein the traffic contribution of each locked bucket is determined based on the current number of users of the locked bucket and the current user activity;

[0128] The curve generating unit 403 is used to generate a cumulative traffic contribution curve based on the traffic contribution of each locked bucket;

[0129] An inflection point determination unit 404 is configured to perform inflection point detection on the cumulative flow contribution curve to obtain inflection points of respective flow contribution degrees;

[0130] The traffic stratification unit 405 is used to stratify each locked bucket based on each traffic contribution inflection point to obtain each locked bucket sampling layer;

[0131] The traffic extraction unit 406 is configured to extract a locked bucket for each experimental group from each locked bucket sampling layer to obtain the target user traffic of each experimental group.

[0132] In one possible implementation, the traffic locking unit 401 is used to randomly extract traffic from each traffic bucket that has not been locked by other AB experimental tasks in the traffic sampling layer based on the number of experimental groups and the traffic proportion of the experimental groups of the current AB experimental task, and use it as the locking bucket for the current AB experimental task.

[0133] In one possible implementation, the contribution acquisition unit 402 is used to obtain the various traffic indicator values ​​of each locked bucket; wherein the various traffic indicator values ​​of the locked bucket are predetermined based on the current number of users of the locked bucket and the current user activity; based on the various traffic indicator values ​​of each locked bucket, the traffic contribution of each locked bucket is determined.

[0134] In one possible implementation, the curve generating unit 403 is used to sort the traffic contribution of each locked bucket in descending order to construct a traffic contribution sequence; determine the first traffic contribution in the traffic contribution sequence as the first data point, and sequentially determine the sum of the i-th traffic contribution in the traffic contribution sequence and each traffic contribution before the i-th traffic contribution as a data point to construct a data point sequence; wherein i is a positive integer greater than 1 and less than or equal to n, and n is the total number of locked buckets; and generate a cumulative traffic contribution curve based on the data point sequence.

[0135] In one possible implementation, the inflection point determination unit 404 is used to cyclically perform an inflection point detection operation on the cumulative traffic contribution curve until it is determined that the termination condition is met, and the candidate traffic contribution inflection point obtained by each execution of the inflection point detection operation is determined as each traffic contribution inflection point; wherein, the inflection point detection operation is: performing inflection point detection on the cumulative traffic contribution curve to be detected to obtain a candidate traffic contribution inflection point; wherein, when the inflection point detection operation is performed for the first time, the cumulative traffic contribution curve to be detected is the cumulative traffic contribution curve; when the inflection point detection operation is not performed for the first time, the cumulative traffic contribution curve to be detected is the two cumulative traffic contribution curve segments segmented when the inflection point detection operation was performed last time; with the candidate traffic contribution inflection point as the segmentation point, the cumulative traffic contribution curve to be detected is divided into two cumulative traffic contribution curve segments.

[0136] In one possible implementation, the traffic stratification unit 405 is used to determine each traffic contribution interval based on each two adjacent traffic contribution inflection points in each traffic contribution inflection point; based on the traffic contribution of each locked bucket, each locked bucket is divided into a matching traffic contribution interval to obtain each locked bucket sampling layer.

[0137] In one possible implementation, the traffic extraction unit 406 is used to randomly extract locked buckets for each experimental group from the locked bucket sampling layer according to the traffic proportion of the experimental group for each locked bucket sampling layer; for each experimental group, the union of the locked buckets randomly extracted for the experimental group in each locked bucket sampling layer is determined as the target user traffic of the experimental group.

[0138] It should be noted that the principle of solving technical problems by the AB experimental sampling device 400 provided in the embodiment of the present application is similar to the AB experimental sampling method provided in the embodiment of the present application. Therefore, the implementation of the AB experimental sampling device 400 provided in the embodiment of the present application can refer to the implementation of the AB experimental sampling method provided in the embodiment of the present application, and the repeated parts will not be repeated.

[0139] After introducing the AB experiment sampling method and device provided in the embodiments of the present application, the electronic device provided in the embodiments of the present application is briefly introduced next.

[0140] See Figure 12 As shown, the electronic device 500 provided in the embodiment of the present application includes at least a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the processor 501 executes the computer program, the above-mentioned AB experimental sampling method provided in the embodiment of the present application is implemented.

[0141] The electronic device 500 provided in the embodiment of the present application may further include a bus 503 connecting different components (including the processor 501 and the memory 502). The bus 503 represents one or more of several types of bus structures, including a memory bus, a peripheral bus, a local bus, etc.

[0142] Memory 502 may include a readable storage medium in the form of volatile memory, such as random access memory (RAM) 5021 and / or cache memory 5022, and may further include read-only memory (ROM) 5023. Memory 502 may also include a program tool 5025 having a set (at least one) of program modules 5024. Program modules 5024 include, but are not limited to, an operating subsystem, one or more application programs, other program modules, and program data. Each of these examples, or some combination thereof, may include an implementation of a network environment.

[0143] Processor 501 can be a single processing element or a collective term for multiple processing elements. For example, processor 501 can be a central processing unit (CPU) or one or more integrated circuits configured to implement the above-mentioned AB experiment sampling method provided in the embodiments of the present application. Specifically, processor 501 can be a general-purpose processor, including but not limited to a CPU, an application-specific integrated circuit (ASIC), an off-the-shelf field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0144] The electronic device 500 can communicate with one or more external devices 504 (e.g., keyboard, remote control, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 500 (e.g., mobile phone, computer, etc.), and / or communicate with a device that enables the electronic device 500 to communicate with one or more other electronic devices 500 (e.g., router, modem, etc.). Such communication can be performed through an input / output (I / O) interface 505. In addition, the electronic device 500 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN) and / or public network, such as the Internet) through a network adapter 506. Figure 12As shown, the network adapter 506 communicates with other modules of the electronic device 500 via the bus 503. Figure 12 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 500, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, disk arrays (Redundant Arrays of Independent Disks, RAID) subsystems, tape drives, and data backup storage subsystems.

[0145] It should be noted that Figure 12 The electronic device 500 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0146] The following describes the computer-readable storage medium provided in the embodiments of the present application. The computer-readable storage medium provided in the embodiments of the present application stores computer instructions that, when executed by a processor, implement the aforementioned AB experimental sampling method provided in the embodiments of the present application. Specifically, the computer instructions may be built into or installed in the processor, so that the processor implements the aforementioned AB experimental sampling method provided in the embodiments of the present application by executing the built-in or installed computer instructions.

[0147] In addition, the above-mentioned AB experimental sampling method provided in the embodiment of the present application can also be implemented as a computer program product, which includes program code. When the program code is run on a processor, it implements the above-mentioned AB experimental sampling method provided in the embodiment of the present application.

[0148] The computer program product provided in the embodiments of the present application may adopt one or more computer-readable storage media, and the computer-readable storage medium may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. Specifically, more specific examples of computer-readable storage media (a non-exhaustive list) include an electrical connection with one or more wires, a portable disk, a hard disk, RAM, ROM, Erasable Programmable Read Only Memory (EPROM), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.

[0149] The computer program product provided in the embodiments of the present application may be a CD-ROM and include program code, and may also be run on electronic devices such as servers and computers. However, the computer program product provided in the embodiments of the present application is not limited thereto. In the embodiments of the present application, the computer-readable storage medium may be any tangible medium that contains or stores program code, and the program code may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0150] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.

[0151] Furthermore, although the operations of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0152] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0153] Obviously, those skilled in the art may make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the embodiments of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalents, the present application also intends to include such modifications and variations.

Claims

1. An AB experiment sampling method, characterized in that: include: Determine the traffic bucket corresponding to the current AB experiment task from each traffic bucket of the traffic sampling layer as the locked bucket; Obtaining a traffic contribution of each locked bucket; wherein the traffic contribution of each locked bucket is determined based on the current number of users of the locked bucket and the current user activity; Generating a cumulative traffic contribution curve based on the traffic contribution of each locked bucket; Performing inflection point detection on the cumulative flow contribution curve to obtain inflection points of each flow contribution; Based on the inflection points of the traffic contribution, the locked buckets are stratified to obtain sampling layers of the locked buckets; Extracting locked buckets for each experimental group from each locked bucket sampling layer to obtain target user traffic for each experimental group; Generating a cumulative traffic contribution curve based on the traffic contribution of each locked bucket includes: Sorting the traffic contribution of each locked bucket in descending order to construct a traffic contribution sequence; Determining the first traffic contribution in the traffic contribution sequence as the first data point, and sequentially determining the sum of the i-th traffic contribution in the traffic contribution sequence and each traffic contribution before the i-th traffic contribution as a data point to construct a data point sequence; wherein i is a positive integer greater than 1 and less than or equal to n, and n is the total number of the locked buckets; Based on the data point sequence, a cumulative flow contribution curve is generated.

2. The AB experiment sampling method according to claim 1, characterized in that: Determine the traffic bucket corresponding to the current AB experiment task from each traffic bucket in the traffic sampling layer as the locked bucket, including: Based on the number of experimental groups and the proportion of experimental group traffic of the current AB experimental task, traffic is randomly sampled from each traffic bucket that is not locked by other AB experimental tasks in the traffic sampling layer as the locked bucket of the current AB experimental task.

3. The AB experiment sampling method according to claim 1, characterized in that: Obtaining the traffic contribution of each locked bucket includes: Obtaining each traffic index value of each locked bucket; wherein each traffic index value of the locked bucket is predetermined based on the current number of users and current user activity of the locked bucket; The traffic contribution of each locked bucket is determined based on the traffic indicator value of each locked bucket.

4. The AB experiment sampling method according to claim 1, wherein: Perform inflection point detection on the cumulative flow contribution curve to obtain inflection points of each flow contribution, including: The inflection point detection operation is cyclically performed on the cumulative flow contribution curve until a termination condition is determined to be satisfied, and the candidate flow contribution inflection points obtained by each execution of the inflection point detection operation are determined as the respective flow contribution inflection points; wherein the inflection point detection operation is: Perform inflection point detection on the cumulative flow contribution curve to be detected to obtain a candidate flow contribution inflection point; wherein, when the inflection point detection operation is performed for the first time, the cumulative flow contribution curve to be detected is the cumulative flow contribution curve; when the inflection point detection operation is not performed for the first time, the cumulative flow contribution curve to be detected is the two cumulative flow contribution curve segments segmented when the inflection point detection operation was performed last time; The candidate flow contribution inflection point is used as a dividing point to divide the to-be-detected cumulative flow contribution curve into two cumulative flow contribution curve segments.

5. The AB experiment sampling method according to claim 1, characterized in that: Based on the inflection points of the traffic contribution, the locked buckets are stratified to obtain sampling layers of the locked buckets, including: Determining each flow contribution interval based on each two adjacent flow contribution inflection points among the flow contribution inflection points; Based on the traffic contribution of each locked bucket, each locked bucket is divided into a matching traffic contribution interval to obtain each locked bucket sampling layer.

6. The AB experiment sampling method according to any one of claims 1 to 5, characterized in that: Extracting a locked bucket for each experimental group from each locked bucket sampling layer to obtain target user traffic for each experimental group includes: For each of the locked bucket sampling layers, randomly select a locked bucket for each of the experimental groups from the locked bucket sampling layer according to the traffic proportion of the experimental group; For each of the experimental groups, the union of the locked buckets randomly selected for the experimental group in each of the locked bucket sampling layers is determined as the target user traffic of the experimental group.

7. An AB experiment sampling device, characterized in that: include: The traffic locking unit is used to determine the traffic bucket corresponding to the current AB experiment task from each traffic bucket of the traffic sampling layer as the locking bucket; A contribution acquisition unit, configured to acquire the traffic contribution of each locked bucket; wherein the traffic contribution of each locked bucket is determined based on the current number of users of the locked bucket and the current user activity; a curve generating unit, configured to generate a cumulative flow contribution curve based on the flow contribution of each locked bucket; an inflection point determination unit, configured to perform inflection point detection on the cumulative flow contribution curve to obtain inflection points of respective flow contribution degrees; a traffic stratification unit, configured to stratify each of the locked buckets based on each of the traffic contribution inflection points to obtain each locked bucket sampling layer; a traffic extraction unit, configured to extract a locked bucket for each experimental group from each locked bucket sampling layer, to obtain target user traffic of each experimental group; Among them, the curve generating unit is used to sort the traffic contribution of each locked bucket in order from high to low to construct a traffic contribution sequence; determine the first traffic contribution in the traffic contribution sequence as the first data point, and sequentially determine the sum of the i-th traffic contribution in the traffic contribution sequence and each traffic contribution before the i-th traffic contribution as a data point to construct a data point sequence; based on the data point sequence, generate a cumulative traffic contribution curve; wherein i is a positive integer greater than 1 and less than or equal to n, and n is the total number of the locked buckets.

8. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the AB experiment sampling method according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the AB experiment sampling method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Flow layer setting method and apparatus based on flow experiment and flow experiment implementing method and apparatus

    CN105243006A

  • Sampling method and device, electronic equipment and storage medium

    CN117493422A