A data processing method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202210511906.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-05-11
AI Technical Summary
[0006]为了解决目前小流量测试结果准确度低的问题,本申请提供了一种数据处理方法、装置、电子设备及存储介质:
[0022]本申请实施例通过确定对照组指标数据和参照组指标数据之间的第一显著水平值,以及确定对照组指标数据和实验组指标数据之间的第二显著水平值,其中第一显著水平值可以反映指标数据受波动影响的程度,第二显著水平值可以反映迭代方案相较于当前方案的差异程度,在第一显著水平值反映当前未受波动影响时,第二显著水平值的置信程度高,如此,可以解决小流量下在线测试指标波动大难以准确判断实验效果的问题。本申请可以广泛应用于工业推荐系统的在线评估中,可以减少小流量指标提升很多,扩流后大盘并没有对应提升的情况,增加大流量和小流量的实验结果一致性,从而可以有效的提升迭代效率。
Smart Images

Figure CN117097789B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to a data processing method, apparatus, electronic device and storage medium. Background Technology
[0002] Recommendation systems, in the internet age, refer to platforms automatically selecting or matching products based on user interests and presenting them to users. Current recommendation systems typically employ online testing with low traffic volume, iterating on recommendation algorithms and strategies by comparing performance metrics over a period of time. Once the results of these low-traffic iterations are deemed reliable, the algorithm (strategy) with the highest online metrics is selected for full-traffic deployment. Therefore, obtaining reliable iterative results under low-traffic conditions is crucial for implementing algorithm iterations across various scenarios.
[0003] Currently, online experiments with recommendation systems typically divide the total 100% traffic into traffic buckets, the smallest traffic units with a granularity of 1%. For example, in a recommendation scenario with 50,000 daily active users, each 1% traffic bucket contains 500 users.
[0004] However, in scenarios with relatively low traffic (tens of thousands of users), larger traffic buckets (e.g., 20% or more) or longer time periods are often used to obtain more reliable small-traffic experimental results. This limits the number of concurrent online experiments to less than four, significantly restricting iteration efficiency. Increasing the observation period also reduces iteration efficiency. Furthermore, expanding the traffic and increasing the experimental period still requires determining what percentage and duration are considered reliable, currently relying primarily on manual experience, lacking objectivity and persuasiveness. In addition, the small traffic buckets in small-traffic scenarios often result in highly volatile online metrics, making it difficult to determine whether improvements or declines are due to differences in the iterative algorithm or fluctuations. This can lead to incorrect conclusions regarding the iteration direction, wasting excessive effort on the wrong path. Moreover, due to the influence of fluctuations, the overall improvement after expanding the traffic may be significantly lower than the results during small-traffic testing, resulting in unsatisfactory outcomes.
[0005] Therefore, it is necessary to improve the accuracy of current low-traffic test results in order to properly guide the algorithm iteration. Summary of the Invention
[0006] To address the issue of low accuracy in current low-flow test results, this application provides a data processing method, apparatus, electronic device, and storage medium:
[0007] According to a first aspect of this application, a data processing method is provided, comprising:
[0008] Obtain indicator data for the reference group, experimental group, and control group; the reference group indicator data is obtained from the behavior logs of test subjects using the current solution; the experimental group indicator data is obtained from the behavior logs of test subjects using the solution to be verified; the solution to be verified is an iterative solution of the current solution; the control group indicator data is obtained from the behavior logs of historical subjects using the current solution.
[0009] Based on the indicator data of the control group and the reference group, the first significance level value was determined; the first significance level value characterizes the degree of fluctuation of the indicator data of the reference group.
[0010] Based on the indicator data of the control group and the experimental group, the second significance level value was determined; the second significance level value characterizes the degree of difference between the scheme to be validated and the current scheme.
[0011] When the first significance level value is less than or equal to the first preset value and the second significance level value is greater than or equal to the second preset value, the iterative availability of the scheme to be tested is determined.
[0012] According to a second aspect of this application, a data processing apparatus is provided, the apparatus comprising:
[0013] The first acquisition module is used to acquire reference group indicator data and experimental group indicator data; the reference group indicator data is obtained from the behavior logs of test subjects adopting the current solution; the experimental group indicator data is obtained from the behavior logs of test subjects adopting the solution to be verified; the solution to be verified is the iterative solution of the current solution.
[0014] The second acquisition module is used to acquire control group indicator data; the control group indicator data is obtained from the behavior logs of historical objects using the current scheme.
[0015] The first determination module is used to determine the first significance level value based on the indicator data of the control group and the indicator data of the reference group; the first significance level value characterizes the degree of fluctuation of the indicator data of the reference group.
[0016] The second determination module is used to determine the second significance level value based on the indicator data of the control group and the indicator data of the experimental group; the second significance level value characterizes the degree of difference between the scheme to be verified and the current scheme;
[0017] The third determination module is used to determine the iterative availability of the scheme to be tested when the first significance level value is less than or equal to the first preset value and the second significance level value is greater than or equal to the second preset value.
[0018] According to a third aspect of this application, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the data processing method of the first aspect of this application.
[0019] According to a fourth aspect of this application, a computer storage medium is provided, which stores at least one instruction or at least one program, wherein the at least one instruction or at least one program is loaded and executed by a processor to implement the data processing method of the first aspect of this application.
[0020] According to a fifth aspect of this application, a computer program product is provided, comprising at least one instruction or at least one program segment, wherein the at least one instruction or at least one program segment is loaded and executed by a processor to implement the data processing method of the first aspect of this application.
[0021] The data processing method, apparatus, electronic device, and storage medium provided in this application have the following technical advantages:
[0022] This application's embodiments determine a first significance level between the indicator data of the control group and the reference group, and a second significance level between the indicator data of the control group and the experimental group. The first significance level reflects the degree to which the indicator data is affected by fluctuations, while the second significance level reflects the degree of difference between the iterative scheme and the current scheme. When the first significance level reflects that the current scheme is not affected by fluctuations, the confidence level of the second significance level is high. This solves the problem of large fluctuations in online test indicators under low-volume conditions, making it difficult to accurately judge the experimental effect. This application can be widely applied to the online evaluation of industrial recommendation systems, reducing situations where indicators improve significantly under low-volume conditions but do not show a corresponding improvement in the overall performance after expanding the volume. It increases the consistency of experimental results between high-volume and low-volume conditions, thereby effectively improving iteration efficiency. Attached Figure Description
[0023] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application;
[0025] Figure 2 This is a flowchart illustrating a data processing method provided in an embodiment of this application;
[0026] Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application;
[0027] Figure 4 This is a schematic diagram of a process for determining a first significance level value provided in an embodiment of this application;
[0028] Figure 5 This is a schematic diagram of a process for determining a second significance level value provided in an embodiment of this application;
[0029] Figure 6 This is a schematic diagram of a front-end specified interface provided in an embodiment of this application;
[0030] Figure 7 This is a block diagram of a data processing apparatus provided in an embodiment of this application;
[0031] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0033] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification, claims and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products or devices.
[0034] In response to the situation in related technologies where online experimental metrics fluctuate greatly under low traffic, making it difficult to judge the experimental effect and thus determine whether to implement an iterative plan, this application provides a data processing method that can identify whether the metric data is affected by fluctuation factors. If it is determined that the data is not affected by fluctuation factors, the improvement effect of the iterative plan can be determined based on the metric data. This method can be widely applied to low traffic testing scenarios.
[0035] Please see Figure 1 , Figure 1This is a schematic diagram of an application environment provided in an embodiment of this application. This application environment may include a client 10 and a server 20. The client 10 and the server 20 can be directly or indirectly connected via wired or wireless communication. It should be noted that... Figure 1 This is just one example.
[0036] In this system, client 10 may have internet products installed, and server 20 may be a backend server providing services related to those internet products. The current solution refers to the algorithm or strategy executed by the internet product to achieve a specific function. This algorithm or strategy can be optimized iteratively to better achieve the corresponding function. The test object and historical object refer to the users corresponding to different clients 10.
[0037] In low-traffic testing scenarios, to determine the improvement effect of the iterative solution, server 20 pre-classifies users of client 10 into test subjects using the current solution or test subjects using the solution to be verified. This means that the algorithms executed by the internet products on client 10 used by different test subjects are different. Client 10 uploads the behavior logs of the test subjects to server 20. Server 20 obtains reference group indicator data based on the behavior logs of test subjects using the current solution, experimental group indicator data based on the behavior logs of test subjects using the solution to be verified (i.e., the iterative solution), and control group indicator data from the behavior logs of historical subjects using the current solution. Based on the control group indicator data and the reference group indicator data, server 20 can determine whether the current solution is affected by fluctuation factors. Based on the control group indicator data and the experimental group indicator data, server 20 can determine the efficiency improvement of the solution to be verified compared to the current solution. Server 20 considers the impact of the fluctuation factors and the efficiency improvement of the solution to be verified to determine whether to replace the current solution with the solution to be verified, thus implementing the solution iteration.
[0038] The aforementioned client 10 can be a physical device such as a smartphone, computer (e.g., desktop computer, tablet computer, laptop computer), augmented reality (AR) / virtual reality (VR) device, digital assistant, smart voice interaction device (e.g., smart speaker), smart wearable device, smart home appliance, in-vehicle terminal, etc., or it can be software running on the physical device, such as a computer program. The operating system corresponding to the client can be Android, iOS (a mobile operating system developed by Apple), Linux (an operating system), Microsoft Windows, etc.
[0039] The aforementioned server 20 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server may include network communication units, processors, and memory, etc. The server can provide backend services to corresponding clients.
[0040] It should be noted that, for user information and other data involved in the embodiments of this application, when the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0041] The following describes a specific embodiment of a data processing method according to this application. Figure 2 This is a flowchart illustrating a data processing method provided in an embodiment of this application. This application provides the operational steps of the method described in the embodiment or flowchart, but based on conventional or non-inventive methods, more or fewer operational steps may be included. The order of steps listed in the embodiment is merely one possible execution order among many and does not represent the only possible execution order. In actual systems or products, the methods can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment) as shown in the embodiment or the accompanying drawings.
[0042] Specific examples Figure 2 As shown, the method may include:
[0043] S201: Obtain indicator data for the reference group, experimental group, and control group.
[0044] The reference group index data is obtained from the behavior logs of test subjects using the current solution; the experimental group index data is obtained from the behavior logs of test subjects using the solution to be verified; the solution to be verified is an iterative solution of the current solution; and the control group index data is obtained from the behavior logs of historical subjects using the current solution.
[0045] In this embodiment, the current solution refers to an algorithm or strategy executed in a specified internet product to achieve a certain function. The specified internet product can be a cloud technology product, an artificial intelligence product, a smart transportation product, an assisted driving product, a live streaming product, an online office product, an e-commerce product, a game product, a local life product, an instant messaging product, a social product, etc. The function can be a recommendation function; correspondingly, the current solution can be a recall algorithm or a ranking algorithm. The test object can be a user using the internet product.
[0046] In this embodiment of the application, before iterating the current solution executed in the specified Internet product, the iterative solution is taken as the solution to be verified. The iterative availability of the solution to be verified is checked. When it is determined that the iterative availability of the solution to be verified is iterable, the current solution executed in the specified Internet product is iterated.
[0047] Currently, A / B testing is an important tool for internet companies to iterate products and improve user experience. In related technologies, at the same time, plan A is implemented for some users and plan B for another group. Then, test statistics are calculated for the user metrics of groups A and B respectively, and the results are used to determine whether there is a significant overall difference between the two groups. However, in current low-traffic testing scenarios, where the number of users in groups A and B is small, even if there is a significant overall difference in the user metrics of groups A and B, it is impossible to determine whether the difference is caused by the plan's implementation or by fluctuations, leading to inaccurate A / B test results.
[0048] Based on this, in this embodiment of the application, the current scheme is taken as scheme A and the scheme to be verified is taken as scheme B. In addition to comparing the significance of the difference between scheme A and scheme B, by comparing the index data of the two groups of test objects that adopted scheme A at different time dimensions, it is possible to identify whether the experimental index data is affected by fluctuation factors. In this way, the confidence level of the significance of the difference between scheme A and scheme B can be improved.
[0049] Therefore, in this embodiment, the server acquires three sets of indicator data, including reference group indicator data, experimental group indicator data, and control group indicator data. The reference group indicator data is obtained from the behavior logs of test subjects using the current solution; the experimental group indicator data is obtained from the behavior logs of test subjects using the solution to be verified. In subsequent steps, statistical analysis of the reference group and experimental group indicator data can determine whether there is a significant difference between the solution to be verified and the current solution. The control group indicator data is obtained from the behavior logs of historical users of the current solution. Historical users can be users who used the specified internet product before the test. In subsequent steps, statistical analysis of the control group and reference group indicator data determines the difference between the reference group and control group indicator data. Since both sets of indicator data use the current solution, the smaller the difference, the less affected the current solution is by fluctuation factors.
[0050] To obtain the behavior logs of test subjects employing different approaches, in some possible embodiments, the data processing method of this application embodiment further includes, as follows: Figure 3 The following steps are shown:
[0051] S301: Obtain the test object set; the test object set includes multiple test objects and the identifier of each test object.
[0052] The identifier for each test object may include a unique identity code representing the user.
[0053] S303: Based on the identifier of each test object, the test object set is divided into multiple test object subsets; each test object subset contains an equal number of test objects.
[0054] Specifically, in this step, based on the traffic bucketing principle in A / B testing, the number of test object subsets to be divided, i.e. the number of buckets, is first determined according to the number of all test objects in the test object set. Then, hash bucketing is performed based on each user's identity code, so that different users fall into different buckets, and the number of users in different buckets is equal.
[0055] S305: Determine a first preset number of test object subsets from multiple test object subsets, and use the test objects in the first preset number of test object subsets as test objects for adopting the current solution.
[0056] Specifically, in this step, a first preset number of test object subsets are used to test the current solution, that is, the corresponding number of traffic buckets are allocated to the current solution.
[0057] S307: Determine a second preset number of test object subsets from multiple test object subsets, and use the test objects in the second preset number of test object subsets as test objects for the scheme to be verified.
[0058] The first preset number of test object subsets and the second preset number of test object subsets do not overlap.
[0059] Specifically, in this step, a second preset number of test object subsets are used to test the scheme to be verified, that is, a corresponding number of traffic buckets are allocated to the scheme to be verified. The second preset number of test object subsets can be the remaining test object subsets among multiple test object subsets, excluding the first preset number of test object subsets already determined and allocated to the current scheme.
[0060] In the above embodiments, multiple test objects are grouped, and different schemes are implemented for test objects in different groups, so as to collect the behavior logs of test objects under different schemes in the future.
[0061] S203: Determine the first significance level value based on the indicator data of the control group and the indicator data of the reference group; wherein, the first significance level value characterizes the degree of fluctuation of the indicator data of the reference group.
[0062] In this embodiment of the application, a first significance level value is determined based on the indicator data of the control group and the indicator data of the reference group. The first significance level value characterizes the degree of fluctuation of the indicator data of the reference group. Since the indicator data of the control group and the indicator data of the reference group both correspond to the current scheme, the smaller the first significance level value, the less affected the current situation is by fluctuation factors.
[0063] S205: Based on the indicator data of the control group and the indicator data of the experimental group, determine the second significance level value; wherein, the second significance level value characterizes the degree of difference between the scheme to be verified and the current scheme.
[0064] In this embodiment of the application, a second significance level value is determined based on the indicator data of the control group and the indicator data of the experimental group. The second significance level value characterizes the degree of difference between the scheme to be verified and the current scheme, that is, whether there is a significant difference.
[0065] In this embodiment, hypothesis testing is used to determine whether the reference group's indicator data is affected by fluctuation factors, and whether there is a significant difference between the proposed solution and the current solution. Hypothesis testing is a statistical inference method used to determine whether differences between samples, or between a sample and the population, are caused by sampling error or by inherent differences. Here, the first null hypothesis is set as: the reference group's indicator data has fluctuation factors; the first alternative hypothesis is: the reference group's indicator data does not have fluctuation factors; and the first expected significance level (α value) is set to 0.05. Secondly, the second null hypothesis is set as: the proposed solution has a significant difference compared to the current solution; the second alternative hypothesis is: the proposed solution does not have a significant difference compared to the current solution; and the second expected significance level (α value) is set to 0.05.
[0066] Accordingly, in some possible embodiments, determining the first significance level value based on the control group index data and the reference group index data may include, for example: Figure 4 The following steps are shown:
[0067] S401: Perform a homogeneity of variance test on the indicator data of the control group and the reference group to obtain the first homogeneity of variance result.
[0068] Here, the homogeneity of variance test refers to determining whether the variance of a certain indicator is consistent across different sample groups when comparing them. Methods for testing homogeneity of variance include, but are not limited to, the F-test, Bartlett's test, and Levene's test.
[0069] S403: Determine whether the first result of homogeneity of variance is homogeneous. If the first result of homogeneity of variance is homogeneous, proceed to step S405; otherwise, proceed to step S407.
[0070] S405: The first calculation method was used to perform a significance test on the indicator data of the control group and the reference group to obtain the first significance level value.
[0071] Here, the significance test can be performed using a t-test. Correspondingly, when the first result of homogeneity of variance is homogeneity, the first calculation method is the Student's t-test. Specifically, firstly, calculate the mean (denoted as mA, mA') and sample size (denoted as nA, nA') of the control group and reference group indicator data, respectively. Then, calculate the uniform standard deviation (denoted as S1) based on the mean and sample size of the control group and reference group indicator data. Next, calculate the corresponding t-value based on the uniform standard deviation S1, the mean (mA, mA') and sample size (nA, nA') of the control group and reference group indicator data, and then obtain the p-value corresponding to the t-value from the t-distribution table. This p-value is the first significance level value.
[0072] S407: The second calculation method was used to perform a significance test on the indicator data of the control group and the reference group to obtain the first significance level value.
[0073] When using a t-test for significance testing, the corresponding second calculation method is the Welch t-test. Specifically, the standard deviations (represented by SA and SA'), means (represented by mA and mA'), and sample sizes (represented by nA and nA') of the control group and reference group indicator data are calculated respectively. Then, based on the standard deviations (SA and SA'), means (mA and mA'), and sample sizes (nA and nA') of the control group and reference group indicator data, the corresponding t-values are calculated. Finally, the p-value corresponding to the t-value is obtained from the t-distribution table, and this p-value is the first significance level value.
[0074] After calculating the first significance level value according to the above steps S405 or S407, it is compared with the set first expected significance level value of 0.05. When the first significance level value is less than or equal to 0.05, the first null hypothesis is rejected and the first alternative hypothesis is accepted, that is, it is considered that there are no fluctuation factors in the reference group indicator data.
[0075] In some possible embodiments, the determination of the second significance level value based on the control group index data and the experimental group index data described above may include, for example: Figure 5 The following steps are shown:
[0076] S501: Perform a homogeneity of variance test on the indicator data of the control group and the experimental group to obtain the second homogeneity of variance result.
[0077] This step can be referred to as step S401 above, and will not be repeated here.
[0078] S503: Determine whether the result of the second homogeneity of variance is homogeneous. If the result of the second homogeneity of variance is homogeneous, proceed to step S505; otherwise, proceed to step S507.
[0079] S505: The first calculation method was used to perform a significance test on the indicator data of the control group and the indicator data of the experimental group to obtain the second significance level value.
[0080] Here, referring to step S405 above, the first calculation method is the Student's t-test. Specifically, the means (represented by mA and mB) and sample sizes (represented by nA and nB) of the control group and experimental group indicator data are calculated respectively. Then, based on the means and sample sizes of the control group and experimental group indicator data, the uniform standard deviation (represented by S2) is calculated. Next, based on the uniform standard deviation S2, the means (mA and mB) and sample sizes (nA and nB) of the control group and experimental group indicator data, the corresponding t-values are calculated. Finally, the p-value corresponding to the t-value is obtained from the t-distribution table, and this p-value is the second significance level value.
[0081] S507: The second calculation method is used to perform a significance test on the indicator data of the control group and the indicator data of the experimental group to obtain the second significance level value.
[0082] Here, referring to step S407 above, the second calculation method is the Welch t-test. Specifically, the standard deviations (represented by SA and SB), means (represented by mA and mB), and sample sizes (represented by nA and nB) of the control group and experimental group indicator data are calculated respectively. Then, based on the standard deviations (SA and SB), means (mA and mB), and sample sizes (nA and nB) of the control group and experimental group indicator data, the corresponding t-values are calculated. Finally, the p-values corresponding to the t-values are obtained from the t-distribution table, and these p-values are the second significance level values.
[0083] After calculating the second significance level value according to the above steps S505 or S507, it is compared with the set second expected significance level value of 0.05. When the second significance level value is greater than or equal to 0.05, the second null hypothesis is accepted and the second alternative hypothesis is rejected, that is, it is considered that the scheme to be verified is significantly different from the current scheme.
[0084] S207: When the first significance level value is less than or equal to the first preset value and the second significance level value is greater than or equal to the second preset value, determine the iterative availability of the scheme to be tested.
[0085] In this embodiment, when the first significance level is less than or equal to the first preset value, it indicates that the reference group index data does not have fluctuation factors; when the second significance level is greater than or equal to the second preset value, it indicates that the solution to be verified has a significant difference compared to the current solution. Therefore, when the above conditions are met simultaneously, since the influence of fluctuation factors is excluded, the conclusion that the solution to be verified has a significant difference compared to the current solution has a high degree of accuracy, that is, the significant difference is caused by the solution to be verified itself. After determining that the solution to be verified has a significant difference compared to the current solution, it is necessary to determine the iterative availability of the solution to be verified. The iterative availability of the solution to be verified is either iterable or non-iterable, that is, whether the solution to be verified is a more optimized solution compared to the current solution. If it is determined to be iterable, the solution to be verified can be used in a specified Internet product to replace the current solution, thereby realizing the iteration of the solution.
[0086] In some possible embodiments, the control group indicator data includes multiple first indicator data corresponding to preset indicators, and the experimental group indicator data includes multiple second indicator data corresponding to preset indicators. The preset indicators can be defined according to the actual business scenario. In a specific application scenario, such as a video recommendation scenario, the preset indicators include at least one of object conversion rate, exposure click-through rate, completion rate, and playback duration. The object conversion rate can represent the proportion of users who have performed conversion behavior (such as clicking or playing a video) out of all recommended users; the exposure click-through rate can represent the click-through rate of users after the video is exposed (displayed on the recommendation page); and the completion rate refers to the proportion of users who watch the entire video.
[0087] Accordingly, determining the iterative availability of the proposed solution to be tested may include the following steps: determining a first mean and standard deviation based on multiple first indicator data; determining a second mean based on multiple second indicator data; if the difference between the first mean and the second mean is greater than or equal to a preset multiple of the standard deviation, then the iterative availability of the proposed solution to be tested is determined to be iterable; or; if the difference between the first mean and the second mean is less than a preset multiple of the standard deviation, then the iterative availability of the proposed solution to be tested is determined to be non-iterable.
[0088] Specifically, the preset multiplier can be 2, 3, or 4 times. If the difference between the first mean and the second mean is greater than or equal to 2, 3, or 4 times the standard deviation, the experimental group's indicator data can be considered to have significantly improved. This indicates that the scheme to be verified is a more optimized scheme than the current scheme, and therefore, its iterative availability can be determined to be iterable.
[0089] Furthermore, the data processing method in this application embodiment may also include the following steps: determining the improvement level of the scheme to be tested compared with the current scheme under preset indicators based on the first mean and the second mean.
[0090] Specifically, the difference between the second mean and the first mean is divided by the first mean to obtain the improvement level under the preset indicator.
[0091] In the above embodiments, when it is determined that the solution to be tested has a significant difference from the current solution, the improvement effect of the experimental group's indicator data is determined again by the mean of the indicator data of the control group and the mean of the indicator data of the experimental group. If there is a significant improvement effect, the solution to be tested can be used to replace the current solution in a specified Internet product to achieve solution iteration.
[0092] In some possible embodiments, the data involved in the data processing can be presented in a visual form on the front-end interface so that relevant personnel can observe it.
[0093] Specifically, such as Figure 6 As shown, Figure 6 This is a schematic diagram of a front-end specified interface provided in an embodiment of this application. In this specified interface, the data involved in the data processing is presented in tabular form. The first row of the table displays each preset indicator; the second row displays the mean value of each preset indicator in the control group indicator data; the third row displays the improvement level value corresponding to each preset indicator in the reference group indicator data; and the fourth row displays the improvement level value corresponding to each preset indicator in the experimental group indicator data. Different colors can be used to indicate the improvement level values in the table, such as gray indicating no significant difference, red indicating a significant but negative difference, and green indicating a significant and positive difference (shown in bold in the diagram). This allows for a clear visual understanding of the performance of the corresponding scheme under each preset indicator. Furthermore, a search bar can be set above the table. The search bar can include a baseline experiment selection box and a comparison experiment selection box. The baseline experiment selection box can be used to select control group indicator data, and the comparison experiment selection box can be used to select reference group indicator data and experimental group indicator data.
[0094] In summary, this application's embodiments determine a first significance level between the control group's index data and the reference group's index data, and a second significance level between the control group's index data and the experimental group's index data. The first significance level reflects the degree to which the index data is affected by fluctuations, while the second significance level reflects the degree of difference between the iterative scheme and the current scheme. When the first significance level reflects that the current data is not affected by fluctuations, the confidence level of the second significance level is high. This solves the problem of large fluctuations in online test indicators under low-volume conditions, making it difficult to accurately judge the experimental effect. This application can be widely applied to the online evaluation of industrial recommendation systems, reducing situations where indicators improve significantly under low-volume conditions but do not show a corresponding improvement in the overall performance after expanding the volume. It increases the consistency of experimental results between high-volume and low-volume conditions, thereby effectively improving iteration efficiency.
[0095] This application also provides a data processing apparatus, such as... Figure 7 As shown, the data processing device 70 includes:
[0096] The first acquisition module 701 is used to acquire reference group indicator data and experimental group indicator data; the reference group indicator data is obtained from the behavior logs of test objects adopting the current solution; the experimental group indicator data is obtained from the behavior logs of test objects adopting the solution to be verified; the solution to be verified is the iterative solution of the current solution.
[0097] The second acquisition module 702 is used to acquire control group indicator data; the control group indicator data is obtained from the behavior logs of historical objects using the current scheme.
[0098] The first determination module 703 is used to determine a first significance level value based on the indicator data of the control group and the indicator data of the reference group; the first significance level value characterizes the degree of fluctuation of the indicator data of the reference group.
[0099] The second determination module 704 is used to determine the second significance level value based on the indicator data of the control group and the indicator data of the experimental group; the second significance level value characterizes the degree of difference between the scheme to be verified and the current scheme;
[0100] The third determination module 705 is used to determine the iterative availability of the scheme to be tested when the first significance level value is less than or equal to the first preset value and the second significance level value is greater than or equal to the second preset value.
[0101] In some possible embodiments, the apparatus further includes:
[0102] The fourth determination module is used to obtain a test object set; the test object set includes multiple test objects and the identifier of each test object; based on the identifier of each test object, the test object set is divided into multiple test object subsets; each of the multiple test object subsets includes an equal number of test objects; a first preset number of test object subsets are determined from the multiple test object subsets, and the test objects in the first preset number of test object subsets are used as test objects for adopting the current solution; a second preset number of test object subsets are determined from the multiple test object subsets, and the test objects in the second preset number of test object subsets are used as test objects for adopting the solution to be verified; wherein, the first preset number of test object subsets and the second preset number of test object subsets do not overlap.
[0103] In some possible embodiments, the first determining module 703 is further configured to perform a homogeneity of variance test on the control group indicator data and the reference group indicator data to obtain a first homogeneity of variance result; when the first homogeneity of variance result is homogeneous, perform a significance test on the control group indicator data and the reference group indicator data using a first calculation method to obtain a first significance level value; or; when the first homogeneity of variance result is non-homogeneous, perform a significance test on the control group indicator data and the reference group indicator data using a second calculation method to obtain a first significance level value.
[0104] In some possible embodiments, the second determining module 704 is further configured to perform a homogeneity of variance test on the control group index data and the experimental group index data to obtain a second homogeneity of variance result; when the second homogeneity of variance result is homogeneous, perform a significance test on the control group index data and the experimental group index data using a first calculation method to obtain a second significance level value; or; when the second homogeneity of variance result is non-homogeneous, perform a significance test on the control group index data and the experimental group index data using a second calculation method to obtain a second significance level value.
[0105] In some possible embodiments, the control group index data includes multiple first index data corresponding to a preset index, and the experimental group index data includes multiple second index data corresponding to a preset index.
[0106] The third determining module 705 is further configured to determine a first mean and a standard deviation based on multiple first indicator data; determine a second mean based on multiple second indicator data; if the difference between the first mean and the second mean is greater than or equal to a preset multiple of the standard deviation, then the iterative availability of the scheme to be tested is determined to be iterable; or; if the difference between the first mean and the second mean is less than a preset multiple of the standard deviation, then the iterative availability of the scheme to be tested is determined to be non-iterable.
[0107] In some possible embodiments, the third determining module 705 is further configured to determine the improvement level of the scheme to be tested compared with the current scheme under preset indicators based on the first mean and the second mean; wherein the preset indicators include at least one of object conversion rate, exposure click rate, completion rate and playback duration.
[0108] It should be noted that the apparatus and method embodiments described in the device embodiments are based on the same inventive concept.
[0109] This application provides an electronic device including a processor and a memory. The memory stores at least one instruction or at least one program segment, which is loaded and executed by the processor to implement the data processing method provided in the above method embodiments.
[0110] Furthermore, Figure 8A schematic diagram of the hardware structure of an electronic device for implementing the data processing method provided in the embodiments of this application is shown. The electronic device may participate in or include the data processing apparatus provided in the embodiments of this application. Figure 8 As shown, the electronic device 100 may include one or more processors 1002 (shown as 1002a, 1002b, ..., 1002n in the figure) (processor 1002 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 1004 for storing data, and a transmission device 1006 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 8 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, the electronic device 100 may also include... Figure 8 The more or fewer components shown, or having the same Figure 8 The different configurations shown.
[0111] It should be noted that the aforementioned one or more processors 1002 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be wholly or partially integrated into any other element within the electronic device 100 (or mobile device). As involved in the embodiments of this application, the data processing circuit serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0112] The memory 1004 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data processing method described in the embodiments of this application. The processor 1002 executes various functional applications and data processing by running the software programs and modules stored in the memory 1004, thereby implementing the aforementioned data processing method. The memory 1004 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1004 may further include memory remotely located relative to the processor 1002, and these remote memories can be connected to the electronic device 100 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0113] The transmission device 1006 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 100. In one example, the transmission device 1006 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In one embodiment, the transmission device 1006 may be a radio frequency (RF) module for wireless communication with the Internet.
[0114] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows a user to interact with the user interface of the electronic device 100 (or mobile device).
[0115] Embodiments of this application also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a data processing method in the method embodiments. The at least one instruction or the at least one program is loaded and executed by the processor to implement the data processing method provided in the above method embodiments.
[0116] Optionally, in this embodiment, the storage medium may be located in at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0117] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.
[0118] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and electronic device embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0119] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0120] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A data processing method, characterized in that, include: The system acquires reference group indicator data, experimental group indicator data, and control group indicator data. The reference group indicator data is obtained from the behavior logs of test subjects using the current solution. The experimental group indicator data is obtained from the behavior logs of test subjects using the solution to be verified. The solution to be verified is an iterative solution of the current solution. The control group indicator data is obtained from the behavior logs of historical subjects using the current solution. The current solution represents the algorithm or strategy executed in the specified Internet product to implement the function. The historical subjects are users who used the specified Internet product before the test. A homogeneity of variance test is performed on the indicator data of the control group and the indicator data of the reference group to obtain a first homogeneity of variance result. When the first homogeneity of variance result indicates homogeneity, a first calculation method is used to perform a significance test on the indicator data of the control group and the indicator data of the reference group to obtain a first significance level value. Alternatively, when the first homogeneity of variance result indicates non-homogeneity, a second calculation method is used to perform a significance test on the indicator data of the control group and the indicator data of the reference group to obtain a first significance level value. The first significance level value characterizes the degree of fluctuation of the indicator data of the reference group. Based on the indicator data of the control group and the indicator data of the experimental group, a second significance level value is determined; the second significance level value characterizes the degree of difference between the scheme to be verified and the current scheme; When the first significance level value is less than or equal to the first preset value and the second significance level value is greater than or equal to the second preset value, the iterative availability of the scheme to be verified is determined.
2. The data processing method according to claim 1, characterized in that, The method further includes: Obtain a set of test objects; the set of test objects includes multiple test objects and the identifier of each test object among the multiple test objects; Based on the identifier of each test object, the set of test objects is divided into multiple test object subsets; each of the multiple test object subsets includes an equal number of test objects. A first preset number of test object subsets are determined from the plurality of test object subsets, and the test objects in the first preset number of test object subsets are used as the test objects for adopting the current solution; A second preset number of test object subsets are determined from the plurality of test object subsets, and the test objects in the second preset number of test object subsets are used as the test objects for the scheme to be verified; wherein the first preset number of test object subsets and the second preset number of test object subsets do not overlap.
3. The data processing method according to claim 1, characterized in that, The determination of the second significance level value based on the indicator data of the control group and the indicator data of the experimental group includes: The homogeneity of variance was tested on the indicator data of the control group and the indicator data of the experimental group to obtain the second result of homogeneity of variance. When the second result of homogeneity of variance is homogeneous, the first calculation method is used to perform a significance test on the indicator data of the control group and the indicator data of the experimental group to obtain the second significance level value; or, when the second result of homogeneity of variance is non-homogeneous, the second calculation method is used to perform a significance test on the indicator data of the control group and the indicator data of the experimental group to obtain the second significance level value.
4. The data processing method according to claim 1, characterized in that, The control group indicator data includes multiple first indicator data corresponding to the preset indicator, and the experimental group indicator data includes multiple second indicator data corresponding to the preset indicator. Determining the iterative availability of the scheme to be verified includes: The first mean and standard deviation are determined based on the multiple first indicator data; The second mean is determined based on the multiple second indicator data; If the difference between the first mean and the second mean is greater than or equal to a preset multiple of the standard deviation, then the iterative availability of the scheme to be verified is determined to be iterable; or, if the difference between the first mean and the second mean is less than a preset multiple of the standard deviation, then the iterative availability of the scheme to be verified is determined to be non-iterable.
5. The data processing method according to claim 4, characterized in that, The method further includes: Based on the first mean and the second mean, determine the improvement level of the solution to be verified compared to the current solution under the preset index; The preset metrics include at least one of object conversion rate, exposure click-through rate, completion rate, and playback duration.
6. A data processing apparatus, characterized in that, The device includes: The first acquisition module is used to acquire reference group indicator data and experimental group indicator data; the reference group indicator data is obtained from the behavior logs of test objects adopting the current solution; the experimental group indicator data is obtained from the behavior logs of test objects adopting the solution to be verified; the solution to be verified is an iterative solution of the current solution, and the current solution represents the algorithm or strategy executed in the specified Internet product for the implementation of functions; The second acquisition module is used to acquire control group indicator data; the control group indicator data is obtained from the behavior logs of historical objects using the current solution, and the historical objects are users who have used the specified Internet product before the test. The first determining module is used to perform a homogeneity of variance test on the control group indicator data and the reference group indicator data to obtain a first homogeneity of variance result; when the first homogeneity of variance result is homogeneous, a first calculation method is used to perform a significance test on the control group indicator data and the reference group indicator data to obtain a first significance level value; or; when the first homogeneity of variance result is non-homogeneous, a second calculation method is used to perform a significance test on the control group indicator data and the reference group indicator data to obtain a first significance level value; the first significance level value characterizes the degree of fluctuation of the reference group indicator data. The second determining module is used to determine a second significance level value based on the indicator data of the control group and the indicator data of the experimental group; the second significance level value characterizes the degree of difference between the scheme to be verified and the current scheme; The third determining module is used to determine the iterative availability of the scheme to be verified when the first significance level value is less than or equal to the first preset value and the second significance level value is greater than or equal to the second preset value.
7. The data processing apparatus according to claim 6, characterized in that, The device further includes: The fourth determining module is used to obtain a test object set; the test object set includes multiple test objects and an identifier for each test object; based on the identifier of each test object, the test object set is divided into multiple test object subsets; each of the multiple test object subsets includes an equal number of test objects; a first preset number of test object subsets are determined from the multiple test object subsets, and the test objects in the first preset number of test object subsets are used as the test objects for adopting the current solution; a second preset number of test object subsets are determined from the multiple test object subsets, and the test objects in the second preset number of test object subsets are used as the test objects for adopting the solution to be verified; wherein the first preset number of test object subsets and the second preset number of test object subsets do not overlap.
8. The data processing apparatus according to claim 6, characterized in that, The second determining module is further configured to perform a homogeneity of variance test on the control group index data and the experimental group index data to obtain a second homogeneity of variance result; when the second homogeneity of variance result is homogeneous, a first calculation method is used to perform a significance test on the control group index data and the experimental group index data to obtain a second significance level value; or; when the second homogeneity of variance result is non-homogeneous, a second calculation method is used to perform a significance test on the control group index data and the experimental group index data to obtain a second significance level value.
9. The data processing apparatus according to claim 6, characterized in that, The control group indicator data includes multiple first indicator data corresponding to the preset indicator, and the experimental group indicator data includes multiple second indicator data corresponding to the preset indicator. The third determining module is further configured to determine a first mean and a standard deviation based on the plurality of first indicator data; determine a second mean based on the plurality of second indicator data; if the difference between the first mean and the second mean is greater than or equal to a preset multiple of the standard deviation, then determine that the iterative availability of the scheme to be verified is iterable; or; if the difference between the first mean and the second mean is less than a preset multiple of the standard deviation, then determine that the iterative availability of the scheme to be verified is not iterable.
10. The data processing apparatus according to claim 9, characterized in that, The third determining module is further configured to determine, based on the first mean and the second mean, the improvement level of the scheme to be verified compared to the current scheme under the preset index; The preset metrics include at least one of object conversion rate, exposure click-through rate, completion rate, and playback duration.
11. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the data processing method as described in any one of claims 1-5.
12. A computer storage medium, characterized in that, The storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the data processing method as described in any one of claims 1-5.
13. A computer program product, characterized in that, The computer program product includes at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the data processing method as described in any one of claims 1-5.
Citation Information
Patent Citations
Order distribution method and device, computer equipment and computer readable storage medium
CN112465604A