Sample grouping method, device, equipment and computer-readable storage medium
By grouping based on the first index value of the sample and the granularity of the partitioning size, the problem of insufficient universality and performance of the sample grouping method in the prior art is solved, efficient sample grouping is achieved, and the efficiency and effect of AB experiments are improved.
Patent Information
- Application Number
- CN202210482235.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-05-05
AI Technical Summary
The prior art is difficult to provide a sample grouping method with strong versatility and high performance, which affects the efficiency and effectiveness of AB experiments.
By obtaining the first index value of the sample to be grouped, the samples are divided into multiple intermediate groups based on these index values and the division granularity, and then the first grouping result is obtained based on the grouping ratio and the samples of the intermediate group to ensure that the degree of difference between the groups is less than or equal to the degree of difference threshold.
It is realized that multiple samples are divided into groups with smaller differences, which improves the universality and performance of sample grouping and enhances the efficiency and effectiveness of AB experiments.
Smart Images

Figure CN114818946B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a method, device, equipment and computer-readable storage medium for sample grouping. Background Art
[0002] In the technical field of data processing, the AB test (also known as the grouped comparison test) is a commonly used means to analyze the effect differences of different solutions. Generally, in the AB test, a large number of experimental samples are first divided into similar multiple groups of samples through a sample grouping method, and then different solutions are experimented on the multiple groups of samples. The effect differences of each solution are analyzed based on the change differences of the experimental indicators among the multiple groups of samples.
[0003] Therefore, the performance of the sample grouping method directly affects the efficiency and effect of the AB test. Since the sample characteristics required for grouping in different AB tests are different, it is particularly important to provide a sample grouping method with strong generality and high performance. Summary of the Invention
[0004] This application provides a method, device, equipment and computer-readable storage medium for sample grouping to improve the generality of sample grouping and the performance of sample grouping.
[0005] In a first aspect, a method for sample grouping is provided, and the method includes:
[0006] Obtain a plurality of samples to be grouped, each sample includes a corresponding first index value, and the first index value of any sample is used to indicate the difference between the any sample and other samples;
[0007] Based on the first index value corresponding to each sample and the division granularity, divide the plurality of samples into a plurality of intermediate groups. The first index values of the samples included in different intermediate groups are located in different continuous value ranges. The division granularity is used to determine the number of the plurality of intermediate groups, and the division granularity is determined based on the previous iteration grouping result;
[0008] According to the grouping ratio and the samples respectively included in the plurality of intermediate groups, obtain a first grouping result corresponding to the plurality of samples. The difference degree of the first index values between the samples included in each group in the first grouping result is less than or equal to the difference degree threshold.
[0009] In a possible implementation, dividing the multiple samples into multiple intermediate groups based on the first metric value corresponding to each sample and the division granularity includes: sorting the multiple samples based on the first metric value corresponding to each sample to obtain a sorted sample sequence; obtaining a difference sequence corresponding to the sample sequence according to the difference between the first metric values of every two adjacent samples in the sample sequence; and dividing the multiple samples into multiple intermediate groups based on the difference sequence and the division granularity.
[0010] In a possible implementation, dividing the multiple samples into multiple intermediate groups based on the difference sequence and the division granularity includes: obtaining a difference threshold based on the difference sequence and the division granularity; and sequentially traversing each difference in the difference sequence, and dividing the multiple samples into multiple intermediate groups according to the relationship between each difference and the difference threshold.
[0011] In a possible implementation, dividing the multiple samples into multiple intermediate groups according to the relationship between each difference and the difference threshold includes: initializing the first intermediate group;
[0012] For the first difference among each difference, when the number of samples included in the first intermediate group is less than the lower threshold, dividing the sample corresponding to the first difference into the first intermediate group;
[0013] When the number of samples included in the first intermediate group is greater than or equal to the lower threshold and less than the upper threshold, if the first difference is less than the difference threshold, dividing the sample corresponding to the first difference into the first intermediate group, and if the first difference is greater than or equal to the difference threshold, determining the samples included in the first intermediate group and initializing the second intermediate group, and dividing the sample corresponding to the first difference into the second intermediate group;
[0014] When the number of samples included in the first intermediate group is greater than or equal to the upper threshold, determining the samples included in the first intermediate group and initializing the second intermediate group, and dividing the sample corresponding to the first difference into the second intermediate group.
[0015] In a possible implementation, the division granularity is a quantile value; obtaining a difference threshold based on the difference sequence and the division granularity includes: sorting the difference sequence to obtain a sorted difference sequence; and using the difference at the quantile value in the sorted difference sequence as the difference threshold.
[0016] In a possible implementation manner, obtaining the first grouping result corresponding to the multiple samples according to the grouping ratio and the samples respectively included in the multiple intermediate groups includes: obtaining the iterative grouping result corresponding to the multiple samples according to the grouping ratio and the samples respectively included in the multiple intermediate groups; when the difference degree of the first index values between the samples included in each group in the iterative grouping result is less than or equal to the difference degree threshold, using the iterative grouping result as the first grouping result; when the difference degree of the first index values between the samples included in each group in the iterative grouping result is greater than the difference degree threshold, updating the division granularity; based on the first index value corresponding to each sample and the updated division granularity, dividing the multiple samples into multiple updated intermediate groups; and obtaining the first grouping result corresponding to the multiple samples according to the grouping ratio and the samples respectively included in the multiple updated intermediate groups.
[0017] In a possible implementation manner, updating the division granularity includes: if the difference degree of the first index values between the samples included in each group in the iterative grouping result is less than or equal to the difference degree of the first index values between the samples included in each group in the grouping result of the previous iteration, updating the division granularity in the reference direction; if the difference degree of the first index values between the samples included in each group in the iterative grouping result is greater than the difference degree of the first index values between the samples included in each group in the grouping result of the previous iteration, updating the division granularity in the anti-reference direction.
[0018] In a possible implementation manner, obtaining the iterative grouping result corresponding to the multiple samples according to the grouping ratio and the samples respectively included in the multiple intermediate groups includes: randomly grouping the samples respectively included in each of the multiple intermediate groups according to the grouping ratio to obtain the intermediate grouping results respectively corresponding to the multiple intermediate groups; merging the intermediate grouping results respectively corresponding to the multiple intermediate groups according to the grouping ratio to obtain the initial grouping result corresponding to the multiple samples; repeatedly performing the above operations until the difference degree of the first index values between the samples included in each group in the initial grouping result is greater than the difference degree threshold, using the initial grouping result as the iterative grouping result, or until the number of loops reaches the loop threshold, using the initial grouping result of the current loop as the iterative grouping result.
[0019] In a possible implementation manner, each sample further includes a corresponding second index value, and the second index value of any sample is used to indicate the difference between the any sample and other samples;
[0020] After obtaining the first grouping result corresponding to the multiple samples according to the grouping ratio and the samples respectively included in the multiple intermediate groups, the method further includes: for any one of the multiple intermediate groups, based on the division granularity and the second index values respectively corresponding to the samples included in the any one of the intermediate groups, dividing the samples included in the any one of the intermediate groups into multiple sub-intermediate groups, where the second index values of the samples included in different sub-intermediate groups are located in different continuous value ranges; obtaining a sub-grouping result corresponding to the samples included in the any one of the intermediate groups according to the grouping ratio and the samples respectively included in the multiple sub-intermediate groups, where the difference degree of the second index values between the samples included in each group in the sub-grouping result is less than or equal to the difference degree threshold; and obtaining a second grouping result corresponding to the multiple samples based on the sub-grouping results respectively corresponding to the multiple intermediate groups.
[0021] In a second aspect, a sample grouping device is provided, where the device includes:
[0022] A first obtaining module, configured to obtain multiple samples to be grouped, each sample includes a corresponding first index value, and the first index value of any one sample is used to indicate the difference between the any one sample and other samples;
[0023] A dividing module, configured to divide the multiple samples into multiple intermediate groups based on the first index value corresponding to each sample and the division granularity, where the first index values of the samples included in different intermediate groups are located in different continuous value ranges, the division granularity is used to determine the number of the multiple intermediate groups, and the division granularity is determined based on the previous iteration grouping result;
[0024] A second obtaining module, configured to obtain a first grouping result corresponding to the multiple samples according to the grouping ratio and the samples respectively included in the multiple intermediate groups, where the difference degree of the first index values between the samples included in each group in the first grouping result is less than or equal to the difference degree threshold.
[0025] In a possible implementation manner, the dividing module is configured to sort the multiple samples based on the first index value corresponding to each sample to obtain a sorted sample sequence; obtain a difference sequence corresponding to the sample sequence according to the difference between the first index values of every two adjacent samples in the sample sequence; and divide the multiple samples into multiple intermediate groups based on the difference sequence and the division granularity.
[0026] In a possible implementation manner, the dividing module is configured to obtain a difference threshold based on the difference sequence and the division granularity; and sequentially traverse each difference in the difference sequence, and divide the multiple samples into multiple intermediate groups according to the relationship between each difference and the difference threshold.
[0027] In a possible implementation, a partitioning module is configured to initialize a first intermediate group; for a first difference among each of the differences, when the number of samples included in the first intermediate group is less than a lower threshold, the samples corresponding to the first difference are partitioned into the first intermediate group;
[0028] when the number of samples included in the first intermediate group is greater than or equal to the lower threshold and less than an upper threshold, if the first difference is less than the difference threshold, the samples corresponding to the first difference are partitioned into the first intermediate group, and if the first difference is greater than or equal to the difference threshold, the samples included in the first intermediate group are determined and a second intermediate group is initialized, and the samples corresponding to the first difference are partitioned into the second intermediate group;
[0029] when the number of samples included in the first intermediate group is greater than or equal to the upper threshold, the samples included in the first intermediate group are determined and a second intermediate group is initialized, and the samples corresponding to the first difference are partitioned into the second intermediate group.
[0030] In a possible implementation, the partitioning granularity is a quantile value; the partitioning module is configured to sort the difference sequence to obtain a sorted difference sequence; and the difference at the quantile value in the sorted difference sequence is used as the difference threshold.
[0031] In a possible implementation, a second acquisition module is configured to obtain an iterative grouping result corresponding to the multiple samples according to a grouping ratio and the samples included in the multiple intermediate groups respectively; when the difference degree of the first index values between the samples included in each group in the iterative grouping result is less than or equal to the difference degree threshold, the iterative grouping result is used as the first grouping result; when the difference degree of the first index values between the samples included in each group in the iterative grouping result is greater than the difference degree threshold, the partitioning granularity is updated; based on the first index value corresponding to each sample and the updated partitioning granularity, the multiple samples are partitioned into multiple updated intermediate groups; and an iterative grouping result corresponding to the multiple samples is obtained according to the grouping ratio and the samples included in the multiple updated intermediate groups respectively.
[0032] In a possible implementation, the second acquisition module is configured to update the partitioning granularity in a reference direction if the difference degree of the first index values between the samples included in each group in the iterative grouping result is less than or equal to the difference degree of the first index values between the samples included in each group in the grouping result of the previous iteration; and update the partitioning granularity in a reverse reference direction if the difference degree of the first index values between the samples included in each group in the iterative grouping result is greater than the difference degree of the first index values between the samples included in each group in the grouping result of the previous iteration.
[0033] In a possible implementation, a second acquisition module is configured to randomly group the samples included in each of the multiple intermediate groups according to the grouping ratio to obtain intermediate grouping results corresponding to the multiple intermediate groups respectively; merge the intermediate grouping results corresponding to the multiple intermediate groups according to the grouping ratio to obtain an initial grouping result corresponding to the multiple samples; repeatedly execute the above operations until the difference degree of the first index values between the samples included in each group in the initial grouping result is greater than the difference degree threshold, and use the initial grouping result as the iterative grouping result, or, when the number of loops reaches the loop threshold, use the initial grouping result of the current loop as the iterative grouping result.
[0034] In a possible implementation, each sample further includes a corresponding second index value, and the second index value of any sample is used to indicate the difference between the any sample and other samples; the apparatus further includes:
[0035] A third acquisition module is configured to, for any one of the multiple intermediate groups, based on the division granularity and the second index values respectively corresponding to the samples included in the any intermediate group, divide the samples included in the any intermediate group into multiple sub-intermediate groups, and the second index values of the samples included in different sub-intermediate groups are located in different continuous value ranges; according to the grouping ratio and the samples respectively included in the multiple sub-intermediate groups, obtain a sub-grouping result corresponding to the samples included in the any intermediate group, and the difference degree of the second index values between the samples included in each group in the sub-grouping result is less than or equal to the difference degree threshold; based on the sub-grouping results respectively corresponding to the multiple intermediate groups, obtain a second grouping result corresponding to the multiple samples.
[0036] In a third aspect, a computer device is further provided, where the computer device includes a processor and a memory, and at least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to enable the computer device to implement the sample grouping method described in any one of the above.
[0037] In a fourth aspect, a computer-readable storage medium is further provided, where at least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor to enable a computer to implement the sample grouping method described in any one of the above.
[0038] In a fifth aspect, a computer program product or a computer program is further provided, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the sample grouping method described in any one of the above.
[0039] The technical solution provided by this application can at least bring the following beneficial effects:
[0040] In the technical solution provided by this application, since the first index value corresponding to a sample is used to indicate the difference between this sample and other samples, and an iterative method is adopted to determine the division granularity based on the grouping results of each iteration, multiple samples are divided into multiple intermediate groups based on the first index value corresponding to each sample and the division granularity. The first index values of the samples included in different intermediate groups are located in different continuous value ranges, so that the differences between the samples in the same intermediate group are small, and then multiple samples are divided into multiple sample groups with small differences, which has strong versatility for samples with different characteristics and improves the performance of sample grouping. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a schematic diagram of the implementation environment of a sample grouping method provided by an embodiment of this application;
[0043] Figure 2 It is a flowchart of a sample grouping method provided by an embodiment of this application;
[0044] Figure 3 It is a schematic diagram of the process of obtaining a second grouping result provided by an embodiment of this application;
[0045] Figure 4 It is a schematic diagram of the process of a sample grouping method provided by an embodiment of this application;
[0046] Figure 5 It is a schematic diagram of the process of traversing a difference sequence provided by an embodiment of this application;
[0047] Figure 6 It is a schematic diagram of a sample grouping device provided by an embodiment of this application;
[0048] Figure 7 It is a schematic diagram of the structure of a computer device provided by an embodiment of this application;
[0049] Figure 8 It is a schematic diagram of the structure of a server provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0051] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) mentioned in the technical solutions of the embodiments of this application are all collected and processed on the basis of complying with relevant policies and regulations and obtaining the consent of the corresponding subjects. After being processed, this data is used in big data application scenarios and cannot be identified to any natural person or have a specific association with their privacy. For example, the multiple samples to be grouped involved in this application are all obtained under the authorization of the user or with the full authorization of all parties.
[0052] Since the performance of the sample grouping method directly affects the effect of the AB test, and the performance of the sample grouping method is determined by the degree of difference between each group in the grouping result, the smaller the degree of difference between each group in the grouping result, the better the performance of the sample grouping method. Therefore, the embodiments of this application provide a sample grouping method, which can divide multiple samples into multiple sample groups with smaller differences and has strong versatility for samples with different characteristics.
[0053] Figure 1 The figure shows a schematic diagram of the implementation environment of the sample grouping method provided by the embodiments of this application. This implementation environment includes: a computer device 101, which can refer to a terminal or a server. Optionally, an application program is installed and running in the computer device 101, and this application program is an application program that supports grouping multiple samples.
[0054] Exemplarily, the terminal can be any electronic product that can perform human-computer interaction with the user in one or more ways such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device. For example, a personal computer (PC), smartphone, personal digital assistant (PDA), wearable device, pocket PC (PPC), tablet computer, intelligent vehicle console, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.
[0055] Those skilled in the art should understand that the above computer device 101 is only an example, and other existing or future possible computer devices can also be applicable to this application and should also be included within the protection scope of this application, and are hereby incorporated herein by reference.
[0056] It should be noted that the sample grouping method provided in the embodiments of this application can be applied to any scenario that requires grouping multiple samples. For example, AB tests conducted in fields such as the Internet, healthcare, or education. It can be understood that the samples corresponding to the Internet field can be user data, such as user identifiers, the samples corresponding to the healthcare field can be patient data, such as patient identifiers, and the samples corresponding to the education field can be student data, such as student identifiers.
[0057] Exemplarily, taking the AB test conducted in the Internet field as an example, when an Internet product needs to be iteratively optimized, for example, the adjustment of the product interface color or the adjustment of the product logic, etc. An AB test can be conducted on the new and old versions of the product. First, according to the grouping index, the users using the product are divided into an experimental group and a control group. In the same time dimension, the users in the experimental group access the new version, and the users in the control group access the old version. Then, the user experience data or business data corresponding to each version is collected, and the collected user experience data or business data is analyzed to determine the effect differences of the experimental indicators corresponding to each version. Finally, the version with better experimental indicator effects is used as the final adopted iterative version.
[0058] Optionally, the experimental object of the AB test can be a version of a product or one or more elements in a product. For example, an AB test can be conducted on the interface color in a product. At this time, the experimental object is the interface color. Among them, the red interface, the blue interface, and the yellow interface are respectively used as the three versions of the AB test. At this time, according to the grouping index, the users using the product need to be divided into experimental group 1, experimental group 2, and experimental group 3. That is to say, the embodiments of this application do not limit the number of groups for sample grouping and can be flexibly adjusted according to the application scenario.
[0059] In the embodiments of the present application, the grouping index is used to determine whether the difference degree among the samples included in each group in the grouping result of sample grouping is less than or equal to the difference degree threshold, that is, an index for measuring the performance of the grouping result. Exemplarily, when the experimental objects of the A / B test are the versions of a product, the experimental index of the A / B test may include the browsing duration. Using the principle of the controlled variable method, the grouping index is an index that can affect the inaccuracy of the experimental index of the A / B test. At this time, the grouping index for sample grouping may include the historical browsing duration of the user and the user age group. Based on the sample grouping method provided by the embodiments of the present application, a large number of users can be divided into an experimental group and a control group according to this grouping index, and the average historical browsing duration and average age of the users between the experimental group and the control group are similar. Thus, the difference in the effect of the experimental index obtained from the A / B test based on this experimental group and control group is mainly caused by the difference in the product version, and is not affected by the difference in the grouping index between the experimental group and the control group, greatly improving the confidence level of the experimental effect.
[0060] Based on the above Figure 1 shown implementation environment, the embodiments of the present application provide a sample grouping method, which is applied to a computer device 101. The computer device 101 may be a terminal or a server, and the embodiments of the present application do not limit this. As Figure 2 shown, the sample grouping method provided by the embodiments of the present application includes the following steps 201-203.
[0061] Step 201, obtain a plurality of samples to be grouped, each sample includes a corresponding first index value, and the first index value of any sample is used to indicate the difference between any sample and other samples.
[0062] The embodiments of the present application do not limit the type of samples. The plurality of samples may be experimental samples corresponding to the A / B test in any application scenario. For example, in the Internet application scenario, the plurality of samples are identifiers corresponding to multiple users. Before obtaining the plurality of samples to be grouped, it is necessary to first determine one or more grouping indexes corresponding to the plurality of samples to be grouped. The embodiments of the present application do not limit the manner of determining the grouping indexes of the plurality of samples. For example, the grouping indexes may be determined according to manual experience, or determined according to the experimental objects and experimental indexes of the A / B test, etc. Generally, the grouping index is an index that can affect the difference in the effect of the experimental index. Optionally, the grouping index may be the same as the experimental index or different from the experimental index, and the grouping index may be one or more.
[0063] Optionally, when there is one grouping indicator corresponding to the multiple samples, the first indicator value is the value of the sample corresponding to the grouping indicator; when there are multiple grouping indicators corresponding to the multiple samples, the first indicator value can be the value of any one of the grouping indicators corresponding to the sample. Since the grouping indicator can be used to determine whether the difference degree between the samples included in each group in the grouping result of the sample grouping method is less than or equal to the difference degree threshold, each indicator in the grouping indicator is used to determine whether the difference degree corresponding to each indicator between the samples included in each group in the grouping result of the sample grouping method is less than or equal to the difference degree threshold. Optionally, the difference degree threshold can be set according to experience or flexibly adjusted according to the application scenario. For example, the difference degree threshold is 0.002.
[0064] It can be understood that the difference degree can also be represented by the similarity degree. Optionally, each indicator in the grouping indicator is used to determine that the similarity degree corresponding to each indicator between the samples included in each group in the grouping result of the sample grouping method is greater than or equal to the similarity degree threshold. Optionally, the similarity degree threshold can be set according to experience or flexibly adjusted according to the application scenario. For example, the similarity degree threshold is 0.99. When the grouping indicator includes the first indicator, the first indicator value is used to determine that the similarity degree of the first indicator value corresponding to each indicator between the samples included in each group in the grouping result is greater than or equal to the similarity degree threshold.
[0065] In a possible implementation manner, the database of the computer device includes the multiple samples and the first indicator value corresponding to each sample, or the database of the server or terminal connected to the computer device includes the multiple samples and the first indicator value corresponding to each sample. Therefore, the computer device can obtain the multiple samples to be grouped and the first indicator value corresponding to each sample from the database. Optionally, the computer device can also analyze and extract the multiple samples to be grouped and the first indicator value corresponding to each sample from the historical record data on the Internet.
[0066] Step 202: Based on the first indicator value corresponding to each sample and the division granularity, divide the multiple samples into multiple intermediate groups. The first indicator values of the samples included in different intermediate groups are located in different continuous value ranges. The division granularity is used to determine the number of intermediate groups, and the division granularity is determined based on the grouping result of the previous iteration.
[0067] In a possible implementation, after obtaining the first metric value corresponding to each sample and the partitioning granularity, multiple samples can be partitioned into different intermediate groups according to the partitioning granularity. The first metric values of the samples included in each intermediate group correspond to a continuous value range, and the value ranges of the first metric values of the samples included in each of the multiple intermediate groups are different. Thus, samples with similar first metric values can be partitioned into the same intermediate group. Also, since the first metric value of any sample is used to indicate the difference between any sample and other samples, partitioning samples with similar first metric values into the same intermediate group means that the difference between the samples in the same intermediate group is small.
[0068] In the embodiments of the present application, the partitioning granularity represents the coarseness or fineness of partitioning multiple intermediate groups. The coarser the partitioning granularity, the fewer the number of intermediate groups obtained by partitioning, and the greater the difference between each intermediate group. The finer the partitioning granularity, the more the number of intermediate groups obtained by partitioning, and the smaller the difference between each intermediate group.
[0069] The embodiments of the present application do not limit the representation form of the partitioning granularity. Optionally, the partitioning granularity can be directly represented by the number of groups. For example, the partitioning granularity is 3 groups; or the partitioning granularity can be represented by a multiple. For example, the partitioning granularity is 1.2 times; or the partitioning granularity can be represented by a quantile value. For example, the partitioning granularity is the 90th quantile. It can be understood that different partitioning granularities can partition different multiple intermediate groups. The partitioning granularity in the embodiments of the present application is updated and determined through the grouping results in the iterative process. Thus, the partitioning granularities updated and determined for different multiple samples may be different and are the most in line with the sample characteristics of the current multiple samples. Optionally, the partitioning granularity in the first iterative process can be set according to experience or adjusted according to the application scenario.
[0070] In a possible implementation, when the partitioning granularity is represented by the number of groups, based on the first metric value corresponding to each sample and the partitioning granularity, partitioning multiple samples into multiple intermediate groups includes: sorting the multiple samples based on the first metric value corresponding to each sample to obtain a sorted sample sequence; and evenly partitioning the sorted sample sequence into corresponding multiple intermediate groups according to the group value and the arrangement order. Exemplarily, the sorted sample sequence includes 20 sample sequence values, the partitioning granularity is 2 groups, the samples corresponding to the first 10 sample sequence values are partitioned into one intermediate group, and the samples corresponding to the last 10 sample sequence values are partitioned into another intermediate group.
[0071] In a possible implementation, when the partitioning granularity is represented by a multiple or a quantile value, based on the first metric value corresponding to each sample and the partitioning granularity, partitioning multiple samples into multiple intermediate groups includes, but is not limited to, the following steps 2021 - step 2023.
[0072] Step 2021: Sort the multiple samples based on the first metric value corresponding to each sample to obtain a sorted sample sequence.
[0073] Optionally, sorting the multiple samples according to the first metric value can be in descending order or ascending order of the first metric value. Exemplarily, taking the number of multiple samples as 5, the first metric values corresponding to Sample 1 - Sample 5 are 5, 7, 3, 2, and 10 respectively. Then, sorting in descending order of the first metric value, the obtained sorted sample sequence can be expressed as {10; 7; 5; 3; 2}. Thus, from the sample sequence, the value range of the multiple samples for the first metric value and the value sparsity within the value range of the multiple samples for the first metric value can be obtained.
[0074] Step 2022: Obtain a difference sequence corresponding to the sample sequence according to the difference between the first metric values of every two adjacent samples in the sample sequence.
[0075] Since the sample sequence is a sorted sequence, the difference between the first metric values of every two adjacent samples in the sample sequence can represent the change amplitude of the multiple samples for the first metric. Optionally, obtain the difference between the first metric values of every two adjacent samples in the same direction. For example, for N sequence values in the sample sequence, where N is a positive integer greater than 2, successively obtain the difference between the latter sequence value and the adjacent former sequence value, or successively obtain the difference between the former sequence value and the adjacent latter sequence value to obtain N - 1 differences, and these N - 1 differences are the obtained difference sequence, where the N - 1 differences maintain the order of the sample sequence.
[0076] In the embodiments of the present application, there is a corresponding relationship between the sample sequence and the difference sequence, that is, one difference sequence value corresponds to one sample sequence value, and one sample sequence value corresponds to one sample. Therefore, one difference sequence value corresponds to one sample. However, since the difference sequence has one less sequence value than the sample sequence, one of the multiple difference sequence values corresponds to two samples. Optionally, the first difference sequence value in the difference sequence corresponds to two samples, or the last difference sequence value in the difference sequence corresponds to two samples.
[0077] Exemplarily, for the sample sequence {10; 7; 5; 3; 2}, the differences between the previous sequence value and the adjacent subsequent sequence value are obtained in sequence as 3, 2, 2, and 1. Then, the difference sequence corresponding to this sample sequence can be expressed as {3; 2; 2; 1}. Optionally, the first difference sequence value 3 corresponds to sample 5, the second difference sequence value 2 corresponds to sample 2, the third difference sequence value 2 corresponds to sample 1, and the last difference sequence value 1 corresponds to samples 3 and 4; or, the first difference sequence value 3 corresponds to samples 5 and 2, the second difference sequence value 2 corresponds to sample 1, the third difference sequence value 2 corresponds to sample 3, and the last difference sequence value 1 corresponds to sample 4.
[0078] Step 2023, based on the difference sequence and the division granularity, divide multiple samples into multiple intermediate groups.
[0079] In a possible implementation manner, since the difference sequence maintains the arrangement order of the sample sequence and there is a corresponding relationship between the difference sequence values and the samples, then, multiple samples can be divided into multiple intermediate groups according to this difference sequence and the division granularity.
[0080] Optionally, when the division granularity is represented by a magnification factor, dividing multiple samples into multiple intermediate groups based on the difference sequence and the division granularity includes: obtaining at least one adjacent pair of difference sequence values in the difference sequence whose difference magnification factor is greater than or equal to the division granularity; using the at least one adjacent pair of difference sequence values as the division boundary to divide multiple samples into multiple intermediate groups. For example, the difference sequence includes 19 difference sequence values. Taking the division granularity of 1.2 times as an example, if the difference sequence values in the difference sequence whose difference magnification factor is greater than or equal to 1.2 times include difference sequence values 5, 6, 9, and 10, then, the samples corresponding to difference sequence 1 - 5 are divided into the first intermediate group, the samples corresponding to difference sequence 6 - 9 are divided into the second intermediate group, and the samples corresponding to difference sequence 10 - 19 are divided into the third intermediate group.
[0081] Optionally, when the division granularity is represented by a quantile value, dividing multiple samples into multiple intermediate groups based on the difference sequence and the division granularity includes: obtaining a difference threshold based on the difference sequence and the quantile value; sequentially traversing each difference in the difference sequence, and dividing multiple samples into multiple intermediate groups according to the relationship between each difference and the difference threshold. It can be understood that the quantile value is a parameter that can be flexibly adjusted. Different quantile values result in different determined difference thresholds. The difference threshold, as a parameter for dividing intermediate groups, can determine the coarseness or fineness of the division granularity of intermediate groups, that is, it can determine the number of intermediate groups. Each difference in the difference sequence is each difference sequence value in the difference sequence.
[0082] In a possible implementation, based on the difference sequence and the quantile value, determining a difference threshold includes: sorting the difference sequence to obtain a sorted difference sequence; using the difference at the quantile value in the sorted difference sequence as the difference threshold. Exemplarily, taking the quantile value as 90, if the sorted difference sequence includes 100 sequence values, the sequence value ranked 90th among the 100 sequence values is used as the difference threshold; if the sorted difference sequence includes 88 sequence values, the sequence value ranked 79th among the 88 sequence values is used as the difference threshold. That is to say, when the product of the length of the difference sequence and the quantile value is a decimal, the rounding method can be used to obtain the integer corresponding to the decimal, so that the difference at the integer position in the sorted difference sequence can be obtained as the difference threshold.
[0083] In the embodiments of the present application, after determining the difference threshold corresponding to the quantile value in the difference sequence, at least one difference sequence value greater than or equal to the difference threshold can be obtained by traversing each difference in the difference sequence, and using the at least one difference sequence value as a dividing boundary, multiple samples are divided into multiple intermediate groups. For example, if the difference sequence includes 19 difference sequence values, and the difference sequence values greater than or equal to the difference threshold in the difference sequence include difference sequences 8 and 15, then the samples corresponding to difference sequences 1-7 are divided into the first intermediate group, the samples corresponding to difference sequences 8-14 are divided into the second intermediate group, and the samples corresponding to difference sequences 15-19 are divided into the third intermediate group.
[0084] Optionally, for the intermediate groups obtained by division, there may be a situation where the number of samples included is too small. For example, if the number of samples included in an intermediate group is 1, it is difficult to group the samples included in this intermediate group according to the grouping ratio. Based on this, the embodiments of the present application limit the number of samples included in the intermediate group by setting an upper threshold and a lower threshold, that is, the number of samples included in the intermediate group is between the upper threshold and the lower threshold, neither more than the upper threshold nor less than the lower threshold. Optionally, the embodiments of the present application do not limit the selection method of the upper threshold and the lower threshold. For example, it can be set according to experience or flexibly adjusted according to the application scenario. Usually, the selection of the lower threshold is related to the grouping ratio to avoid the above problem that the samples included in the intermediate group cannot be grouped according to the grouping ratio. For example, if the grouping ratio is 1:2, the lower threshold is at least 3.
[0085] In a possible implementation, when the intermediate group includes a corresponding upper threshold and a lower threshold, multiple samples are divided into multiple intermediate groups according to the relationship between each difference and the difference threshold, including: initializing the first intermediate group; for the first difference among each difference, when the number of samples included in the first intermediate group is less than the lower threshold, dividing the samples corresponding to the first difference into the first intermediate group; when the number of samples included in the first intermediate group is greater than or equal to the lower threshold and less than the upper threshold, if the first difference is less than the difference threshold, dividing the samples corresponding to the first difference into the first intermediate group, if the first difference is greater than or equal to the difference threshold, determining the samples included in the first intermediate group and initializing the second intermediate group, and dividing the samples corresponding to the first difference into the second intermediate group; when the number of samples included in the first intermediate group is greater than or equal to the upper threshold, determining the samples included in the first intermediate group and initializing the second intermediate group, and dividing the samples corresponding to the first difference into the second intermediate group.
[0086] It can be understood that after determining the samples included in the first intermediate group and initializing the second intermediate group, for the second difference among each difference, when the number of samples included in the second intermediate group is less than the lower threshold, dividing the samples corresponding to the second difference into the second intermediate group; when the number of samples included in the second intermediate group is greater than or equal to the lower threshold and less than the upper threshold, if the second difference is less than the difference threshold, dividing the samples corresponding to the second difference into the second intermediate group, if the second difference is greater than or equal to the difference threshold, determining the samples included in the second intermediate group and initializing the third intermediate group, and dividing the samples corresponding to the second difference into the third intermediate group; when the number of samples included in the second intermediate group is greater than or equal to the upper threshold, determining the samples included in the second intermediate group and initializing the third intermediate group, and dividing the samples corresponding to the second difference into the third intermediate group. And so on, until each difference in the difference sequence is traversed. Thus, multiple divided intermediate groups and the samples included in each of the multiple intermediate groups can be obtained.
[0087] Step 203, according to the grouping ratio and the samples respectively included in the multiple intermediate groups, obtain the first grouping result corresponding to the multiple samples, and the difference degree of the first index values between the samples included in each group in the first grouping result is less than or equal to the difference degree threshold.
[0088] In an exemplary embodiment, after obtaining the samples included in each of the multiple intermediate groups, the samples included in each intermediate group can be grouped according to the grouping ratio to obtain the intermediate grouping results corresponding to the samples included in each intermediate group. Based on the intermediate grouping results corresponding to the multiple intermediate groups respectively, the first grouping results corresponding to the multiple samples can be obtained. Since, compared with the value range corresponding to the multiple samples, the samples included in each intermediate group are located within a smaller value interval of the first index value, and the value intervals corresponding to each intermediate group are different. Since the value interval of each intermediate group is smaller than the value interval of the multiple samples, therefore, each intermediate group is grouped first, and then each intermediate grouping result is combined. Thus, the performance of the obtained first grouping results is better than the performance of directly grouping according to the multiple samples.
[0089] Optionally, the grouping ratio can be flexibly adjusted according to the requirements of the application scenario. For example, the grouping ratio can be 1:1, then the intermediate group is divided into group 1 and group 2, and group 1 and group 2 respectively include 50% of the samples in the intermediate group; the grouping ratio can also be 2:3:5, then the intermediate group is divided into group 1, group 2 and group 3, group 1 includes 20% of the samples in the intermediate group, group 2 includes 30% of the samples in the intermediate group, and group 3 includes 50% of the samples in the intermediate group.
[0090] In a possible implementation manner, by means of an iterative update of the partitioning granularity, a more accurate first grouping result under the partitioning granularity can be found. Optionally, obtaining the first grouping results corresponding to the multiple samples according to the grouping ratio and the samples included in each of the multiple intermediate groups includes, but is not limited to, the following steps 2031 and 2032.
[0091] Step 2031, obtaining the iterative grouping results corresponding to the multiple samples according to the grouping ratio and the samples included in each of the multiple intermediate groups.
[0092] Optionally, the samples included in each intermediate group are grouped according to the grouping ratio to obtain the intermediate grouping results corresponding to the samples included in each intermediate group. By combining the intermediate grouping results corresponding to the multiple intermediate groups respectively, the iterative grouping results corresponding to the multiple samples are obtained. The embodiments of the present application do not limit the manner of grouping according to the grouping ratio. Optionally, random grouping is performed according to the grouping ratio, or grouping is performed according to the grouping rules according to the grouping ratio. Among them, the grouping rules can be set according to experience. For example, the grouping rule is that two adjacent samples in the sample sequence cannot be assigned to the same group.
[0093] In a possible implementation manner, obtaining an iterative grouping result corresponding to multiple samples according to a grouping ratio and samples respectively included in multiple intermediate groups includes: randomly grouping the samples included in each intermediate group among the multiple intermediate groups according to the grouping ratio to obtain intermediate grouping results respectively corresponding to the multiple intermediate groups; merging the intermediate grouping results respectively corresponding to the multiple intermediate groups according to the grouping ratio to obtain an initial grouping result corresponding to the multiple samples; repeatedly performing the above operations until the difference degree of the first index values between the samples included in each group in the initial grouping result is greater than a difference degree threshold, taking the initial grouping result as the iterative grouping result, or when the number of repetitions reaches a repetition threshold, taking the initial grouping result of the current repetition as the iterative grouping result.
[0094] Optionally, merging the intermediate grouping results respectively corresponding to the multiple intermediate groups according to the grouping ratio includes: merging the combinations with the same grouping ratio in the intermediate grouping results corresponding to each intermediate group according to the grouping ratio occupied by each group in the intermediate grouping result. For example, taking the grouping ratio of 2:3:5, the intermediate grouping result includes group 1, group 2, and group 3, group 1 includes 20% of the samples in the intermediate group, group 2 includes 30% of the samples in the intermediate group, and group 3 includes 50% of the samples in the intermediate group as an example, merging group 1 in the intermediate grouping results corresponding to each intermediate group, merging group 2 in the intermediate grouping results corresponding to each intermediate group, and merging group 3 in the intermediate grouping results corresponding to each intermediate group to obtain the merged group 1, group 2, and group 3, and the merged group 1, group 2, and group 3 are the initial grouping result.
[0095] Step 2032, when the difference degree of the first index values between the samples included in each group in the iterative grouping result is less than or equal to the difference degree threshold, taking the iterative grouping result as the first grouping result; when the difference degree of the first index values between the samples included in each group in the iterative grouping result is greater than the difference degree threshold, updating the partitioning granularity, and obtaining the first grouping result based on the updated partitioning granularity, where the difference degree of the first index values between the samples included in each group in the first grouping result is less than or equal to the difference degree threshold.
[0096] In the embodiments of the present application, obtaining the first grouping result based on the updated partitioning granularity, where the difference degree of the first index values between the samples included in each group in the first grouping result is less than or equal to the difference degree threshold includes: dividing the multiple samples into multiple updated intermediate groups based on the first index value corresponding to each sample and the updated partitioning granularity; obtaining the first grouping result corresponding to the multiple samples according to the grouping ratio and the samples respectively included in the multiple updated intermediate groups.
[0097] It can be understood that the operations performed after updating the partitioning granularity can refer to the operations performed in steps 202 and 203, which will not be elaborated here. Since an updated partitioning granularity is generated in each iteration, each iteration corresponds to an iterative grouping result until the difference degree of the first index values between the samples included in each group in the iterative grouping result is less than or equal to the difference degree threshold, and the iterative grouping result is used as the first grouping result, and the iteration ends.
[0098] In the embodiment of the present application, for the iterative grouping result in the current iteration, updating the partitioning granularity includes: if the difference degree of the first index values between the samples included in each group in the iterative grouping result is less than or equal to the difference degree of the first index values between the samples included in each group in the grouping result of the previous iteration, update the partitioning granularity in the reference direction; if the difference degree of the first index values between the samples included in each group in the iterative grouping result is greater than the difference degree of the first index values between the samples included in each group in the grouping result of the previous iteration, update the partitioning granularity in the reverse reference direction.
[0099] Optionally, the update direction of the partitioning granularity includes increase and decrease. The reference direction can be either increase or decrease, and the reverse reference direction is the direction opposite to the reference direction. By determining different update directions based on the change trend of the grouping results of each iteration, the speed of obtaining the first grouping result, where the difference degree of the first index values between the samples included in each group in the first grouping result is less than or equal to the difference degree threshold, is accelerated. In the embodiment of the present application, after updating the partitioning granularity in each iteration, the value of the updated partitioning granularity will also be recorded, and filtering will be performed according to the recorded value of the updated partitioning granularity when updating the partitioning granularity in each iteration. For example, if the value of the partitioning granularity updated in the current iteration belongs to the recorded value of the updated partitioning granularity, the partitioning granularity is updated again. This avoids obtaining duplicate iterative grouping results based on duplicate partitioning granularities, further accelerating the speed of obtaining the first grouping result, where the difference degree of the first index values between the samples included in each group in the first grouping result is less than or equal to the difference degree threshold.
[0100] In a possible implementation manner, updating the partitioning granularity in the reference direction includes: updating the partitioning granularity in the reference direction and according to the reference step size; updating the partitioning granularity in the reverse reference direction includes: updating the partitioning granularity in the reverse reference direction and according to the reverse step size. Optionally, both the reference step size and the reverse step size can be set according to experience or flexibly adjusted according to the application scenario, and the reference step size and the reverse step size can be the same or different.
[0101] The embodiments of the present application do not limit the manner of calculating the difference degree of the first index values among the samples included in each group. Optionally, the difference degree among the groups can be calculated by calculating the mean or variance of the first index values among the samples included in each group. Exemplarily, taking the mean as an example, the formula (mean max -mean min ) / mean max can be used to calculate the difference degree among the groups. Wherein, mean max represents the mean of the largest first index values among the samples included in each group, and mean min represents the mean of the smallest first index values among the samples included in each group.
[0102] Thus, by iteratively updating the partitioning granularity, the first grouping results under different partitioning granularities can be obtained until a first grouping result with higher performance is found. In addition, by updating the partitioning granularity to adjust the number of intermediate groups obtained, the number of intermediate groups does not need to be set according to manual experience for a specific sample set, but the number of groups adapted to the current sample set is automatically found through an iterative update method, making the sample grouping method provided by the embodiments of the present application highly versatile.
[0103] Through the above steps 201 - step 203, based on the first index values respectively corresponding to multiple samples, the first grouping results corresponding to the multiple samples are obtained, and the difference degree of the first index values among the samples included in each group in the first grouping result is less than or equal to the difference degree threshold. In a possible implementation manner, there may be multiple grouping indicators for the multiple samples. For the case where there are multiple grouping indicators, after obtaining the first grouping result and the difference degree of the first index values among the samples included in each group in the first grouping result is less than or equal to the difference degree threshold, the samples included in each intermediate group among the multiple intermediate groups obtained by partitioning are respectively used as the overall sample set, and steps 201 - step 203 are re - adopted to obtain the corresponding grouping results.
[0104] Exemplarily, taking the second indicator included in the grouping indicator as an example, each sample among the multiple samples obtained also includes a corresponding second indicator value. Optionally, after obtaining the first grouping results corresponding to the multiple samples according to the grouping ratio and the samples respectively included in the multiple intermediate groups, it further includes: for any one of the multiple intermediate groups, based on the partitioning granularity and the second indicator values respectively corresponding to the samples included in any one of the intermediate groups, the samples included in any one of the intermediate groups are divided into multiple sub - intermediate groups, and the second indicator values of the samples included in different sub - intermediate groups are located in different continuous value intervals; according to the grouping ratio and the samples respectively included in the multiple sub - intermediate groups, the sub - grouping results corresponding to the samples included in any one of the intermediate groups are obtained, and the difference degree of the second indicator values among the samples included in each group in the sub - grouping result is less than or equal to the difference degree threshold.
[0105] Optionally, the second index is any index in the grouping indices other than the first index. Similarly, the second index value of any sample is also used to indicate the difference between any sample and other samples.
[0106] In the embodiments of the present application, after obtaining the sub-grouping results corresponding to the samples included in any intermediate group, and the difference degree of the second index values among the samples included in each group in the sub-grouping results is less than or equal to the difference degree threshold, the second grouping results corresponding to multiple samples can be obtained based on the sub-grouping results corresponding to multiple intermediate groups respectively. Optionally, obtaining the second grouping results corresponding to multiple samples based on the sub-grouping results corresponding to multiple intermediate groups respectively includes: combining the sub-grouping results corresponding to the multiple intermediate groups according to the grouping ratio to obtain the second grouping results corresponding to multiple samples. Exemplarily, taking the grouping ratio of 1:1 as an example, for 3 intermediate groups, if the sub-grouping result corresponding to intermediate group 1 is subgroup 11 and subgroup 12, the sub-grouping result corresponding to intermediate group 2 is subgroup 21 and subgroup 22, and the sub-grouping result corresponding to intermediate group 3 is subgroup 31 and subgroup 32, combine subgroup 11, subgroup 21, and subgroup 31, and combine subgroup 12, subgroup 22, and subgroup 32 to obtain the second grouping result.
[0107] In a possible implementation manner, refer to Figure 3 , Figure 3 which is a schematic diagram of a process for obtaining the second grouping result provided by the embodiments of the present application. As Figure 3 shown, for multiple samples to be allocated, first obtain the first grouping result based on the first index value corresponding to each sample. In this process, the multiple samples can be divided into the first intermediate group, the second intermediate group,..., the Mth intermediate group, where M is a positive integer greater than 2.
[0108] Exemplarily, for the samples included in the second intermediate group, further divide the samples included in the first intermediate group into the first sub-intermediate group, the second sub-intermediate group,..., the Kth sub-intermediate group based on the second index value corresponding to each sample included in the second intermediate group, where K is a positive integer greater than 2. Next, divide each sub-intermediate group into an experimental group and a control group according to the grouping ratio. Finally, combine the experimental groups of each sub-intermediate group to obtain the experimental group corresponding to the second intermediate group, and combine the control groups of each sub-intermediate group to obtain the control group corresponding to the second intermediate group.
[0109] Based on the same principle, the experimental group and the control group corresponding to each intermediate group can be obtained. Further, combine the experimental groups corresponding to each intermediate group to obtain the experimental group corresponding to the multiple samples, and combine the control groups corresponding to each intermediate group to obtain the control group corresponding to the multiple samples. The experimental group and the control group corresponding to the multiple samples are the second grouping results.
[0110] In the sample grouping method provided by the embodiments of the present application, since the first index value corresponding to the sample is used to indicate the difference between the sample and other samples, and an iterative manner is adopted to determine the division granularity based on the grouping result of each iteration, multiple samples are divided into multiple intermediate groups based on the first index value corresponding to each sample and the division granularity. The first index values of the samples included in different intermediate groups are located in different continuous value ranges, so that the difference between the samples in the same intermediate group is small. Furthermore, multiple samples are divided into multiple sample groups with small differences, which has strong versatility for samples with different characteristics and improves the performance of sample grouping.
[0111] Exemplarily, taking the division granularity as the quantile value as an example, refer to Figure 4 , Figure 4 which is a schematic process diagram of a sample grouping method provided by the embodiments of the present application. As Figure 4 shown, the sample grouping method includes the following steps 1-step 13.
[0112] Step 1, obtain multiple samples to be assigned.
[0113] Optionally, each sample in the multiple samples includes a corresponding grouping index value. The grouping index value includes the value of at least one index used to determine whether the difference degree between the samples included in each group in the grouping result is less than or equal to the difference degree threshold.
[0114] Step 2, determine whether grouping needs to be performed based on the index; when grouping needs to be performed based on the index, execute Step 3, and when grouping does not need to be performed based on the index, execute Step 13.
[0115] Optionally, traverse each index in the grouping index, and grouping needs to be performed for each index in the grouping index. After traversing each index in the grouping index, it is determined that grouping no longer needs to be performed based on the index.
[0116] Step 3, sort the multiple samples according to the index value to obtain a sorted sample sequence.
[0117] Step 4, based on the difference between the first index values of every two adjacent samples in the sample sequence, obtain the difference sequence corresponding to the sample sequence.
[0118] Step 5, based on the difference sequence and the quantile value, obtain the difference threshold.
[0119] Step 6, traverse the difference sequence based on the difference threshold, and divide the multiple samples into multiple intermediate groups.
[0120] Optionally, refer to Figure 5 , Figure 5It is a schematic diagram of a process for traversing a difference sequence provided by an embodiment of the present application. Step 6 includes the following steps 61 to 66:
[0121] Step 61, initialize the upper limit threshold, lower limit threshold, and intermediate group list.
[0122] Step 62, determine whether the number of samples included in the current intermediate group is less than the lower limit threshold; when the number of samples included in the current intermediate group is less than the lower limit threshold, execute Step 63, and when the number of samples included in the current intermediate group is greater than or equal to the lower limit threshold, execute Step 64.
[0123] Step 63, divide the currently traversed difference into the current intermediate group.
[0124] Step 64, determine whether the number of samples included in the current intermediate group is less than the upper limit threshold; when the number of samples included in the current intermediate group is less than the upper limit threshold, execute Step 65, and when the number of samples included in the current intermediate group is greater than or equal to the lower limit threshold, execute Step 66.
[0125] Step 65, determine whether the difference is less than the difference threshold; when the difference is less than the difference threshold, execute Step 63, and when the difference is greater than or equal to the difference threshold, execute Step 66.
[0126] Step 66, add the current intermediate group to the intermediate group list, initialize the next intermediate group, and divide the currently traversed difference into the next intermediate group.
[0127] By executing the above Steps 61 to 66 for each difference in the difference sequence, after the traversal ends, multiple divided intermediate groups and the samples included in each intermediate group can be obtained through the intermediate group list.
[0128] Step 7, shuffle the samples included in each intermediate group and randomly group them according to the grouping ratio to obtain the intermediate grouping results corresponding to each intermediate group.
[0129] Step 8, merge the intermediate grouping results corresponding to each intermediate group to obtain the initial grouping result corresponding to the multiple samples.
[0130] Step 9, determine whether the difference degree between each group in the initial grouping result is less than or equal to the difference degree threshold; when the difference degree between each group in the initial grouping result is greater than the difference degree threshold, execute Step 10, and when the difference degree between each group in the initial grouping result is less than or equal to the difference degree threshold, execute Step 12.
[0131] Step 10, determine whether the number of loops has reached the loop threshold; when the number of loops has reached the loop threshold, execute Step 11, and when the number of loops has not reached the loop threshold, return to execute Step 7.
[0132] Step 11: Update the quantile value based on the initial grouping result in the current loop, and then return to execute Step 5.
[0133] Step 12: Obtain the initial grouping result in the current loop as the target grouping result corresponding to this metric, and then return to execute Step 2.
[0134] Step 13: Return the optimal grouping result based on the target grouping results corresponding to each metric.
[0135] This sample grouping method sorts the samples according to the metrics, obtains the corresponding difference sequence, and automatically searches for the division results of multiple optimal intermediate groups for the iteratively updated quantile values, making the number of groups of the obtained multiple intermediate groups more in line with the sample characteristics, enhancing the universality of the sample grouping method for different samples, and improving the performance of the grouping results.
[0136] See Figure 6 , this embodiment of the present application provides a sample grouping device, which includes:
[0137] The first acquisition module 601 is configured to acquire multiple samples to be grouped, each sample includes a corresponding first metric value, and the first metric value of any sample is used to indicate the difference between any sample and other samples;
[0138] The division module 602 is configured to divide the multiple samples into multiple intermediate groups based on the first metric value corresponding to each sample and the division granularity. The first metric values of the samples included in different intermediate groups are located in different continuous value ranges. The division granularity is used to determine the number of groups of the multiple intermediate groups, and the division granularity is determined based on the previous iteration grouping result;
[0139] The second acquisition module 603 is configured to obtain the first grouping result corresponding to the multiple samples according to the grouping ratio and the samples included in the multiple intermediate groups. The difference degree of the first metric values between the samples included in each group in the first grouping result is less than or equal to the difference degree threshold.
[0140] In a possible implementation manner, the division module 602 is configured to sort the multiple samples based on the first metric value corresponding to each sample to obtain a sorted sample sequence; obtain a difference sequence corresponding to the sample sequence according to the difference between the first metric values of every two adjacent samples in the sample sequence; and divide the multiple samples into multiple intermediate groups based on the difference sequence and the division granularity.
[0141] In a possible implementation manner, the division module 602 is configured to obtain a difference threshold based on the difference sequence and the division granularity; traverse each difference in the difference sequence in turn, and divide the multiple samples into multiple intermediate groups according to the relationship between each difference and the difference threshold.
[0142] In a possible implementation, a partitioning module 602 is configured to initialize a first intermediate group; for a first difference among each difference, when the number of samples included in the first intermediate group is less than a lower threshold, partition the samples corresponding to the first difference into the first intermediate group;
[0143] When the number of samples included in the first intermediate group is greater than or equal to the lower threshold and less than an upper threshold, if the first difference is less than a difference threshold, partition the samples corresponding to the first difference into the first intermediate group, and if the first difference is greater than or equal to the difference threshold, determine the samples included in the first intermediate group and initialize a second intermediate group, and partition the samples corresponding to the first difference into the second intermediate group;
[0144] When the number of samples included in the first intermediate group is greater than or equal to the upper threshold, determine the samples included in the first intermediate group and initialize a second intermediate group, and partition the samples corresponding to the first difference into the second intermediate group.
[0145] In a possible implementation, the partitioning granularity is a quantile value; the partitioning module 602 is configured to sort a difference sequence to obtain a sorted difference sequence; and use the difference at the quantile value in the sorted difference sequence as the difference threshold.
[0146] In a possible implementation, a second obtaining module 603 is configured to obtain an iterative grouping result corresponding to multiple samples according to a grouping ratio and the samples included in multiple intermediate groups respectively; when the difference degree of the first index values among the samples included in each group in the iterative grouping result is less than or equal to a difference degree threshold, use the iterative grouping result as a first grouping result; when the difference degree of the first index values among the samples included in each group in the iterative grouping result is greater than the difference degree threshold, update the partitioning granularity; based on the first index value corresponding to each sample and the updated partitioning granularity, partition the multiple samples into multiple updated intermediate groups; and obtain a first grouping result corresponding to the multiple samples according to the grouping ratio and the samples included in the multiple updated intermediate groups respectively.
[0147] In a possible implementation, the second obtaining module 603 is configured to update the partitioning granularity in a reference direction if the difference degree of the first index values among the samples included in each group in the iterative grouping result is less than or equal to the difference degree of the first index values among the samples included in each group in the grouping result of the previous iteration; and update the partitioning granularity in a reverse reference direction if the difference degree of the first index values among the samples included in each group in the iterative grouping result is greater than the difference degree of the first index values among the samples included in each group in the grouping result of the previous iteration.
[0148] In a possible implementation manner, the second acquisition module 603 is configured to randomly group the samples included in each intermediate group among a plurality of intermediate groups according to a grouping ratio to obtain intermediate grouping results respectively corresponding to the plurality of intermediate groups; merge the intermediate grouping results respectively corresponding to the plurality of intermediate groups according to the grouping ratio to obtain an initial grouping result corresponding to the plurality of samples; repeatedly execute the above operations until the difference degree of the first index values between the samples included in each group in the initial grouping result is greater than a difference degree threshold, and use the initial grouping result as the iterative grouping result, or, when the number of loops reaches a loop threshold, use the initial grouping result of the current loop as the iterative grouping result.
[0149] In a possible implementation manner, each sample further includes a corresponding second index value, and the second index value of any sample is used to indicate the difference between any sample and other samples; refer to Figure 6 , the apparatus further includes:
[0150] The third acquisition module 604 is configured to, for any one of the plurality of intermediate groups, based on a division granularity and the second index values respectively corresponding to the samples included in any one of the intermediate groups, divide the samples included in any one of the intermediate groups into a plurality of sub-intermediate groups, and the second index values of the samples included in different sub-intermediate groups are located in different continuous value ranges; obtain a sub-grouping result corresponding to the samples included in any one of the intermediate groups according to the grouping ratio and the samples respectively included in the plurality of sub-intermediate groups, and the difference degree of the second index values between the samples included in each group in the sub-grouping result is less than or equal to the difference degree threshold; obtain a second grouping result corresponding to the plurality of samples based on the sub-grouping results respectively corresponding to the plurality of intermediate groups.
[0151] In the sample grouping apparatus provided in the embodiments of the present application, since the first index value corresponding to a sample is used to indicate the difference between the sample and other samples, and an iterative manner is adopted to determine the division granularity based on the grouping result of each iteration, a plurality of samples are divided into a plurality of intermediate groups based on the first index value corresponding to each sample and the division granularity, and the first index values of the samples included in different intermediate groups are located in different continuous value ranges, so that the difference between the samples in the same intermediate group is small, and further, a plurality of samples are divided into a plurality of sample groups with small differences, which has strong versatility for samples with different characteristics and improves the performance of sample grouping.
[0152] It should be understood that when the apparatus provided in the above embodiments implements its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions may be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.
[0153] Please refer to Figure 7 , which shows a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device may be a terminal, for example, it may be: a smart phone, a tablet computer, a vehicle-mounted terminal, a notebook computer or a desktop computer. The terminal may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0154] Generally, the terminal includes: a processor 701 and a memory 702.
[0155] The processor 701 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 701 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 701 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 701 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0156] The memory 702 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 701 to implement the sample grouping method provided by the method embodiment of the present application.
[0157] In some embodiments, the terminal may further optionally include: a peripheral device interface 703 and at least one peripheral device. The processor 701, the memory 702, and the peripheral device interface 703 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 703 via a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 704, a display screen 705, a camera assembly 706, an audio circuit 707, and a power supply 709.
[0158] The peripheral device interface 703 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 701 and the memory 702. In some embodiments, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0159] The radio frequency circuit 704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 704 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 704 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 704 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a Wireless Fidelity (WiFi) network. In some embodiments, the radio frequency circuit 704 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0160] The display screen 705 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 705 is a touch display screen, the display screen 705 also has the ability to collect touch signals on or above the surface of the display screen 705. The touch signals can be input to the processor 701 as control signals for processing. At this time, the display screen 705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 705, which is provided on the front panel of the terminal; in some other embodiments, there can be at least two display screens 705, which are respectively provided on different surfaces of the terminal or are in a folding design; in still some other embodiments, the display screen 705 can be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal. Even, the display screen 705 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 705 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0161] The camera module 706 is used to capture images or videos. Optionally, the camera module 706 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are respectively any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 706 can also include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. A dual-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0162] The audio circuit 707 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 701 for processing, or input to the radio frequency circuit 704 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 701 or the radio frequency circuit 704 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 707 may also include a headphone jack.
[0163] The power supply 709 is used to supply power to each component in the terminal. The power supply 709 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 709 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0164] In some embodiments, the terminal further includes one or more sensors 710. The one or more sensors 710 include but are not limited to: an acceleration sensor 711, a gyroscope sensor 712, a pressure sensor 713, an optical sensor 715, and a proximity sensor 716.
[0165] The acceleration sensor 711 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor 711 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 701 can control the display screen 705 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 711. The acceleration sensor 711 can also be used for collecting game or user's motion data.
[0166] The gyroscope sensor 712 can detect the body direction and rotation angle of the terminal. The gyroscope sensor 712 can cooperate with the acceleration sensor 711 to collect the 3D actions of the user on the terminal. According to the data collected by the gyroscope sensor 712, the processor 701 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0167] The pressure sensor 713 can be disposed on the side frame of the terminal and / or the lower layer of the display screen 705. When the pressure sensor 713 is disposed on the side frame of the terminal, it can detect the holding signal of the user for the terminal, and the processor 701 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 713. When the pressure sensor 713 is disposed on the lower layer of the display screen 705, the processor 701 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 705. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0168] The optical sensor 715 is used to collect the ambient light intensity. In one embodiment, the processor 701 can control the display brightness of the display screen 705 according to the ambient light intensity collected by the optical sensor 715. Specifically, when the ambient light intensity is high, the display brightness of the display screen 705 is increased; when the ambient light intensity is low, the display brightness of the display screen 705 is decreased. In another embodiment, the processor 701 can also dynamically adjust the shooting parameters of the camera assembly 706 according to the ambient light intensity collected by the optical sensor 715.
[0169] The proximity sensor 716, also known as a distance sensor, is usually disposed on the front panel of the terminal. The proximity sensor 716 is used to collect the distance between the user and the front of the terminal. In one embodiment, when the proximity sensor 716 detects that the distance between the user and the front of the terminal is gradually decreasing, the processor 701 controls the display screen 705 to switch from the lit state to the off state; when the proximity sensor 716 detects that the distance between the user and the front of the terminal is gradually increasing, the processor 701 controls the display screen 705 to switch from the off state to the lit state.
[0170] Those skilled in the art can understand that Figure 7 the structure shown in
[0171] does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout. Figure 8 , Figure 8FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 800 may vary greatly due to different configurations or performances, and may include one or more processors 801 and one or more memories 802. Among them, at least one program instruction is stored in the one or more memories 802, and the at least one program instruction is loaded and executed by the one or more processors 801 to implement the sample grouping method provided by each of the above method embodiments. Of course, the server 800 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server 800 may also include other components for implementing the functions of the device, which will not be elaborated here.
[0172] In an exemplary embodiment, a computer device is further provided. The computer device includes a processor and a memory, and at least one program code is stored in the memory. The at least one program code is loaded and executed by one or more processors to enable the computer device to implement any one of the above sample grouping methods.
[0173] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor of a computer device to enable the computer to implement any one of the above sample grouping methods.
[0174] Optionally, the above computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0175] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute any one of the above sample grouping methods.
[0176] In the description, claims and drawings of the present application, terms such as "first", "second", "third" and "fourth" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0177] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included within the protection scope of the present application.
Claims
1. A method for sample grouping, which is applied to the field of Internet A / B testing, and is characterized in that, The method includes: Obtaining a plurality of samples to be grouped, each sample including a corresponding first index value, the first index value of any sample being used to indicate the difference between the any sample and other samples, the sample being a user identifier, and the grouping indicators for sample grouping including the historical browsing duration of the user and the user age group; Based on the first index value corresponding to each sample and the partitioning granularity, partitioning the plurality of samples into a plurality of intermediate groups, the first index values of the samples included in different intermediate groups being located in different continuous value ranges, the partitioning granularity being used to determine the number of the plurality of intermediate groups, and the partitioning granularity being determined based on the previous iteration grouping result; According to the grouping ratio and the samples included in the plurality of intermediate groups respectively, obtaining a first grouping result corresponding to the plurality of samples, the difference degree between the first index values of the samples included in each group in the first grouping result being less than or equal to a difference degree threshold, and the difference degree between each group being obtained by calculating the mean or variance of the first index values of the samples included in each group; The obtaining a first grouping result corresponding to the plurality of samples according to the grouping ratio and the samples included in the plurality of intermediate groups respectively includes: According to the grouping ratio and the samples included in the plurality of intermediate groups respectively, obtaining an iterative grouping result corresponding to the plurality of samples; When the difference degree between the first index values of the samples included in each group in the iterative grouping result is less than or equal to the difference degree threshold, taking the iterative grouping result as the first grouping result; When the difference degree between the first index values of the samples included in each group in the iterative grouping result is greater than the difference degree threshold, updating the partitioning granularity; based on the first index value corresponding to each sample and the updated partitioning granularity, partitioning the plurality of samples into a plurality of updated intermediate groups; according to the grouping ratio and the samples included in the plurality of updated intermediate groups respectively, obtaining a first grouping result corresponding to the plurality of samples.
2. The method according to claim 1, wherein The partitioning the plurality of samples into a plurality of intermediate groups based on the first index value corresponding to each sample and the partitioning granularity includes: Sorting the plurality of samples based on the first index value corresponding to each sample to obtain a sorted sample sequence; According to the difference between the first index values of every two adjacent samples in the sample sequence, obtaining a difference sequence corresponding to the sample sequence; Based on the difference sequence and the partitioning granularity, partitioning the plurality of samples into a plurality of intermediate groups.
3. The method according to claim 2, wherein The partitioning the plurality of samples into a plurality of intermediate groups based on the difference sequence and the partitioning granularity includes: Based on the difference sequence and the partitioning granularity, obtaining a difference threshold; Traversing each difference in the difference sequence in turn, and partitioning the plurality of samples into a plurality of intermediate groups according to the relationship between each difference and the difference threshold.
4. The method according to claim 3, wherein The partitioning the plurality of samples into a plurality of intermediate groups according to the relationship between each difference and the difference threshold includes: Initializing the first intermediate group; For the first difference in each difference, when the number of samples included in the first intermediate group is less than a lower limit threshold, partitioning the sample corresponding to the first difference into the first intermediate group; When the number of samples included in the first intermediate group is greater than or equal to the lower threshold and less than the upper threshold, if the first difference is less than the difference threshold, divide the samples corresponding to the first difference into the first intermediate group; if the first difference is greater than or equal to the difference threshold, determine the samples included in the first intermediate group and initialize the second intermediate group, and divide the samples corresponding to the first difference into the second intermediate group; When the number of samples included in the first intermediate group is greater than or equal to the upper threshold, determine the samples included in the first intermediate group and initialize the second intermediate group, and divide the samples corresponding to the first difference into the second intermediate group.
5. The method according to claim 3, wherein The division granularity is the quantile value; obtaining the difference threshold based on the difference sequence and the division granularity includes: Sort the difference sequence to obtain the sorted difference sequence; Take the difference at the quantile value in the sorted difference sequence as the difference threshold.
6. The method according to claim 1, wherein Updating the division granularity includes: If the difference degree of the first index values between the samples included in each group in the iterative grouping result is less than or equal to the difference degree of the first index values between the samples included in each group in the grouping result of the previous iteration, update the division granularity in the reference direction; If the difference degree of the first index values between the samples included in each group in the iterative grouping result is greater than the difference degree of the first index values between the samples included in each group in the grouping result of the previous iteration, update the division granularity in the anti-reference direction.
7. The method according to claim 1, wherein Obtaining the iterative grouping result corresponding to the multiple samples according to the grouping ratio and the samples included in the multiple intermediate groups includes: Randomly group the samples included in each intermediate group among the multiple intermediate groups according to the grouping ratio to obtain the intermediate grouping results corresponding to the multiple intermediate groups respectively; Merge the intermediate grouping results corresponding to the multiple intermediate groups according to the grouping ratio to obtain the initial grouping result corresponding to the multiple samples; Loop to execute the above operations until the difference degree of the first index values between the samples included in each group in the initial grouping result is greater than the difference degree threshold, take the initial grouping result as the iterative grouping result, or when the number of loops reaches the loop threshold, take the initial grouping result of the current loop as the iterative grouping result.
8. The method according to any one of claims 1-7, characterized in that Each sample also includes a corresponding second index value, and the second index value of any sample is used to indicate the difference between the any sample and other samples; After obtaining the first grouping result corresponding to the multiple samples according to the grouping ratio and the samples included in the multiple intermediate groups, it further includes: For any intermediate group among the multiple intermediate groups, based on the division granularity and the second index values corresponding to the samples included in the any intermediate group, divide the samples included in the any intermediate group into multiple sub-intermediate groups, and the second index values of the samples included in different sub-intermediate groups are located in different continuous value ranges; Obtain the sub-grouping result corresponding to the samples included in any one of the intermediate groups according to the grouping ratio and the samples respectively included in the multiple sub-intermediate groups, where the difference degree of the second index values between the samples included in each group in the sub-grouping result is less than or equal to the difference degree threshold; Based on the sub-grouping results respectively corresponding to the multiple intermediate groups, obtain the second grouping result corresponding to the multiple samples.
9. A sample grouping device, which is applied to the field of Internet A / B testing, is characterized in that The device includes: A first acquisition module, configured to acquire a plurality of samples to be grouped, each sample includes a corresponding first index value, and the first index value of any one sample is used to indicate the difference between the any one sample and other samples. The sample is a user identifier, and the grouping indicators for sample grouping include the user's historical browsing duration and user age group; A division module, configured to divide the plurality of samples into a plurality of intermediate groups based on the first index value corresponding to each sample and the division granularity. The first index values of the samples included in different intermediate groups are located in different continuous value ranges. The division granularity is used to determine the number of the plurality of intermediate groups, and the division granularity is determined based on the previous iteration grouping result; A second acquisition module, configured to obtain the first grouping result corresponding to the plurality of samples according to the grouping ratio and the samples respectively included in the multiple intermediate groups. The difference degree of the first index values between the samples included in each group in the first grouping result is less than or equal to the difference degree threshold, and the difference degree between each group is obtained by calculating the mean or variance of the first index values between the samples included in each group; The obtaining the first grouping result corresponding to the plurality of samples according to the grouping ratio and the samples respectively included in the multiple intermediate groups includes: Obtain the iterative grouping result corresponding to the plurality of samples according to the grouping ratio and the samples respectively included in the multiple intermediate groups; When the difference degree of the first index values between the samples included in each group in the iterative grouping result is less than or equal to the difference degree threshold, use the iterative grouping result as the first grouping result; When the difference degree of the first index values between the samples included in each group in the iterative grouping result is greater than the difference degree threshold, update the division granularity; based on the first index value corresponding to each sample and the updated division granularity, divide the plurality of samples into a plurality of updated intermediate groups; obtain the first grouping result corresponding to the plurality of samples according to the grouping ratio and the samples respectively included in the multiple updated intermediate groups.
10. A computer device, characterized in that, The computer device includes a processor and a memory. At least one computer program or instruction is stored in the memory, and the at least one computer program or instruction is loaded and executed by the processor to enable the computer device to implement the sample grouping method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by the processor to enable the computer to implement the sample grouping method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and device for identifying user type
CN107015993A
Service strategy evaluation method, device and electronic device
CN109308552A