Sample data processing method and device

By pregrouping the sample data and building a decision tree, and assigning the sample data to the experimental group and the control group, the problem of inaccurate test results caused by the difference between the virtual data samples and the real samples is solved, and the authenticity and accuracy of test results are achieved.

CN120180245APending Publication Date: 2025-06-20BEIJING JINGDONG QIANSHITECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311754628.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In scenarios with small sample sizes, the difference between virtual data samples and real samples affects the authenticity and accuracy of the test results.

Method used

By pregrouping the sample data, assigning category tags, and building multiple decision trees based on multiple sample features, the sample data is assigned to the experimental group and the control group.

Benefits of technology

It improves the authenticity and accuracy of product test results, and even when the sample number is small, it can ensure that the sample data in the experimental group and the control group are similar.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180245A_ABST
    Figure CN120180245A_ABST
Patent Text Reader

Abstract

The invention discloses a sample data processing method and device, and relates to the technical field of testing. The embodiment of the method comprises the following steps: pre-grouping sample data, and allocating a category label to each sample in the sample data according to a pre-grouping result; constructing a plurality of decision trees based on a plurality of sample features included in the sample data and category labels of samples to which the sample features belong; and distributing sample data to an experimental group and a control group for testing by utilizing the constructed decision trees. The authenticity and the accuracy of a test result for a product are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of testing, and in particular, to a method and apparatus for processing sample data. Background Art

[0002] Before a developed product such as a logistics distribution management device, a route planning device, or a distribution area division device is put into use, it generally needs to be tested with data first. Currently, in the testing process, the real data of the business is mainly divided into similar experimental group samples and control group samples, and the existing product is used to process the control group samples, the developed product is used to process the experimental group samples, and the processing results of the existing product and the developed product are compared to evaluate the developed product. Among them, the similarity between the experimental group samples and the control group samples is an important factor to ensure the accuracy of the evaluation results.

[0003] Currently, in scenarios where the sample size is relatively small, the sample data is mainly expanded by generating virtual data samples to ensure the similarity of the two groups of samples. However, there are inevitably differences between the virtual data samples and the real samples, and the existence of such differences will affect the authenticity and accuracy of the test results. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method and apparatus for processing sample data, which can effectively improve the authenticity and accuracy of the test results for products.

[0005] To achieve the above object, in a first aspect, an embodiment of the present invention provides a method for processing sample data, including:

[0006] Pre-grouping the sample data, and according to the result of the pre-grouping, assigning a category label to each sample in the sample data;

[0007] Based on multiple sample features included in the sample data and the category label of the sample to which the sample feature belongs, constructing multiple decision trees;

[0008] Using the multiple constructed decision trees, allocating the sample data to an experimental group and a control group for testing.

[0009] Optionally, the pre-grouping the sample data includes:

[0010] Extracting multiple sample features from the sample data;

[0011] Using the multiple sample features to perform clustering processing on the multiple samples included in the sample data;

[0012] According to the result of the clustering processing, pre-dividing the sample data into two sample groups for testing.

[0013] Optionally, extracting a plurality of sample features from the sample data includes:

[0014] Extracting a plurality of initial sample features from the sample data according to a predefined feature extraction strategy;

[0015] Constructing linear regression relationships between the plurality of initial sample features and a preset value index and a preset experimental item respectively;

[0016] According to the constructed linear regression relationships, screening out a plurality of sample features that are irrelevant to the experimental item and relevant to the value index from the plurality of initial sample features.

[0017] Optionally, clustering the plurality of samples included in the sample data includes:

[0018] Clustering the plurality of samples included in the sample data into two sample classes.

[0019] Optionally, pre-dividing the sample data into two sample groups for testing includes:

[0020] Directly exchanging some samples in the two sample classes to obtain two sample groups for testing;

[0021] Or,

[0022] For each of the sample classes, perform the following operations:

[0023] According to the sample features, calculating the distance from each sample in the sample class to the center point of the sample class;

[0024] Sorting the plurality of samples in the sample class according to the calculated distances;

[0025] Swapping the samples corresponding to the even sorting positions or the odd sorting positions in the two sample classes to obtain two sample groups for testing.

[0026] Optionally, the above sample data processing method further includes:

[0027] When the number of samples included in the sample data is even, controlling the number of samples included in the two sample groups to be equal;

[0028] When the number of samples included in the sample data is odd, controlling the number of samples in the first sample group to be 1 more than the number of samples in the second sample group, where the first sample group corresponds to the experimental group and the second sample group corresponds to the control group.

[0029] Optionally, the above sample data processing method further includes:

[0030] Steps for iteratively optimizing the construction of multiple decision trees and steps for allocating the sample data to an experimental group and a control group for testing until an iterative optimization stop condition is met, and determining the allocation result corresponding to the iterative optimization stop condition as the processing result of the sample data.

[0031] Optionally, for each iteration cycle other than the first iteration cycle,

[0032] The steps for constructing multiple decision trees include:

[0033] According to the allocation result of the previous iteration cycle, update the class labels of the samples in the previous iteration cycle or re-allocate class labels to the samples.

[0034] Based on the multiple sample features included in the sample data and the updated class labels of the samples to which the sample features belong or the re-allocated class labels, construct multiple decision trees.

[0035] Optionally, the construction of multiple decision trees includes:

[0036] Perform sampling with replacement on the sample data to construct multiple sample sets.

[0037] For each of the sample sets, perform the following operations:

[0038] Randomly select a set number of target sample features in random order from the samples included in the sample set.

[0039] Use the set number of randomly selected target sample features and the class labels of the samples to which the target sample features belong to construct a decision tree.

[0040] Optionally, the allocation of the sample data to an experimental group and a control group for testing includes:

[0041] For each sample in the sample data, perform the following operations:

[0042] Use multiple of the decision trees to calculate the score of the sample.

[0043] Sort the multiple samples in the sample data according to the scores of the samples.

[0044] According to the sorting result, allocate the multiple samples in the sample data to the experimental group and the control group.

[0045] Optionally, the calculation of the score of the sample includes:

[0046] Determine the output result of each of the decision trees for the sample, where the output result of each of the decision trees for the sample is the first category label corresponding to the experimental group or the second category label corresponding to the control group;

[0047] According to the output results of multiple decision trees for the sample, count the number of labels corresponding to the first category label and / or the number of labels corresponding to the second category label;

[0048] Using the statistical results, calculate the score of the sample.

[0049] Optionally, the step of using the statistical results to calculate the score of the sample includes:

[0050] In the case where the number of labels corresponding to the first category label is counted, directly determine that the ratio between the number of labels corresponding to the first category label for the sample and the number of decision trees is the score of the sample;

[0051] Or,

[0052] In the case where the number of labels corresponding to the second category label is counted, directly determine that the ratio between the number of labels corresponding to the second category label for the sample and the number of decision trees is the score of the sample;

[0053] Or,

[0054] In the case where the number of labels corresponding to the first category label and the number of labels corresponding to the second category label are both counted,

[0055] Compare the number of labels corresponding to the first category label with the number of labels corresponding to the second category label;

[0056] In the case where the comparison result indicates that the number of labels corresponding to the first category label is greater than or equal to the number of labels corresponding to the second category label, determine that the ratio between the number of labels corresponding to the first category label for the sample and the number of decision trees is the score of the sample;

[0057] In the case where the comparison result indicates that the number of labels corresponding to the second category label is greater than the number of labels corresponding to the first category label, determine that the ratio between the number of labels corresponding to the second category label for the sample and the number of decision trees is the score of the sample.

[0058] Optionally, the above sample data processing method further includes:

[0059] In the current iteration cycle, using the output results of multiple decision trees for each sample, count the number of labels of the first type of label corresponding to the experimental group and the number of labels of the second type of label corresponding to the control group for each sample;

[0060] According to the counted number of labels of the first type of label and the number of labels of the second type of label for each sample, calculate the gap between the experimental group and the control group assigned in the current iteration cycle;

[0061] In the case where the gap is less than a preset gap threshold, determine that the current iteration cycle meets the iteration optimization stop condition.

[0062] Optionally, the above sample data processing method further includes:

[0063] Calculate the gap between the experimental group and the control group assigned in each iteration cycle;

[0064] In the case where the number of iterations is not less than a preset iteration threshold,

[0065] Calculate the first gap average value between the gaps corresponding to the first set number of consecutive iteration cycles;

[0066] Calculate the second gap average value between the gaps corresponding to the second set number of consecutive iteration cycles;

[0067] According to the first gap average value and the second gap average value, calculate the gap change rate;

[0068] In the case where the gap change rate is less than a preset change rate threshold, determine that the current iteration cycle meets the iteration optimization stop condition.

[0069] In a second aspect, an embodiment of the present invention provides a sample data processing device, which is characterized by including: a pre-grouping module, a model processing module, and a data processing module, where,

[0070] The pre-grouping module is used to pre-group sample data and assign category labels to each sample in the sample data according to the pre-grouping result;

[0071] The model processing module is used to construct multiple decision trees based on multiple sample features included in the sample data and the category labels of the samples to which the sample features belong;

[0072] The data processing module is used to use the multiple constructed decision trees to allocate the sample data to an experimental group and a control group for testing.

[0073] One embodiment of the above invention has the following advantages or beneficial effects: By pre-grouping the sample data to assign a class label to each sample according to the result of the pre-grouping, where the class label indicates whether the sample belongs to the experimental group or the control group. Since the sample features can better reflect the characteristics and essence of the samples, the experimental group, and the control group, when constructing a decision tree based on multiple sample features and class labels included in the sample data subsequently, the decision tree can provide a more accurate grouping. Then, using the constructed multiple decision trees to allocate the sample data to the experimental group and the control group for testing, even with a relatively small number of samples, it can ensure that the sample data in the experimental group is similar to the sample data in the control group. Subsequently, using similar experimental and control groups for testing can effectively improve the authenticity and accuracy of the test results for the product.

[0074] The further effects of the above non-conventional optional manner will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:

[0076] Figure 1 is a schematic diagram of the main process of a method for processing sample data according to an embodiment of the present invention;

[0077] Figure 2 is a schematic diagram of the main process of extracting multiple sample features from sample data according to an embodiment of the present invention;

[0078] Figure 3 is a schematic diagram of the main process of allocating sample data to the experimental group and the control group for testing according to an embodiment of the present invention;

[0079] Figure 4 is a schematic diagram of the main process of calculating the score of a sample according to an embodiment of the present invention;

[0080] Figure 5 is a schematic diagram of the main process framework of sample data processing according to an embodiment of the present invention;

[0081] Figure 6 is a schematic diagram of the main process of another method for processing sample data according to an embodiment of the present invention;

[0082] Figure 7 is a schematic diagram of the curve relationship between the number of iterations and the gap between the experimental group and the control group according to an embodiment of the present invention;

[0083] Figure 8 is a schematic diagram of the main modules of a sample data processing device according to an embodiment of the present invention;

[0084] Figure 9 is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;

[0085] Figure 10 is a schematic structural diagram of a computer system of a terminal device suitable for implementing the embodiments of the present invention. Detailed implementation manners

[0086] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.

[0087] Figure 1 is a main flowchart of a sample data processing method according to an embodiment of the present invention. As Figure 1 shown, the sample data processing method may include the following steps:

[0088] Step S101: Pre-group the sample data, and assign a class label to each sample in the sample data according to the result of the pre-grouping;

[0089] Among them, sample data generally refers to the data that needs to be processed by the product selected for testing the performance or effect of the product. The sample data may contain a certain amount of samples. Among them, a sample is a set of data that can be analyzed by the product. For example, for the product to be tested as a logistics distribution management device, the sample data generally comes from the data generated by logistics distribution, and the data included in a logistics distribution order is a sample. Another example is that for the product to be tested as a route planning device, the sample data generally comes from data such as the delivery address of a logistics order, and a logistics order or a delivery address is a sample.

[0090] Among them, the class label refers to a mark used to indicate that the sample belongs to the experimental group or the control group, and the class label can be customized. For example, the class label indicating the experimental group is 1, and the class label indicating the control group is 0. Another example is that the class label indicating the experimental group is A, and the class label indicating the control group is B.

[0091] Step S102: Construct multiple decision trees based on multiple sample features included in the sample data and the class labels of the samples to which the sample features belong;

[0092] The sample features are obtained from the data included in the samples. They are mainly related to the samples and can be determined by the business side. The sample features may include: basic features that can directly describe the characteristics of the samples and are obtained directly from the samples, features that are intuitive and highly interpretable, features that cover different categories of samples, and features that can incorporate business logic and facilitate iterative optimization. These features can be directly obtained from the samples or can be calculated through direct statistics (the features obtained by direct statistics are data features that can be directly obtained based on the original data of the samples through calculation and statistics, such as the number of warehouse SKUs and the freight volume of a route), data transformation statistics (the features obtained by data transformation statistics are features calculated based on the original data of the samples through function transformation processing, and common calculation functions include taking the logarithm of the freight volume and calculating the mean of the warehouse SKU shipment volume), and other calculation methods (such as calculating variance and unbiased estimation). For example, for samples related to warehousing planning, the sample features may include the number of SKU inventories, the types of SKUs, and whether there are Tianlang and Dilang robots. For example, for samples related to route planning, the sample features may include the freight volume of consumer goods, personal items, business documents, and office supplies. For another example, for samples related to logistics distribution, the sample features may include unit freight, single-piece weight, and transportation timeliness. It should be noted that each sample generally contains at least one sample feature.

[0093] Step S103: Use the constructed multiple decision trees to allocate the sample data to the experimental group and the control group for testing.

[0094] In Figure 1 In the illustrated embodiment, by pre-grouping the sample data to assign a category label to each sample according to the result of the pre-grouping, the category label indicates whether the sample belongs to the experimental group or the control group. Since the sample features can better reflect the characteristics and essence of the samples, the experimental group, and the control group, constructing a decision tree based on the multiple sample features and category labels included in the sample data can enable the decision tree to provide more accurate grouping. Then, use the constructed multiple decision trees to allocate the sample data to the experimental group and the control group for testing. Even with a relatively small number of samples, it can ensure that the sample data in the experimental group is similar to the sample data in the control group. Then, subsequent testing using similar experimental groups and control groups can effectively improve the authenticity and accuracy of the test results for the product.

[0095] Among them, there are various specific implementation manners for the above-mentioned step S101. For example, the first specific implementation manner can be to pre-group the sample data manually according to experience. The second specific implementation manner can be to randomly pre-group the sample data. The third specific implementation manner includes: extracting multiple sample features from the sample data; using the multiple sample features to perform clustering processing on multiple samples included in the sample data; and according to the result of the clustering processing, pre-dividing the sample data into two sample groups for testing. Among them, one of the two pre-divided sample groups is the pre-divided experimental group, and the other is the pre-divided test group. Correspondingly, the specific implementation of assigning category labels to each sample in the sample data is: marking the experimental group label for each sample pre-divided into one sample group, and marking the control group label for each sample pre-divided into the other sample group. The sample features can better reflect the characteristics of the samples. By clustering the samples in the sample data through the sample features, the distances between similar samples and the clustering centers can be basically the same, so that the samples in the two sample groups pre-divided by the pre-grouping are relatively similar.

[0096] In a preferred embodiment, step S101 is implemented by selecting the above-mentioned third specific implementation manner to perform pre-grouping according to the sample features and reduce the manual intervention in the grouping.

[0097] Among them, for the step of clustering multiple samples included in the sample data in the third specific implementation manner of the above step S101, the first specific implementation manner may include: clustering with one clustering center to obtain the distances between each sample and the clustering center. Thus, relatively similar samples can be directly determined according to the distances between the samples and the clustering center point, and the relatively similar samples can be divided into two sample groups. That is, for the first specific implementation manner of the clustering process, the subsequent specific implementation manner of dividing the sample data into two sample groups for testing may be: sorting each sample in the sample data according to the distance between the sample and the clustering center point, and dividing the samples located at odd positions in the sorting into one sample group, and dividing the samples sorted at even positions into another sample group. When the number of samples included in the two sample groups is the same, randomly select one sample group to assign the experimental group label, and assign the control group label to the other sample group. When the number of samples included in the two sample groups is different (the number of samples in one sample group is one more than that in the other sample group), assign the experimental group label to the sample group with more samples, and assign the control group label to the sample group with fewer samples. For example, for the sample order of the samples included in the sample data after the above clustering and sorting: A1, A2, A3, A4, A5, A6, A7, A8, …, the samples A1, A3, A5, A7, … located at odd positions are assigned to one sample group, and the samples A2, A4, A6, A8, … located at even positions are assigned to another sample group. To ensure that the number of samples included in the control group and the experimental group is relatively balanced, and at the same time ensure the similarity of the samples included in the control group and the experimental group.

[0098] In addition, the second specific implementation manner of clustering multiple samples included in the sample data may include: directly clustering the multiple samples included in the sample data into two sample classes. For the manner of clustering into two sample classes, during the clustering process, it is necessary to constrain the difference between the number of samples included in the two sample classes not to exceed 1 to ensure the balance of clustering. Since the samples in each sample class obtained by clustering are samples with relatively large similarity, through the second specific implementation manner, the samples clustered into the two sample classes are quite different.

[0099] Furthermore, on the basis of the second specific implementation manner of clustering multiple samples included in the sample data above, there are two implementation manners for pre-dividing the sample data into two sample groups for testing, among which,

[0100] Method 1: Directly exchange some samples in two sample classes to obtain two sample groups for testing. For example, for sample data including samples B1, B2, B3, B4, B5, B6, B7, B8, B9, …, through the clustering process of the above-mentioned second specific implementation manner, two sample classes obtained are respectively: sample class 1: B1, B2, B3, B4, …; sample class 2: B5, B6, B7, B8, B9, …. Then through this method 1, B1, B2, … in sample class 1 can be directly exchanged with B5, B6, … in sample class 2, thereby obtaining two sample groups. In a preferred embodiment, the number of samples involved in the exchange is half of the number of samples in sample class 1 and half of the number of samples in sample class 2, so as to better reflect the balance of grouping in the pre-grouping as much as possible.

[0101] Method 2: For each sample class, perform the following operations: Calculate the distance from each sample in the sample class to the center point of the sample class according to the sample characteristics; Sort the multiple samples in the sample class according to the calculated distance. Further, based on the sorting result, exchange the samples corresponding to the even sorting positions or the odd sorting positions in the two sample classes to obtain two sample groups for testing.

[0102] Among them, the distance from each sample in the sample class to the center point of the sample class can be calculated by the following calculation formula (1).

[0103]

[0104] Among them, d(x,x k ) represents the distance from the sample x belonging to the k-th sample class to the central sample point x k of the k-th sample class; x t represents the t-th sample feature of the sample x; represents the t-th sample feature of the central sample point x k ; n represents the number of sample features.

[0105] In addition, the above-mentioned sorting of multiple samples in the sample class according to the calculated distance can also be replaced by sorting all samples according to the calculated distance. This sorting is to first sort the samples in one sample class, and then continue to sort the samples in the other sample class subsequently. For example, the sorting result Among them, belongs to the same sample class, then belongs to another sample class. Based on this sorting result, the specific implementation manner for subsequently obtaining two sample groups for testing is: Select the samples in the odd positions of the sorting result to form a sample group, and select the samples in the even positions in the sorting result to form another sample group.

[0106] Further, as Figure 2 shown, the specific implementation of extracting multiple sample features from the sample data may include:

[0107] Step S201: Extract multiple initial sample features from the sample data according to a predefined feature extraction strategy;

[0108] Among them, the predefined feature extraction strategy can be set accordingly according to specific business requirements. The initial sample features are directly obtained from the samples, or can be obtained through direct statistics (the features obtained by direct statistics are data features that can be directly obtained based on the original data of the samples, such as the number of warehouse SKUs, the freight volume of the line, etc.), data transformation statistics (the features obtained by data transformation statistics are features calculated based on the original data of the samples through function transformation processing, and common calculation functions such as taking the logarithm of the freight volume and calculating the mean of the warehouse SKU shipment volume), and other calculation methods (such as calculating variance, unbiased estimation, etc.).

[0109] Step S202: Construct linear regression relationships between multiple initial sample features and a preset value index and a preset experimental item respectively;

[0110] The value index refers to an index related to evaluating or assessing the business, such as the distribution efficiency, distribution distance, distribution cost, etc. for distribution, and the storage volume, storage cost, etc. for warehousing.

[0111] Among them, the experimental item is predefined. For example, for the logistics business department using the pickup and delivery algorithm, the corresponding experimental item label is 1, and for the logistics business department not using the pickup and delivery algorithm, the corresponding experimental item label is 0. Then, select multiple initial sample features in the sample data of the logistics business department with the experimental item label of 1 to construct a linear regression relationship with the pickup and delivery algorithm.

[0112] Among them, the linear regression relationship reflects the correlation between the initial sample features and the value index and / or the experimental item. If there is a linear regression relationship, it means that the initial sample features are relevant to the value index and / or the experimental item. If there is no linear regression relationship, it means that the initial sample features are not relevant to the value index and / or the experimental item.

[0113] Step S203: According to the constructed linear regression relationship, screen out multiple sample features that are irrelevant to the experimental item and relevant to the value index from multiple initial sample features. By screening out multiple sample features that are irrelevant to the experimental item and relevant to the value index through the above process, the balance and similarity of subsequent grouping of sample data can be improved.

[0114] Further, the above sample data processing method may further include: when the number of samples included in the sample data is even, controlling the number of samples included in the two sample groups to be equal; when the number of samples included in the sample data is odd, controlling the number of samples in the first sample group to be 1 more than the number of samples in the second sample group, where the first sample group corresponds to the experimental group and the second sample group corresponds to the control group. Among them, the process of controlling the number of samples in the sample group is generally controlled by constructing a constraint condition for clustering during the above clustering process.

[0115] Further, based on the above various embodiments, the above sample data processing method may further include: iteratively optimizing the steps of constructing multiple decision trees and the steps of iteratively optimizing the allocation of sample data to the experimental group and the control group for testing until the iterative optimization stop condition is met, and determining the allocation result corresponding to the iterative optimization stop condition as the processing result of the sample data. The iterative optimization stop condition may be that the number of iterations reaches a preset iteration number threshold, or that the difference between the control group and the experimental group after iteration is within a preset difference range. Through iterative optimization, the similarity between the experimental group and the control group can be relatively high, so as to effectively improve the accuracy of the evaluation of the subsequent product to be tested by the experimental group and the control group.

[0116] Specifically, the iterative optimization process may include: for each iteration cycle except the first iteration cycle,

[0117] The steps of constructing multiple decision trees include: according to the allocation result of the previous iteration cycle, updating the class labels of the samples in the previous iteration cycle or reassigning class labels to the samples; based on multiple sample features included in the sample data and the updated class labels or reallocated class labels of the samples to which the sample features belong, constructing multiple decision trees. By updating the class labels of the samples in the previous iteration cycle or reassigning class labels to the samples, the class labels on which the current iteration cycle is based can be made better, and the accuracy of the constructed decision trees can be improved through the iterative optimization process.

[0118] Among them, the iterative construction of multiple decision trees can be implemented by a random forest model, that is, the solution provided by the embodiments of the present invention utilizes the learning ability of the random forest model to gradually optimize the process of allocating the experimental group and the control group. For example, after the first model training (i.e., the first iteration cycle) and the allocation of the experimental group and the control group, the sample uniformity will be better than the pre-grouping result after the above clustering. After the second iteration, the model will be trained according to the result of the first iteration, which is equivalent to obtaining a better training set, thereby training a better model and gradually increasing the reliability of the model, so as to effectively improve the reliability of the allocated experimental group and control group.

[0119] Specifically, there are two implementation manners for stopping iterative optimization.

[0120] The first implementation manner of stopping iterative optimization may include: in the current iteration cycle, using the output results of multiple decision trees for each sample, counting the number of labels of the first type corresponding to the experimental group and the number of labels of the second type corresponding to the control group for each sample; calculating the gap between the experimental group and the control group allocated in the current iteration cycle according to the counted number of labels of the first type and the number of labels of the second type of each sample; and determining that the current iteration cycle meets the iterative optimization stop condition when the gap is less than a preset gap threshold.

[0121] Among them, using the output results of multiple decision trees for each sample means that each decision tree will give each sample a corresponding type label (the first type label corresponding to the experimental group or the second type label corresponding to the control group), that is, the number of decision trees is the same as the number of type labels obtained for each sample. Based on this, the specific implementation manner of calculating the gap between the experimental group and the control group allocated in the current iteration cycle can be calculated by the following calculation formulas (2) and (3).

[0122]

[0123] Among them, represents the sample gap between the sample x being allocated to the experimental group and the control group in the s-th iteration cycle; A x represents the counted number of labels of the first type obtained by the sample x in the S-th iteration cycle; B x represents the counted number of labels of the second type obtained by the sample x in the s-th iteration cycle; M represents the total number of decision trees.

[0124] When is 0, it means the ideal situation where the sample x is exactly the same in the experimental group and the control group.

[0125]

[0126] Among them, K s represents the gap between the experimental group and the control group allocated in the s-th iteration cycle; represents the sample gap between the sample x being allocated to the experimental group and the control group in the s-th iteration cycle calculated by the above calculation formula (2); N represents the total number of samples.

[0127] When K s = 0, it means that the s-th iteration cycle is the most ideal i-grouping situation, that is, the experimental group and the control group are exactly the same before the experiment. However, in many cases, the result of K s = 0 cannot be obtained. Therefore, it is necessary to control K sThe smaller, the better. To avoid the iteration entering an infinite loop, the iteration optimization can be stopped by setting a gap threshold. This gap threshold can be set or adjusted according to experimental experience.

[0128] The second implementation method for stopping the iteration optimization may include: calculating the gap between the experimental group and the control group assigned in each iteration cycle; when the number of iterations is not less than a preset iteration threshold, calculating the first gap mean between the gaps corresponding to a continuous first set number of iteration cycles; calculating the second gap mean between the gaps corresponding to a continuous second set number of iteration cycles; calculating a gap change rate according to the first gap mean and the second gap mean; and when the gap change rate is less than a preset change rate threshold, determining that the current iteration cycle meets the iteration optimization stop condition.

[0129] Among them, calculating the gap between the experimental group and the control group assigned in each iteration cycle can be implemented using the above formula (3), which will not be elaborated here.

[0130] Among them, when the first set number is less than the second set number, the preset iteration threshold is generally equal to the second set number. That is, after the number of iteration cycles reaches the second set number, the above process of calculating the first gap mean and the second gap mean is carried out to reduce the consumption of computing resources. It should be noted that the essence of the second implementation method for stopping the iteration optimization is that after the number of iterations reaches the second set number, for each additional iteration cycle, the process of calculating the first gap mean, calculating the second gap mean, and calculating the gap change rate needs to be executed based on this additional iteration cycle. For example, if the second set number is 5, starting from the 5th iteration cycle, the 5th iteration cycle, the subsequent added 6th iteration cycle, the 7th iteration cycle,... all need to execute the process of calculating the first gap mean, calculating the second gap mean, and calculating the gap change rate.

[0131] Among them, the first set number, the second set number, and the preset iteration threshold can be set or adjusted according to requirements.

[0132] Among them, both the consecutive first set number of iteration cycles and the consecutive second set number of iteration cycles are determined forward from the current iteration cycle for consecutive iteration cycles. For example, if the first set number is 3 and the current iteration cycle is the 3rd iteration cycle, then starting from these 3 iteration cycles, 3 consecutive iteration cycles are determined forward (these 3 consecutive iteration cycles are the 3rd iteration cycle, the 2nd iteration cycle, and the 1st iteration cycle); for example, if the first set number is 3 and the current iteration cycle is the 6th iteration cycle, then starting from these 6 iteration cycles, 3 consecutive iteration cycles are determined forward (these 3 consecutive iteration cycles are the 6th iteration cycle, the 5th iteration cycle, and the 4th iteration cycle); for example, if the second set number is 5 and the current iteration cycle is the 6th iteration cycle, then starting from these 6 iteration cycles, 5 consecutive iteration cycles are determined forward (these 5 consecutive iteration cycles are the 6th iteration cycle, the 5th iteration cycle, the 4th iteration cycle, the 3rd iteration cycle, and the 2nd iteration cycle).

[0133] Among them, calculating the first gap mean value between the gaps corresponding to the consecutive first set number of iteration cycles can be achieved through the following calculation formula (4).

[0134] K g =(K t +K t-1 +K t-2 +…K t-g+1 ) / g (4)

[0135] Among them, K g represents the first gap mean value between the gaps corresponding to the consecutive first set number of iteration cycles; g represents the first set number; K t , K t-1 , K t-2 ,…, K t-g+1 represent the gaps corresponding to the consecutive first set number g of iteration cycles starting from the current iteration cycle t;

[0136] Among them, calculating the second gap mean value between the gaps corresponding to the consecutive second set number of iteration cycles can be achieved through the following calculation formula (5).

[0137]

[0138] Among them, K w represents the second gap mean value between the gaps corresponding to the consecutive second set number of iteration cycles; w represents the second set number; K t +K t-1 +K t-2 +…K t-w+1 represents the gaps corresponding to the consecutive second set number w of iteration cycles starting from the current iteration cycle t.

[0139] Among them, the above-mentioned calculation of the difference change rate can be achieved through the following calculation formula (6).

[0140]

[0141] Among them, Q represents the difference change rate; K g represents the first difference mean value between the differences corresponding to the first set number of consecutive iteration cycles; K w represents the second difference mean value between the differences corresponding to the second set number of consecutive iteration cycles.

[0142] Among them, the preset change rate threshold can be set according to the actual situation. For example, it can be set to 0.1.

[0143] The above two implementation methods for stopping iterative optimization both reduce the interference of manual grouping, and can effectively evaluate the similarity between the experimental group and the control group, so as to ensure that the finally obtained experimental group and control group can effectively evaluate the product to be tested.

[0144] It should be noted that for the above two implementation methods for stopping iterative optimization, the user can choose any one of the implementation methods according to the needs to achieve stopping iterative optimization.

[0145] Furthermore, as Figure 3 shown, for each iteration cycle, the specific implementation method of constructing multiple decision trees may include: performing sampling with replacement on the sample data to construct multiple sample sets; and for each sample set, performing the following operations:

[0146] Randomly select a set number of target sample features in random order from the samples included in the sample set; use the set number of randomly selected target sample features and the class labels of the samples to which the target sample features belong to construct a decision tree.

[0147] The essence of the above-mentioned sampling with replacement on the sample data to construct multiple sample sets is that for the construction process of each sample set, sampling with replacement is performed multiple times (the number of times is the same as the number of samples included in the sample data). For example, if M sample sets need to be constructed and the sample data contains N samples, then the sample data is sampled with replacement N times, and the number of samples included in each constructed sample set can be the same or different.

[0148] Among them, the number of the set number of randomly selected target sample features is much lower than the total number of sample features included in the samples. By sequential selection, it is ensured that the set number of selected target sample features is not repeated. Through the random forest training process, at each split of the target sample features, a decision tree is constructed.

[0149] Among them, each of the constructed M sample sets corresponds to a decision tree, and M decision trees are obtained through the above process. By encapsulating the M decision trees into the same model, the category label of the sample is determined by using the model to call the decision tree.

[0150] Further, the specific implementation of allocating the sample data to the experimental group and the control group for testing may include: as Figure 3 shown, for each sample in the sample data, the following steps S301 to S303 are executed:

[0151] Step S301: Calculate the score of the sample by using multiple decision trees;

[0152] For example, based on the score of the sample calculated in this step, the obtained sample set is: {(x i , s i ) | i ∈ {1, 2,..., N}}, where x i represents the i-th sample, s i represents the score of the i-th sample; N represents the total number of samples in the sample data.

[0153] Step S302: Sort the multiple samples in the sample data according to the score of the sample;

[0154] This sorting can be in ascending or descending order, and the sorted sample set is obtained: {(x j , s j ) | j ∈ {1, 2,..., N}}.

[0155] Step S303: Allocate the multiple samples in the sample data to the experimental group and the control group according to the sorting result.

[0156] Specifically, the samples sorted in the odd positions can be allocated to the experimental group, and the samples sorted in the even positions can be allocated to the control group.

[0157] That is, the samples in the sample data are allocated to the experimental group and the control group through the decision tree, making the whole allocation more accurate.

[0158] Specifically, as Figure 4 shown, the specific implementation of the above step S301 may include the following steps:

[0159] Step S401: Determine the output result of each decision tree for the sample. Among them, the output result of each decision tree for the sample is the first category label corresponding to the experimental group or the second category label corresponding to the control group;

[0160] For example, there are M decision trees, and the output results for the sample x are as follows: A decision trees give the output result for the sample x as the first category label corresponding to the experimental group, and B decision trees give the output result for the sample x as the second category label corresponding to the control group, where A + B = M.

[0161] Step S402: According to the output results of multiple decision trees for the sample, count the number of labels corresponding to the first category label and / or the number of labels corresponding to the second category label;

[0162] For example, for the above example, for the sample x, the number of labels corresponding to the first category label is A, and the number of labels corresponding to the second category label is B.

[0163] Step S403: Use the statistical results to calculate the score of the sample.

[0164] For different situations, the above step S403 can have the following several processing methods.

[0165] Processing method 1 of step S403: For the situation where the number of labels corresponding to the first category label is statistically obtained, directly determine that the ratio between the number of labels of the first category label for the sample and the number of decision trees is the score of the sample.

[0166] Processing method 2 of step S403: For the situation where the number of labels corresponding to the second category label is statistically obtained, directly determine that the ratio between the number of labels of the second category label for the sample and the number of decision trees is the score of the sample.

[0167] Processing method 3 of step S403: For the situation where the number of labels corresponding to the first category label and the number of labels corresponding to the second category label are statistically obtained, compare the number of labels of the first category label with the number of labels of the second category label; in the case where the comparison result indicates that the number of labels of the first category label is greater than or equal to the number of labels of the second category label, determine that the ratio between the number of labels of the first category label for the sample and the number of decision trees is the score of the sample; in the case where the comparison result indicates that the number of labels of the second category label is greater than the number of labels of the first category label, determine that the ratio between the number of labels of the second category label for the sample and the number of decision trees is the score of the sample.

[0168] That is: for the above sample x, when A ≥ B, s(x) = A / M; when A < B, s(x) = B / M; where A represents the number of labels of the first category label obtained by the sample x; B represents the number of labels of the second category label obtained by the sample x; s(x) represents the score of the sample x; M represents the total number of decision trees.

[0169] In addition, the samples for the control group and the experimental group are mainly allocated in a 1:1 manner. If the number of samples is large, a 1:k allocation method can also be used based on the above scheme, that is, one experimental group sample is matched with multiple control group samples to use as much data as possible.

[0170] In summary, as Figure 5 shown, the solution provided by the embodiment of the present invention can pre-group the sample data through clustering, and then train the random forest model through an iterative optimization process to obtain multiple decision trees in each iteration cycle. Multiple decision trees are used to determine the type labels for the samples, and based on the type labels determined by the multiple decision trees for the samples, the scores of the samples are calculated, and then the samples are sorted. According to the sorting results, the control group and the experimental group are re-allocated for the samples, and the samples are re-labeled, and then enter the next iteration cycle. After the iterative optimization is completed, the control group and the experimental group allocated in the last iteration cycle are obtained.

[0171] The following uses a specific embodiment of a logistics network project to illustrate in detail the process of allocating the experimental group and the control group for the sample data and the impact of the grouping results on subsequent product testing.

[0172] Among them, in this logistics network project, the samples in the sample data are relatively independent small networks in the entire large logistics network. However, in the experiments in 7 regions across the country, the number of small networks in each region is only more than 200. Based on these small numbers of samples, the following scheme is used to allocate the experimental group and the control group for the small number of samples.

[0173] Step S601: Extract multiple sample features from the sample data;

[0174] For example, for the average volume of the entities in the small network involved in the above logistics network project, it is used as the volume feature actual_compute_volume of the small network, the average value of the mileage features of the entities in the small network is used as the mileage feature mileage of the small network, and the average value of the mean square cost per 100 kilometers of the entities in the small network is used as the mean square cost per 100 kilometers unit_price of the small network. Among them, the mean square cost per 100 kilometers is also the value index of the small network.

[0175] Step S602: Use multiple sample features to perform clustering processing on multiple samples included in the sample data;

[0176] Step S603: According to the results of the clustering process, pre-divide the sample data into two sample groups for testing, and according to the results of the pre-grouping, assign category labels to each sample in the sample data;

[0177] The following operations are repeatedly executed until the loop stop condition is met:

[0178] Step S604: Based on multiple sample features included in the sample data and the class labels of the samples to which the sample features belong, construct multiple decision trees;

[0179] Step S605: Use the multiple decision trees to calculate the scores of the samples;

[0180] Step S606: Sort the multiple samples in the sample data according to the scores of the samples;

[0181] Step S607: According to the sorting result, allocate the multiple samples in the sample data to the experimental group and the control group;

[0182] Step S608: Calculate the gap between the allocated experimental group and the control group; if the number of iterations is not less than the second set quantity, execute the following Step S609; if the number of iterations is less than the second set quantity, execute the following Step S604;

[0183] Among them, the change curve of the gap between the experimental group and the control group in each iteration cycle is as Figure 7 shown. From Figure 7 it can be clearly seen that as the number of iterations increases, the gap between the experimental group and the control group gradually shrinks and fluctuates within a certain range.

[0184] Step S609: Calculate the first gap mean value between the gaps corresponding to the first set quantity of consecutive iteration cycles;

[0185] Step S610: Calculate the second gap mean value between the gaps corresponding to the second set quantity of consecutive iteration cycles;

[0186] Step S611: Calculate the gap change rate according to the first gap mean value and the second gap mean value; if the gap change rate is less than 0.1, determine that the current iteration cycle meets the iteration optimization stop condition and execute Step S612; if the current iteration cycle does not meet the iteration optimization stop condition, execute Step S604;

[0187] Step S612: Output the control group and the experimental group allocated in the current iteration cycle.

[0188] Subsequently, the above-mentioned logistics network project also uses the traditional random diversion method to divide the control group and the experimental group.

[0189] By analyzing the means of the sample characteristics of the control group and the experimental group divided by the traditional random diversion method, it is found that there is a large difference in the first and second characteristics of the samples in the control group and the experimental group divided by the random diversion method, and the difference in the third characteristic, that is, the diversion rationality index, is 3%. However, in fact, the true experimental value of this project is about 1%-2%, and 3% is already twice the deviation of the true value. Therefore, generally speaking, the random diversion result is not reasonable enough at this time.

[0190] Next, analyze the experimental group and the control group obtained by the solution provided in the embodiment of the present invention, and it is found that the solution provided in the embodiment of the present invention can well balance the characteristics of the experimental group and the control group. Among them, the difference in the value index item unit_price between the experimental group and the control group is less than 0.1%, which provides a better experimental basis for the true experimental effect of 1%-2%.

[0191] Figure 8 It is a schematic structural diagram of a sample data processing device provided in an embodiment of the present invention. As Figure 8 shown, the sample data processing device 800 may include: a pre-grouping module 801, a model processing module 802, and a data processing module 803, where

[0192] The pre-grouping module 801 is used to pre-group the sample data and assign a category label to each sample in the sample data according to the pre-grouping result;

[0193] The model processing module 802 is used to construct multiple decision trees based on multiple sample characteristics included in the sample data and the category labels of the samples to which the sample characteristics belong;

[0194] The data processing module 803 is used to use the constructed multiple decision trees to assign the sample data to the experimental group and the control group for testing.

[0195] In the embodiment of the present invention, the pre-grouping module 801 is further used to extract multiple sample characteristics from the sample data; use the multiple sample characteristics to perform clustering processing on the multiple samples included in the sample data; and pre-divide the sample data into two sample groups for testing according to the clustering processing result.

[0196] In the embodiment of the present invention, the model processing module 802 is further used to extract multiple initial sample characteristics from the sample data according to a predefined feature extraction strategy; construct a linear regression relationship between the multiple initial sample characteristics and a preset value index and a preset experimental item; and screen out multiple sample characteristics that are irrelevant to the experimental item and relevant to the value index from the multiple initial sample characteristics according to the constructed linear regression relationship.

[0197] In an embodiment of the present invention, the pre-grouping module 801 is further configured to cluster a plurality of samples included in the sample data into two sample classes.

[0198] In an embodiment of the present invention, the pre-grouping module 801 is further configured to directly exchange some samples in the two sample classes to obtain two sample groups for testing.

[0199] In an embodiment of the present invention, the pre-grouping module 801 is further configured to perform an operation for each sample class: calculate the distance from each sample in the sample class to the center point of the sample class according to the sample features; sort the plurality of samples in the sample class according to the calculated distance; swap the samples corresponding to the even sorting positions or the odd sorting positions in the two sample classes to obtain two sample groups for testing.

[0200] In an embodiment of the present invention, the data processing module 803 is further configured to, when the number of samples included in the sample data is even, control the number of samples included in the two sample groups to be equal; when the number of samples included in the sample data is odd, control the number of samples in the first sample group to be 1 more than the number of samples in the second sample group, where the first sample group corresponds to the experimental group and the second sample group corresponds to the control group.

[0201] In an embodiment of the present invention, the model processing module 802 is further configured to iteratively optimize the steps of constructing multiple decision trees and the steps of iteratively optimizing the allocation of the sample data to the experimental group and the control group for testing until the iterative optimization stop condition is met, and determine the allocation result corresponding to the iterative optimization stop condition as the processing result of the sample data.

[0202] In an embodiment of the present invention, the model processing module 802 is further configured to, for each iteration cycle other than the first iteration cycle, update the class labels of the samples in the previous iteration cycle or reassign class labels to the samples according to the allocation result of the previous iteration cycle; construct multiple decision trees based on the multiple sample features included in the sample data and the updated class labels or reallocated class labels of the samples to which the sample features belong.

[0203] In an embodiment of the present invention, the model processing module 802 is further configured to perform sampling with replacement on the sample data to construct a plurality of sample sets; for each sample set, perform an operation: randomly select a set number of target sample features in a random order from the samples included in the sample set; construct a decision tree by using the set number of randomly selected target sample features and the class labels of the samples to which the target sample features belong.

[0204] In an embodiment of the present invention, the data processing module 803 is further configured to perform operations for each sample in the sample data: calculate the score of the sample by using multiple decision trees; sort the multiple samples in the sample data according to the scores of the samples; and allocate the multiple samples in the sample data to the experimental group and the control group according to the sorting result.

[0205] In an embodiment of the present invention, the data processing module 803 is further configured to determine the output result of each decision tree for the sample, where the output result of each decision tree for the sample is a first category label corresponding to the experimental group or a second category label corresponding to the control group; count the number of labels corresponding to the first category label and / or the number of labels corresponding to the second category label according to the output results of the multiple decision trees for the sample; and calculate the score of the sample by using the counted result.

[0206] In an embodiment of the present invention, when the data processing module 803 counts the number of labels corresponding to the first category label, the data processing module 803 is further configured to directly determine that the ratio between the number of labels corresponding to the first category label of the sample and the number of decision trees is the score of the sample.

[0207] In an embodiment of the present invention, when the data processing module 803 counts the number of labels corresponding to the second category label, the data processing module 803 is further configured to directly determine that the ratio between the number of labels corresponding to the second category label of the sample and the number of decision trees is the score of the sample.

[0208] In an embodiment of the present invention, when the data processing module 803 counts the number of labels corresponding to the first category label and the number of labels corresponding to the second category label, the data processing module 803 is further configured to compare the number of labels corresponding to the first category label with the number of labels corresponding to the second category label; when the comparison result indicates that the number of labels corresponding to the first category label is greater than or equal to the number of labels corresponding to the second category label, determine that the ratio between the number of labels corresponding to the first category label of the sample and the number of decision trees is the score of the sample; and when the comparison result indicates that the number of labels corresponding to the second category label is greater than the number of labels corresponding to the first category label, determine that the ratio between the number of labels corresponding to the second category label of the sample and the number of decision trees is the score of the sample.

[0209] In an embodiment of the present invention, the model processing module 802 is further configured to, in the current iteration cycle, use the output results of multiple decision trees for each sample to count the number of labels of the first type corresponding to the experimental group and the number of labels of the second type corresponding to the control group for each sample; calculate the gap between the experimental group and the control group allocated in the current iteration cycle according to the counted number of labels of the first type and the number of labels of the second type of each sample; and determine that the current iteration cycle meets the iteration optimization stop condition when the gap is less than a preset gap threshold.

[0210] In an embodiment of the present invention, the model processing module 802 is further configured to calculate the gap between the experimental group and the control group allocated in each iteration cycle; calculate the first gap mean between the gaps corresponding to the first set number of consecutive iteration cycles when the number of iterations is not less than a preset iteration threshold; calculate the second gap mean between the gaps corresponding to the second set number of consecutive iteration cycles; calculate the gap change rate according to the first gap mean and the second gap mean; and determine that the current iteration cycle meets the iteration optimization stop condition when the gap change rate is less than a preset change rate threshold.

[0211] Figure 9 An exemplary system architecture 900 to which the sample data processing method or sample data processing apparatus according to the embodiments of the present invention can be applied is shown.

[0212] As Figure 9 shown, the system architecture 900 may include terminal devices 901, 902, a network 903, a sample processing server 904, a database 905, and a test server 906. The network 903 is used to provide a medium for communication links between the terminal devices 901, 902 and the sample processing server 904, between the sample processing server 904 and the database 905, between the database 905 and the test server 906, and between the test server 906 and the terminal devices 901, 902. The network 903 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0213] The sample processing server 904 may obtain sample data matching the user demand information from the database according to the user demand information sent by the terminal devices 901, 902, and process the sample data according to the technical solutions provided in the above embodiments to evenly distribute the samples in the sample data into the experimental group and the control group, and store the samples divided into the experimental group and the control group in the database 905 or directly send them to the corresponding test server 906. Further, the samples divided into the experimental group and the control group may also be sent to the terminal devices 901, 902 so that the user can obtain the samples divided into the experimental group and the control group through the terminal devices 901, 902.

[0214] The test server 906 can be set with one or more virtual machines, and each virtual machine is equipped with the program to be tested and the already launched program used as a control. The test server 906 can read the samples divided into the experimental group and the control group from the database 905, and process the samples of the experimental group and the control group respectively through the program to be tested and the already launched program installed on the test server 906, so as to evaluate the program to be tested, and display and provide the evaluation result of the test to the user through the terminal devices 901 and 902.

[0215] The terminal devices 901 and 902 can be various electronic devices with a display screen and supporting web browsing, including but not limited to desktop computers, smart phones, tablet computers, etc.

[0216] It should be noted that the sample data processing method provided by the embodiments of the present invention is generally completed by the sample processing server 904. Correspondingly, each module of the sample data processing device can be set in the sample processing server 904.

[0217] It should be understood, Figure 9 the numbers of the terminal devices, network, sample processing server, test server and database in

[0218] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, network, sample processing server, test server and database. Figure 10 which shows a schematic structural diagram of a computer system 1000 of a server or a terminal device suitable for implementing the embodiments of the present invention. Figure 10 The server or terminal device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0219] As Figure 10 shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the system 1000 are also stored. The CPU 1001, ROM 1002 and RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.

[0220] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as required. A removable medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1010 as required so that a computer program read therefrom is installed into the storage section 1008 as required.

[0221] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by a central processing unit (CPU) 1001, the above-described functions defined in the system of the present invention are executed.

[0222] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0223] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0224] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a pre-grouping module, a model processing module, and a data processing module. Among them, the names of these modules do not constitute a limitation on the module itself in some cases. For example, the data processing module can also be described as "the module that allocates sample data to the experimental group and the control group for testing".

[0225] As another aspect, the present invention also provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device includes: pre-grouping sample data, and based on the result of the pre-grouping, assigning a class label to each sample in the sample data; constructing multiple decision trees based on multiple sample features included in the sample data and the class label of the sample to which the sample feature belongs; using the constructed multiple decision trees to allocate the sample data to the experimental group and the control group for testing.

[0226] According to the technical solution of the embodiments of the present invention, by pre-grouping the sample data to assign a class label to each sample based on the result of the pre-grouping, the class label indicates that the sample belongs to the experimental group or the control group. Since the sample features can better reflect the characteristics and essence of the sample, the experimental group, and the control group, constructing a decision tree based on multiple sample features and class labels included in the sample data can enable the decision tree to provide a more accurate grouping. Then, using the constructed multiple decision trees to allocate the sample data to the experimental group and the control group for testing can ensure that the sample data in the experimental group is similar to the sample data in the control group even with a relatively small number of samples. Therefore, subsequent testing using similar experimental groups and control groups can effectively improve the authenticity and accuracy of the test results for the product.

[0227] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for processing sample data, characterized in that, Including: Pre-grouping the sample data, and assigning class labels to each sample in the sample data according to the results of the pre-grouping; Constructing multiple decision trees based on multiple sample features included in the sample data and class labels of samples to which the sample features belong; Using the multiple constructed decision trees to assign the sample data to an experimental group and a control group for testing.

2. The method for processing sample data according to claim 1, characterized in that, The pre-grouping of the sample data includes: Extracting multiple sample features from the sample data; Using the multiple sample features to perform clustering processing on multiple samples included in the sample data; According to the results of the clustering processing, pre-dividing the sample data into two sample groups for testing.

3. The method for processing sample data according to claim 2, characterized in that, The extracting multiple sample features from the sample data includes: Extracting multiple initial sample features from the sample data according to a predefined feature extraction strategy; Constructing linear regression relationships between the multiple initial sample features and a preset value index and a preset experimental item respectively; According to the constructed linear regression relationships, screening out multiple sample features that are irrelevant to the experimental item and relevant to the value index from the multiple initial sample features; and / or The performing clustering processing on multiple samples included in the sample data includes: Clustering multiple samples included in the sample data into two sample classes.

4. The method for processing sample data according to claim 3, characterized in that, For the case where the clustering processing is to cluster multiple samples included in the sample data into two sample classes, The pre-dividing the sample data into two sample groups for testing includes: Directly exchanging some samples in the two sample classes to obtain two sample groups for testing; Or For each of the sample classes, perform the following operations: According to the sample features, calculating the distance from each sample in the sample class to the center point of the sample class; Sorting the multiple samples in the sample class according to the calculated distances; Exchanging samples corresponding to even sorting positions or odd sorting positions in the two sample classes to obtain two sample groups for testing; and / or The sample data processing method further includes: When the number of samples included in the sample data is even, controlling the number of samples included in the two sample groups to be equal; When the number of samples included in the sample data is odd, controlling the number of samples in the first sample group to be 1 more than the number of samples in the second sample group, where the first sample group corresponds to the experimental group and the second sample group corresponds to the control group.

5. The method for processing sample data according to claim 1, characterized in that, It also includes: Iteratively optimizing the steps of constructing multiple decision trees and the steps of iteratively optimizing the assignment of the sample data to the experimental group and the control group for testing until the iterative optimization stop condition is met, and determining the assignment result corresponding to the iterative optimization stop condition as the processing result of the sample data, where For each iteration cycle except the first iteration cycle, The steps of constructing multiple decision trees include: According to the assignment result of the previous iteration cycle, updating the class labels of the samples in the previous iteration cycle or re-assigning class labels to the samples; Construct multiple decision trees based on multiple sample features included in the sample data and the updated class labels of the samples to which the sample features belong or the class labels reassigned to them.

6. The sample data processing method according to claim 1 or 5, wherein The constructing of multiple decision trees includes: Perform sampling with replacement on the sample data to construct multiple sample sets; For each of the sample sets, perform the following operations: Randomly select a set number of target sample features from the samples included in the sample set in a random order; Construct a decision tree using the set number of randomly selected target sample features and the class labels of the samples to which the target sample features belong.

7. The sample data processing method according to claim 1 or 5, wherein The allocating of the sample data to the experimental group and the control group for testing includes: For each sample in the sample data, perform the following operations: Calculate the score of the sample using multiple decision trees; Sort the multiple samples in the sample data according to the scores of the samples; Allocate the multiple samples in the sample data to the experimental group and the control group according to the sorting result.

8. The sample data processing method according to claim 7, wherein The calculating of the score of the sample includes: Determine the output result of each decision tree for the sample, where the output result of each decision tree for the sample is the first class label corresponding to the experimental group or the second class label corresponding to the control group; According to the output results of multiple decision trees for the sample, count the number of labels corresponding to the first class label and / or the number of labels corresponding to the second class label; Calculate the score of the sample using the counted results.

9. The sample data processing method according to claim 5, wherein The sample data processing method further includes: In the current iteration cycle, use the output results of multiple decision trees for each sample to count the number of labels of the first type corresponding to the experimental group and the number of labels of the second type corresponding to the control group for each sample; Calculate the gap between the experimental group and the control group allocated in the current iteration cycle according to the counted number of labels of the first type and the number of labels of the second type for each sample; When the gap is less than a preset gap threshold, determine that the current iteration cycle meets the iteration optimization stop condition; and / or The sample data processing method further includes: Calculate the gap between the experimental group and the control group allocated in each iteration cycle; When the number of iterations is not less than a preset iteration threshold, Calculate the first gap mean between the gaps corresponding to the first set number of consecutive iteration cycles; Calculate the second gap mean between the gaps corresponding to the second set number of consecutive iteration cycles; Calculate the gap change rate according to the first gap mean and the second gap mean; When the gap change rate is less than a preset change rate threshold, determine that the current iteration cycle meets the iteration optimization stop condition.

10. A sample data processing device, wherein It includes: A pre-grouping module, a model processing module, and a data processing module, where The pre-grouping module is used to pre-group the sample data and, according to the pre-grouping result, assign class labels to each sample in the sample data; The model processing module is configured to construct multiple decision trees based on multiple sample features included in the sample data and class labels of the samples to which the sample features belong. The data processing module is configured to use the multiple constructed decision trees to allocate the sample data to an experimental group and a control group for testing.

11. An electronic device, wherein Comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-9.

12. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-9 is implemented.

Citation Information

Cited By

  • Sample distribution method and device and computer equipment

    CN122112575A