Method and system for determining experimental data, storage medium and computer program product

By constructing a multi-dimensional spatial data structure and generating the nearest control group user samples, the problem of the control group and the experimental group not maintaining the same distribution in the prior art is solved, and the efficiency and accuracy of business activities and business product effect evaluation are improved.

CN120145173APending Publication Date: 2025-06-13CHINA UNIONPAY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411715905.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When evaluating business activities and business product effects, the user of the control group and the experimental group performs a stratified sampling, which may lead to the control group being unable to maintain the same distribution as the experimental group in terms of various indicators, and the sampling time consumed exponentially increases with the increase of users and the number of indicators, affecting the accuracy and efficiency of the assessment.

Method used

By constructing a multi-dimensional spatial data structure based on the user characteristics of the experimental group and the control group, the user samples of the control group are determined to determine the closest control group to the user sample of each experimental group, the control group is generated, and the generation of the control group is optimized through distance information and sample mapping information to ensure that the control group and the experimental group maintain the same distribution in each user characteristics.

Benefits of technology

The efficiency and accuracy of the generation of the control group is improved, the accuracy and efficiency of business activities and business product effectiveness evaluation are ensured, and the consumption of computing resources is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145173A_ABST
    Figure CN120145173A_ABST
Patent Text Reader

Abstract

The invention relates to a method and a system for determining experimental data, a computer readable storage medium for implementing the method and a computer program product. According to one aspect of the invention, the method for determining experimental data comprises the following steps: obtaining a set of first user samples based on user features of a plurality of first users in an experimental group, and obtaining a set of second user samples based on user features of a plurality of second users in a candidate control group; constructing a multi-dimensional spatial data structure based on the set of the second user samples; determining a second user sample closest to each first user sample in the multi-dimensional spatial data structure to generate a first control group of the experimental group; and determining a plurality of second user samples associated with each first user sample in the multi-dimensional spatial data structure, and determining a second control group of the experimental group based on the distance information between each first user sample and the plurality of second user samples and the sample mapping information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a method for determining experimental data, a system, a computer-readable storage medium for implementing the above method, and a computer program product. Background Art

[0002] With the rapid development of information technology, the means of promoting business using various business activities and business products are gradually shifting towards refinement and intelligence, and the evaluation of the effects of various business activities and business products is an important task. Currently, for the evaluation of the effects of business activities and business products, it is usually necessary to divide users into an experimental group and a control group. In the same time dimension, users in the experimental group can participate in business activities or use business products, while users in the control group do not participate in business activities or do not use business products. By statistically analyzing the differences between the users in the experimental group and the control group in various indicators (such as the number of transactions, transaction amount, etc.), the effects of business activities and business products are evaluated.

[0003] Generally speaking, the control group and the experimental group need to be identically distributed in various indicators to ensure accurate analysis and evaluation of the effects of business activities and business products. Currently, artificial rule sampling methods and stratified sampling methods are generally used to obtain the control group. In the artificial rule sampling method, for example, if it is required that users be identically distributed in terms of age and the number of transactions, then user samples are respectively drawn from the user sample set according to the user distribution in each age group and each transaction number interval to obtain the control group. In the stratified sampling method, user characteristics can be divided into multiple interval segments and sampling can be carried out according to the proportion of users in each interval segment. For example, age can be divided into segments of under 30 years old, 30 - 60 years old, and over 60 years old. If the proportion of users over 60 years old is 0.2, then 20% of the users over 60 years old are drawn from the user sample set. If the proportion of users under 30 years old is 0.3, then 30% of the users under 30 years old are drawn from the user sample set, and so on to obtain the control group.

[0004] However, the control group obtained by the above sampling methods may not be identically distributed with the experimental group in various indicators, and the time consumption of the above sampling methods shows exponential growth with the increase in the number of users and the number of indicators, thus affecting the accuracy and efficiency of analyzing and evaluating the effects of business activities and business products. Summary of the Invention

[0005] In order to solve or at least alleviate one or more of the above problems, the following technical solutions are provided.

[0006] According to a first aspect of the present application, there is provided a method for determining experimental data, the method comprising the following steps: obtaining a set of first user samples based on the user characteristics of a plurality of first users in an experimental group and obtaining a set of second user samples based on the user characteristics of a plurality of second users in a candidate control group; constructing a multi-dimensional space data structure based on the set of second user samples; determining, in the multi-dimensional space data structure, the second user sample closest to each first user sample to generate a first control group of the experimental group; and determining, in the multi-dimensional space data structure, a plurality of second user samples associated with each first user sample and determining a second control group of the experimental group based on the distance information and sample mapping information between each first user sample and the plurality of second user samples.

[0007] According to the method for determining experimental data according to an embodiment of the present application, wherein obtaining a set of first user samples based on the user characteristics of a plurality of first users in an experimental group comprises: determining abnormal characteristics in the user characteristics of the plurality of first users and processing the abnormal characteristics to obtain the user characteristics of the plurality of preprocessed first users; converting the user characteristics of the plurality of preprocessed first users into first numerical characteristics; and performing a normalization process on the first numerical characteristics to obtain the set of first user samples.

[0008] According to the method for determining experimental data according to an embodiment of the present application or any of the above embodiments, wherein obtaining a set of second user samples based on the user characteristics of a plurality of second users in a candidate control group comprises: determining abnormal characteristics in the user characteristics of the plurality of second users and processing the abnormal characteristics to obtain the user characteristics of the plurality of preprocessed second users; converting the user characteristics of the plurality of preprocessed second users into second numerical characteristics; and performing a normalization process on the second numerical characteristics to obtain the set of second user samples.

[0009] According to the method for determining experimental data according to an embodiment of the present application or any of the above embodiments, wherein the user characteristics of the first user and the user characteristics of the second user each comprise one or more of the following: user natural attributes, user transaction attributes, user activity attributes.

[0010] According to the method for determining experimental data according to an embodiment of the present application or any of the above embodiments, wherein constructing a multi-dimensional space data structure based on the set of second user samples comprises: taking each second user sample in the set of second user samples as a data point in a multi-dimensional space to obtain a data set; determining dividing planes of a plurality of dimensions based on the data points and dividing the data set based on the division of the multi-dimensional space by the dividing planes; and constructing the multi-dimensional space data structure based on the division result of the data set.

[0011] The method for determining experimental data according to an embodiment of the present application or any one of the above embodiments, wherein the multi-dimensional space data structure is a multi-dimensional space partition tree structure.

[0012] The method for determining experimental data according to an embodiment of the present application or any one of the above embodiments, wherein determining a second user sample closest to each first user sample in the multi-dimensional space data structure to generate a first control group of the experimental group includes: for each first user sample, searching in parallel in the multi-dimensional space data structure for a second user sample closest to each first user sample; and generating a first control group of the experimental group based on the second user sample closest to each first user sample.

[0013] The method for determining experimental data according to an embodiment of the present application or any one of the above embodiments, wherein searching in parallel in the multi-dimensional space data structure for a second user sample closest to each first user sample includes: determining a search path for each first user sample and searching in parallel in the multi-dimensional space data structure for a candidate second user sample closest to each first user sample according to the search path; determining a distance between each first user sample and the candidate second user sample closest to each first user sample; and selectively updating the search path based on the distance and searching in parallel in the multi-dimensional space data structure for a second user sample closest to each first user sample according to the updated search path.

[0014] The method for determining experimental data according to an embodiment of the present application or any one of the above embodiments, wherein determining a plurality of second user samples associated with each first user sample in the multi-dimensional space data structure includes: for each first user sample, searching in parallel in the multi-dimensional space data structure for a predetermined number of second user samples adjacent to each first user sample; and determining the predetermined number of second user samples adjacent to each first user sample searched as the plurality of second user samples associated with each first user sample.

[0015] The method for determining experimental data according to an embodiment of the present application or any one of the above embodiments, wherein the distance information and sample mapping information between each of the first user samples and the plurality of second user samples are determined by the following method: determining the distance between each of the first user samples and each of the second user samples in the multi-dimensional space data structure and determining the distance matrix generated based on the distance as the distance information; and determining the mapping relationship between the first user sample and the second user sample corresponding to each element in the distance matrix and determining the sample mapping matrix generated based on the mapping relationship as the sample mapping information.

[0016] The method for determining experimental data according to an embodiment of the present application or any one of the above embodiments, wherein determining the second control group of the experimental group based on the distance information and sample mapping information between each of the first user samples and the plurality of second user samples includes repeatedly performing the following steps until the distance matrix is an empty matrix: determining the element with the smallest element value in the distance matrix; determining the first target user sample and the second target user sample corresponding to the element with the smallest element value based on the sample mapping matrix; determining the second control group of the experimental group based on the first target user sample and the second target user sample corresponding to the element with the smallest element value; deleting the rows associated with the first target user sample in the distance matrix and the sample mapping matrix respectively; and setting the value of the element associated with the second target user sample in the distance matrix to infinity.

[0017] The method for determining experimental data according to an embodiment of the present application or any one of the above embodiments, wherein the method further includes: constructing a first loss function based on the distance mean between each of the first user samples in the set of first user samples and each of the second user samples in the first control group; constructing a second loss function based on the distance mean between each of the first user samples in the set of first user samples and each of the second user samples in the second control group; and performing a same distribution test on each of the first user samples in the set of first user samples and each of the second user samples in the first control group based on the value of the first loss function, and performing a same distribution test on each of the first user samples in the set of first user samples and each of the second user samples in the second control group based on the value of the second loss function.

[0018] According to a second aspect of the present application, there is provided a system for determining experimental data, the system comprising: a memory; a processor coupled to the memory; and a computer program stored on the memory and running on the processor, the running of the computer program causing the following operations: obtaining a set of first user samples based on the user characteristics of a plurality of first users within an experimental group and obtaining a set of second user samples based on the user characteristics of a plurality of second users within a candidate control group; constructing a multi-dimensional space data structure based on the set of second user samples; determining, in the multi-dimensional space data structure, the second user sample closest to each first user sample to generate a first control group for the experimental group; and determining, in the multi-dimensional space data structure, a plurality of second user samples associated with each first user sample and determining a second control group for the experimental group based on the distance information and sample mapping information between each first user sample and the plurality of second user samples.

[0019] According to a third aspect of the present application, there is provided a computer-readable storage medium comprising instructions that, when run, execute the steps of the method for determining experimental data according to the first aspect of the present application.

[0020] According to a fourth aspect of the present application, there is provided a computer program product comprising instructions that, when executed by a processor, implement the steps of the method for determining experimental data according to the first aspect of the present application.

[0021] The solution for determining experimental data proposed according to one or more embodiments of the present application can construct a multi-dimensional space data structure based on a set of user samples within a candidate control group and generate a control group for the experimental group through the multi-dimensional space data structure, ensuring that the control group has the same distribution as the experimental group in terms of each user characteristic, and by means of the multi-dimensional space data structure, it is possible to search in parallel for the second user samples within the candidate control group that are closest or associated with each first user sample within the experimental group to generate the control group, improving the efficiency of generating the control group and saving computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and / or other aspects and advantages of the present application will become clearer and easier to understand through the following description of various aspects in conjunction with the accompanying drawings, in which the same or similar units are denoted by the same reference numerals. The drawings include:

[0023] Figure 1 A flowchart showing a method for determining experimental data according to one or more embodiments of the present application.

[0024] Figure 2 A flowchart showing a method for performing a same-distribution test on an experimental group and a control group according to one or more embodiments of the present application.

[0025] Figure 3 A block diagram of a system for determining experimental data according to one or more embodiments of the present application is shown. Detailed implementation manners

[0026] The present application will be described more fully hereinafter with reference to the accompanying drawings, in which illustrative embodiments of the present application are shown. However, the present application may be implemented in different forms and should not be construed as limited to the embodiments set forth herein. The above-described embodiments are provided to make the disclosure of the present application thorough and complete, and to more fully convey the scope of protection of the present application to those skilled in the art.

[0027] In this specification, terms such as "comprising" and "including" mean that the technical solutions of the present application do not exclude the presence of other elements and steps that are not directly or explicitly recited in addition to the elements and steps directly and explicitly recited in the specification and claims.

[0028] Unless otherwise specified, terms such as "first" and "second" do not denote an order of elements in terms of time, space, size, etc., but are merely used to distinguish between elements.

[0029] In the context of the present application, an experimental group refers to a group of participants or samples that receive an experimental intervention or treatment. The samples in the experimental group can receive specific intervention measures, new methods, or experimental conditions, and the results of the experimental group are used to evaluate the effect of the experimental intervention or treatment. A control group refers to a group of participants or samples that do not receive an experimental intervention or treatment, and the control group can be used as a benchmark to compare with the results of the experimental group to determine the effect of the experimental intervention or treatment. As an example, the above method can be used to analyze the effect of a new product, algorithm, or marketing campaign. For example, users participating in a marketing campaign can be used as the experimental group and a control group can be generated based on the experimental group, and the effect of the marketing campaign can be analyzed by analyzing the behavioral data of the users in the experimental group and the control group.

[0030] Hereinafter, various exemplary embodiments of the present application will be described in detail with reference to the accompanying drawings.

[0031] Figure 1 A flowchart of a method for determining experimental data according to one or more embodiments of the present application is shown.

[0032] As Figure 1 shown, in step S101, a set of first user samples is obtained based on the user characteristics of a plurality of first users in the experimental group, and a set of second user samples is obtained based on the user characteristics of a plurality of second users in the candidate control group.

[0033] Optionally, in step S101, for the user characteristics of multiple first users in the experimental group, abnormal characteristics in the user characteristics can be detected and processed to obtain preprocessed user characteristics, the preprocessed user characteristics can be converted into numerical characteristics, and the numerical characteristics can be normalized to obtain a set of first user samples. Optionally, in step S101, for the user characteristics of multiple second users in the candidate control group, abnormal characteristics in the user characteristics can be detected and processed to obtain preprocessed user characteristics, the preprocessed user characteristics can be converted into numerical characteristics, and the numerical characteristics can be normalized to obtain a set of second user samples.

[0034] In one embodiment, the user characteristics of the first user and the user characteristics of the second user may each include user natural attributes (such as gender, age, native place, usual location, etc.), user transaction attributes (such as transaction amount, transaction time, number of transaction pens, transaction scenario, etc.), user activity attributes (such as user satisfaction, user consumption willingness, last use time, etc.), and the like. In the detection and processing of abnormal characteristics, taking the transaction amount as an example of user characteristics, the interquartile method can be used to detect abnormal transaction amounts in the transaction amount, and the detected abnormal transaction amount can be used as an abnormal judgment threshold to remove the transaction amount above or below the abnormal judgment threshold from the transaction amount, or map the transaction amount above or below the abnormal judgment threshold to a transaction amount interval, for example, map the transaction amount above the abnormal judgment threshold to more than one hundred thousand yuan. Exemplarily, when using the interquartile method to detect abnormal transaction amounts in the transaction amount, the transaction amounts can be sorted by size and the first quartile and the third quartile of the transaction amount can be determined, the interquartile range can be determined based on the difference between the third quartile and the first quartile, and the abnormal transaction amounts in the transaction amount can be detected based on the first quartile, the third quartile, and the interquartile range.

[0035] In one embodiment, user features can be classified into categorical features, ordinal features, and numerical features. Among them, categorical features can be understood as features for classifying users, ordinal features can be understood as features that can be sorted according to degree, and numerical features can be understood as features that can be represented by numerical values. Exemplarily, the user's gender, native place, usual location, and transaction scenario can be classified as categorical features, the user's age, transaction amount, number of transactions, user usage duration, and last usage time can be classified as numerical features, and user satisfaction and user consumption willingness can be classified as ordinal features. In one embodiment, categorical features and ordinal features can be converted into numerical features. For example, one-hot encoding processing can be used for categorical features to convert them into numerical features, and ordinal features can be mapped to different intensity numbers according to intensity to be converted into numerical features. Exemplarily, "very satisfied", "satisfied", "average", "dissatisfied", and "very dissatisfied" in user satisfaction can be mapped to 5, 4, 3, 2, and 1 respectively. In one embodiment, the mean and variance of numerical features can be determined and the numerical features can be standardized based on the determined mean and variance, so as to scale the numerical features. For example, the numerical features can be mapped to between 0 and 1. In one embodiment, different weights can be assigned to different user features, so as to, for example, focus on one or more user features when analyzing marketing activities.

[0036] In one embodiment, the set of first user samples and the set of second user samples can be respectively represented as sample matrices. Each row in the sample matrix can represent a user in the set of user samples, and each column in the sample matrix can represent a user feature. Exemplarily, assuming that there are 10 users in the set of user samples, and the user features of each user include age, gender, and transaction amount, then the set of user samples can be represented as a 10*3 sample matrix, where each row represents a user, and each column represents the numerical features of the user's age, gender, and transaction amount after standardization processing.

[0037] In step S103, a multi-dimensional space data structure is constructed based on the set of second user samples.

[0038] Optionally, in step S103, each second user sample in the set of second user samples can be used as a data point in a multi-dimensional space to obtain a data set. Based on the data points, splitting planes for multiple dimensions are determined, and the multi-dimensional space is divided based on these splitting planes to partition the data set. Then, a multi-dimensional space data structure is constructed based on the partitioning result of the data set. Exemplarily, the number of dimensions of the multi-dimensional space can be set to the number of types of user features of the second user. For example, assume there are 10 users in the set of second user samples, and the user features of each user include age, gender, and transaction amount. Then the data set can include 10 data points in a three-dimensional space, and each data point represents the normalized numerical features of a user's age, gender, and transaction amount. In one embodiment, the multi-dimensional space data structure can be implemented as a multi-dimensional space partitioning tree structure, i.e., the structure of a KD (K-Dimension) tree. By constructing a multi-dimensional space data structure based on the set of second user samples, that is, partitioning the multi-dimensional space in multiple dimensions to divide the search space into smaller regions, it is possible to avoid searching the user samples of the entire data set, thereby reducing the number of data points that need to be searched, which is beneficial to improving the efficiency of generating a control group in the subsequent process and saving computing resources. Since the KD tree is a tree structure for storing points in a K-dimensional space for spatial search, its average time complexity is O(logN), which is much smaller than the computational time of linear partitioning. Moreover, the search for candidate control group user samples for each experimental group's user samples does not affect each other, that is, it supports parallel computing and cluster operations, thus significantly reducing the computational time consumed in generating the control group, especially when the number of user samples in the experimental group and the number of their user features are large.

[0039] In step S105, the second user sample closest to each first user sample is determined in the multi-dimensional space data structure to generate the first control group of the experimental group.

[0040] Optionally, in step S105, for each first user sample, the second user sample closest to each first user sample can be searched in parallel in the multi-dimensional space data structure, and the second user sample closest to each first user sample can be used as the first control group of the experimental group. Optionally, in step S105, the search path for each first user sample can be determined and the candidate second user sample closest to each first user sample can be searched in parallel in the multi-dimensional space data structure according to the search path, the distance between each first user sample and the candidate second user sample closest to each first user sample can be determined, and the search path can be selectively updated based on the distance and the second user sample closest to each first user sample can be searched in parallel in the multi-dimensional space data structure according to the updated search path. For example, in the multi-dimensional space partition tree structure, the tree can be recursively traversed downward from the root node by the nearest neighbor search method until the leaf node is found, and the search path can be updated during the backtracking process to search for the second user sample closest to each first user sample.

[0041] In one embodiment, after generating the first control group of the experimental group, a first loss function can be constructed based on the average distance (e.g., the average Euclidean distance, the average absolute value distance, etc.) between each first user sample in the set of first user samples and each second user sample in the first control group, and a same distribution test can be performed on each first user sample in the set of first user samples and each second user sample in the first control group based on the value of the first loss function, that is, whether each first user sample in the set of first user samples and each second user sample in the first control group have the same probability distribution in terms of each user feature. Exemplarily, the first loss function Loss can be constructed by the following formula (1) X,Y :

[0042]

[0043] where n represents the number of first user samples in the experimental group, m represents the number of user features, x i,dim represents the user feature of the first user, and y j,dim represents the user feature of the second user in the first control group.

[0044] When the value of the first loss function is less than the loss threshold (e.g., set to 2 or 3), it can be determined that each first user sample in the set of first user samples and each second user sample in the first control group satisfy the same distribution; conversely, it can be determined that each first user sample in the set of first user samples and each second user sample in the first control group do not satisfy the same distribution. In one embodiment, when it is determined that each first user sample in the set of first user samples and each second user sample in the first control group do not satisfy the same distribution, multiple second users in the candidate control group can be updated, and the first control group of the experimental group can be regenerated based on the updated candidate control group, for example, referring to the descriptions of steps S101, S103, and S105 above.

[0045] In step S107, multiple second user samples associated with each first user sample are determined in the multi-dimensional space data structure, and the second control group of the experimental group is determined based on the distance information and sample mapping information between each first user sample and the multiple second user samples.

[0046] Optionally, in step S107, for each first user sample, a predetermined number of second user samples adjacent to each first user sample can be searched in parallel in the multi-dimensional space data structure, and the predetermined number of second user samples adjacent to each first user sample found can be determined as the multiple second user samples associated with each first user sample. For example, in the multi-dimensional space partition tree structure, a predetermined number of second user samples adjacent to each first user sample can be searched in parallel by the k-nearest neighbor search method. Exemplarily, the predetermined number can be set based on the number of first user samples and the minimum threshold number, for example, set to the number between the minimum threshold number (e.g., 100) and the number of first user samples at a predetermined ratio (e.g., 10%).

[0047] Optionally, in step S107, the distance between each first user sample and each second user sample among the multiple second user samples associated with the first user sample can be determined in the multi-dimensional space data structure, and the distance matrix generated based on the distance can be determined as the distance information, and the mapping relationship between the first user sample and the second user sample corresponding to each element in the distance matrix can be determined, and the sample mapping matrix generated based on the mapping relationship can be determined as the sample mapping information.

[0048] Exemplarily, assume that there are 5 users in the set of the first user samples of the experimental group and the user characteristics of each user include age, gender, and transaction amount, and there are 50 users in the set of the second user samples of the candidate control group and the user characteristics of each user include age, gender, and transaction amount. For each first user sample, a predetermined number (e.g., 3) of second user samples adjacent to each first user sample can be searched in parallel in the multi-dimensional space data structure as the multiple second user samples associated with each first user sample, that is, for each user sample of the experimental group, 3 user samples of the candidate control group adjacent to it are determined, and the distance between each user sample of the experimental group and each of the 3 user samples of the candidate control group is determined in the multi-dimensional space data structure and a distance matrix is generated based on the distance. For example, the generated distance matrix Distance can be expressed as:

[0049]

[0050] Among them, the elements in the first row of the distance matrix Distance respectively represent the distances between the first user of the experimental group and each of the 3 user samples of the candidate control group associated with the first user, the elements in the second row of the distance matrix Distance respectively represent the distances between the second user of the experimental group and each of the 3 user samples of the candidate control group associated with the second user, and so on.

[0051] Exemplarily, the mapping relationship between the user samples of the experimental group and the user samples of the candidate control group corresponding to each element in the distance matrix can be determined and a sample mapping matrix is generated based on the mapping relationship. For example, the generated sample mapping matrix Index can be expressed as:

[0052]

[0053] Among them, the element 10 in the first row and the first column of the sample mapping matrix Index indicates that the element 0.2 in the first row and the first column of the distance matrix Distance represents the distance between the first user of the experimental group and the 10th user of the candidate control group, the element 33 in the third row and the second column of the sample mapping matrix Index indicates that the element 0.6 in the third row and the second column of the distance matrix Distance represents the distance between the third user of the experimental group and the 33rd user of the candidate control group, and so on.

[0054] Optionally, in step S107, after determining the distance matrix and the sample mapping matrix, the following steps may be cyclically executed until the distance matrix becomes an empty matrix: Determine the element with the smallest element value in the distance matrix; Based on the sample mapping matrix, determine the first target user sample and the second target user sample corresponding to the element with the smallest element value; Determine the second control group of the experimental group based on the first target user sample and the second target user sample corresponding to the element with the smallest element value; Delete the rows associated with the first target user sample in the distance matrix and the sample mapping matrix respectively; And set the value of the element associated with the second target user sample in the distance matrix to infinity. Exemplarily, taking the above distance matrix Distance and sample mapping matrix Index as examples, the control group set can be first set to an empty set, and then the element with the smallest element value 0.1 in the distance matrix Distance can be determined. Based on the sample mapping matrix Index, it can be determined that the element 0.1 represents the distance between the second user in the experimental group and the 37th user in the candidate control group. Thus, the 37th user in the candidate control group can be used as the control group user of the second user in the experimental group and added to the control group set. Then, the second row is deleted in the distance matrix Distance and the sample mapping matrix Index respectively to indicate that the determination of the control group user of the second user in the experimental group is completed, and the value of the element associated with the 37th user in the candidate control group in the distance matrix Distance is set to infinity. For example, based on the sample mapping matrix Index, the value of the element in the 5th row and the 3rd column in the distance matrix Distance is set to infinity. Exemplarily, the above operations can be cyclically executed to reduce the order of the distance matrix Distance and the sample mapping matrix Index and generate corresponding control group users for each user in the experimental group.

[0055] In one embodiment, after generating the second control group of the experimental group, a second loss function may be constructed based on the distance mean (e.g., Euclidean distance mean, absolute value distance mean, etc.) between each first user sample in the set of first user samples and each second user sample in the second control group, and a same distribution test may be performed on each first user sample in the set of first user samples and each second user sample in the second control group based on the value of the second loss function, that is, whether each first user sample in the set of first user samples and each second user sample in the second control group have the same probability distribution in terms of each user feature. Similarly, the second loss function can be constructed with reference to the above formula (1), for example, replacing y j,dim with the user feature of the second user in the second control group.

[0056] When the value of the second loss function is less than the loss threshold (e.g., set to 2 or 3), it can be determined that each first user sample in the set of first user samples and each second user sample in the second control group satisfy the same distribution; conversely, it can be determined that each first user sample in the set of first user samples and each second user sample in the second control group do not satisfy the same distribution. In one embodiment, when it is determined that each first user sample in the set of first user samples and each second user sample in the second control group do not satisfy the same distribution, multiple second users in the candidate control group can be updated, and the second control group of the experimental group can be regenerated based on the updated candidate control group, for example, referring to the descriptions of steps S101, S103, and S107 above.

[0057] It should be noted that the first control group of the experimental group generated in step S105 may include duplicate user samples in the candidate control group, and the second control group of the experimental group generated in step S107 does not include duplicate user samples in the candidate control group. It should be noted that the execution order of step S105 and step S107 is not limited to Figure 1 the order shown in. Without departing from the spirit and scope of the present application, step S105 and step S107 can be executed in parallel or step S107 can be executed before step S105.

[0058] In one embodiment, users participating in the marketing activity can be used as the experimental group and by means of Figure 1 the method shown to generate the first control group and the second control group according to the experimental group, where the number of user samples in the experimental group, the random sample group, the first control group, and the second control group is the same, and hypothesis tests such as the KS test, the T test, and the F test are respectively performed on the experimental group and the random sample group, the first control group, and the second control group. The results of this hypothesis test (e.g., >0.05 represents passing the hypothesis test) can be compared with the results of the same distribution test using the loss function, as shown in Table 1 below:

[0059] Table 1

[0060]

[0061]

[0062] Based on the hypothesis test results of the KS test, T test, and F test in Table 1 above and the same distribution test results of the loss function, compared with the random sample group (i.e., randomly selected user samples), the control group generated by the method for determining experimental data according to one or more embodiments of the present application can ensure the same distribution as the experimental group. Experiments show that when the number of user samples in the experimental group is 475,692, the time consumed to generate the first control group and the second control group with a sample size of 475,692 according to the method for determining experimental data proposed in one or more embodiments of the present application is approximately 300 seconds, and it can pass the hypothesis tests of the KS test, T test, and F test.

[0063] In one embodiment, users participating in the marketing activity can be used as the experimental group and the second control group can be generated based on the experimental group by means of the Figure 1 method shown, where the number of user samples in the experimental group is different from that in the random sample group and the second control group, and the hypothesis tests of the KS test, T test, and F test are respectively performed on the experimental group, the random sample group, and the second control group. The hypothesis test results (for example, >0.05 represents passing the hypothesis test) can be compared with the same distribution test results using the loss function, as shown in Table 2 below:

[0064] Table 2

[0065]

[0066] Based on the hypothesis test results of the KS test, T test, and F test in Table 2 above and the same distribution test results of the loss function, compared with the random sample group (i.e., randomly selected user samples), the control group generated by the method for determining experimental data according to one or more embodiments of the present application can ensure the same distribution as the experimental group. Experiments show that when the number of user samples in the experimental group is 193,012, the time consumed to generate the second control group with a sample size of 38,602 according to the method for determining experimental data proposed in one or more embodiments of the present application is approximately 30 seconds, and it can pass the hypothesis tests of the KS test, T test, and F test.

[0067] In one embodiment, the first control group can be generated based on the experimental group by means of the Figure 1 method shown and the same distribution tests of the KS test, T test, and F test are performed on the experimental group and the first control group in terms of multiple user characteristics. The same distribution test results are shown in Table 3 below:

[0068] Table 3

[0069]

[0070] Based on the results of the same distribution test in Table 3 above, compared with the random sample group (i.e., randomly selected user samples), the control group generated by the method for determining experimental data according to one or more embodiments of the present application can ensure the same distribution as the experimental group in terms of multiple user characteristics.

[0071] The method for determining experimental data proposed according to one or more embodiments of the present application can construct a multi-dimensional space data structure based on the set of user samples in the candidate control group and generate a control group for the experimental group through the multi-dimensional space data structure, ensuring that the control group has the same distribution as the experimental group in terms of each user characteristic. Moreover, by means of the multi-dimensional space data structure, for each first user sample in the experimental group, the second user sample in the candidate control group that is closest or associated with it can be searched in parallel to generate the control group, improving the efficiency of generating the control group and saving computing resources. The method for determining experimental data proposed according to one or more embodiments of the present application can quickly generate a control group that has the same distribution as the experimental group in terms of multiple user characteristics when the number of user samples in the experimental group is large. The generated control group can be used to analyze the effects of new products, algorithms, and marketing activities. For example, the effects of marketing activities can be analyzed by analyzing the behavior data of users in the experimental group and the control group (such as the number of times of participating in marketing activities, consumption amount, user satisfaction, etc.).

[0072] Figure 2 The flowchart of the method for performing the same distribution test on the experimental group and the control group according to one or more embodiments of the present application is shown.

[0073] As Figure 2 shown, in step S201, a control group for the experimental group is generated. Optionally, the first control group for the experimental group and the second control group for the experimental group can be generated with reference to the above Figure 1 description.

[0074] In step S203, a loss function is constructed and it is determined whether the value of the loss function is less than a loss threshold. When the value of the loss function is less than the loss threshold, step S205 is entered; otherwise, step S207 is entered. In one embodiment, after the first control group of the experimental group is generated, a first loss function can be constructed based on the mean distance (e.g., the mean Euclidean distance, the mean absolute value distance, etc.) between each first user sample in the set of first user samples and each second user sample in the first control group, and it is determined whether the value of the first loss function is less than the loss threshold. In one embodiment, after the second control group of the experimental group is generated, a second loss function can be constructed based on the mean distance (e.g., the mean Euclidean distance, the mean absolute value distance, etc.) between each first user sample in the set of first user samples and each second user sample in the second control group, and it is determined whether the value of the second loss function is less than the loss threshold. Exemplarily, the first loss function and the second loss function can be constructed with reference to the above formula (1). By constructing the loss function to perform a same-distribution test on the experimental group and the control group in terms of each user feature, the same-distribution test efficiency can be improved compared to the hypothesis test, saving computing resources.

[0075] In step S205, it is determined that the experimental group and the control group satisfy the same distribution.

[0076] In step S207, it is determined that the experimental group and the control group do not satisfy the same distribution, and multiple users within the candidate control group are updated.

[0077] Figure 3 A block diagram of a system for determining experimental data according to one or more embodiments of the present application is shown.

[0078] As Figure 3 shown, the system 300 for determining experimental data includes a memory 310, a processor 320, and a computer program 330 stored on the memory 310 and executable on the processor 320. The processor 320 runs the computer program 330 to implement a method for determining experimental data according to one aspect of the present application.

[0079] The present application can also be implemented as a computer-readable storage medium, the computer storage medium including instructions that, when running, execute a method for determining experimental data according to one aspect of the present application. Additionally, the present application can also be implemented as a computer-readable storage medium, the computer storage medium including instructions that, when running, execute a method for determining experimental data according to one aspect of the present application.

[0080] In applicable cases, various embodiments provided by this application can be implemented using hardware, software, or a combination of hardware and software. Moreover, in applicable cases, without departing from the scope of this application, various hardware components and / or software components described herein can be combined into composite components including software, hardware, and / or both. In applicable cases, without departing from the scope of this application, various hardware components and / or software components described herein can be divided into sub-components including software, hardware, or both. Additionally, in applicable cases, it is contemplated that software components can be implemented as hardware components, and vice versa.

[0081] Software according to this application (such as program code and / or data) can be stored on one or more computer storage media. It is also contemplated that one or more general-purpose or special-purpose computers and / or systems, either networked and / or otherwise, can be used to implement the software identified herein. In applicable cases, the order of the various steps described herein can be changed, combined into composite steps, and / or divided into sub-steps to provide the features described herein.

[0082] The embodiments and examples presented herein are provided to best illustrate the embodiments in accordance with this application and its specific applications, and thereby enable those skilled in the art to implement and use this application. However, those skilled in the art will know that the above description and examples are provided for ease of illustration and example only. The presented description is not intended to cover all aspects of this application or to limit this application to the precise form disclosed.

Claims

1. A method for determining experimental data, characterized in that The method comprises the following steps: Acquire a set of first user samples based on user characteristics of a plurality of first users in the experimental group, and acquire a set of second user samples based on user characteristics of a plurality of second users in the candidate control group; Constructing a multidimensional space data structure based on the set of the second user samples; Determining in the multidimensional space data structure the second user sample that is most adjacent to each first user sample to generate a first control group of the experimental group; as well as A plurality of second user samples associated with each of the first user samples are determined in the multidimensional space data structure, and a second control group of the experimental group is determined based on distance information and sample mapping information between each of the first user samples and the plurality of second user samples.

2. The method according to claim 1, wherein obtaining a set of first user samples based on user characteristics of a plurality of first users in the experimental group comprises: determining abnormal features among the user features of the plurality of first users and processing the abnormal features to obtain preprocessed user features of the plurality of first users; converting the preprocessed user features of the plurality of first users into first numerical features; and The first numerical feature is standardized to obtain a set of the first user samples.

3. The method according to claim 1, wherein obtaining a set of second user samples based on user characteristics of a plurality of second users in the candidate control group comprises: determining abnormal features among the user features of the plurality of second users and processing the abnormal features to obtain preprocessed user features of the plurality of second users; converting the preprocessed user features of the plurality of second users into second numerical features; and The second numerical feature is normalized to obtain a set of the second user samples.

4. The method according to claim 1, wherein the user characteristics of the first user and the user characteristics of the second user each include one or more of the following: user natural attributes, user transaction attributes, and user active attributes.

5. The method according to claim 1, wherein constructing a multidimensional space data structure based on the set of the second user samples comprises: Taking each second user sample in the set of the second user samples as a data point in the multidimensional space to obtain a data set; Determining a segmentation plane of multiple dimensions based on the data points and dividing the data set based on the segmentation plane dividing the multi-dimensional space; as well as The multidimensional space data structure is constructed based on the partition result of the data set. The method according to claim 1 , wherein the multidimensional space data structure is a multidimensional space partitioning tree structure.

7. The method according to claim 1, wherein determining the second user sample closest to each first user sample in the multidimensional space data structure to generate the first control group of the experimental group comprises: For each of the first user samples, searching in parallel in the multidimensional space data structure for a second user sample that is closest to each of the first user samples; as well as A first control group of the experimental group is generated based on a second user sample that is most adjacent to each of the first user samples.

8. The method according to claim 7, wherein searching in parallel in the multidimensional space data structure for the second user samples that are closest to each of the first user samples comprises: Determine a search path for each of the first user samples and search in parallel in the multidimensional space data structure for a candidate second user sample that is closest to each of the first user samples according to the search path; Determine a distance between each of the first user samples and a candidate second user sample that is most adjacent to each of the first user samples; as well as The search path is selectively updated based on the distance and second user samples that are closest to each of the first user samples are searched in parallel in the multi-dimensional space data structure according to the updated search path.

9. The method according to claim 1, wherein determining a plurality of second user samples associated with each of the first user samples in the multidimensional space data structure comprises: For each of the first user samples, searching in parallel in the multidimensional space data structure for a predetermined number of second user samples adjacent to each of the first user samples; as well as A predetermined number of searched second user samples adjacent to each first user sample are determined as a plurality of second user samples associated with each first user sample.

10. The method according to claim 1, wherein the distance information and sample mapping information between each first user sample and the plurality of second user samples are determined by: Determining the distance between each of the first user samples and each of the second user samples in the plurality of second user samples in the multidimensional space data structure and determining a distance matrix generated based on the distance as the distance information; and A mapping relationship between a first user sample and a second user sample corresponding to each element in the distance matrix is ​​determined, and a sample mapping matrix generated based on the mapping relationship is determined as the sample mapping information.

11. The method according to claim 10, wherein determining the second control group of the experimental group based on the distance information and sample mapping information between each of the first user samples and the plurality of second user samples comprises cyclically executing the following steps until the distance matrix is ​​an empty matrix: Determine the element with the smallest element value in the distance matrix; Determine, based on the sample mapping matrix, a first target user sample and a second target user sample corresponding to the element with the smallest element value; Determine a second control group of the experimental group based on the first target user sample and the second target user sample corresponding to the element with the smallest element value; Deleting rows associated with the first target user sample in the distance matrix and the sample mapping matrix respectively; and The value of the element associated with the second target user sample in the distance matrix is ​​set to infinity.

12. The method according to claim 1, wherein the method further comprises: constructing a first loss function based on the mean of the distances between each first user sample in the set of the first user samples and each second user sample in the first control group; constructing a second loss function based on the mean of the distance between each first user sample in the set of the first user samples and each second user sample in the second control group; as well as Based on the value of the first loss function, an identical distribution test is performed on each first user sample in the set of the first user samples and each second user sample in the first control group, and based on the value of the second loss function, an identical distribution test is performed on each first user sample in the set of the first user samples and each second user sample in the second control group.

13. A system for determining experimental data, characterized in that The system comprises: Memory; a processor coupled to the memory; and A computer program stored on the memory and running on the processor, the execution of the computer program causing the following operations: Acquire a set of first user samples based on user characteristics of a plurality of first users in the experimental group, and acquire a set of second user samples based on user characteristics of a plurality of second users in the candidate control group; Constructing a multidimensional space data structure based on the set of the second user samples; Determining in the multidimensional space data structure the second user sample that is most adjacent to each first user sample to generate a first control group of the experimental group; and A plurality of second user samples associated with each of the first user samples are determined in the multidimensional space data structure, and a second control group of the experimental group is determined based on distance information and sample mapping information between each of the first user samples and the plurality of second user samples.

14. The system according to claim 13, wherein the execution of the computer program results in obtaining a set of first user samples based on user features of a plurality of first users in the experimental group, comprising: determining abnormal features among the user features of the plurality of first users and processing the abnormal features to obtain preprocessed user features of the plurality of first users; converting the preprocessed user features of the plurality of first users into first numerical features; and The first numerical feature is standardized to obtain a set of the first user samples.

15. The system according to claim 13, wherein the execution of the computer program results in obtaining a set of second user samples based on user features of a plurality of second users in the candidate control group, comprising: determining abnormal features among the user features of the plurality of second users and processing the abnormal features to obtain preprocessed user features of the plurality of second users; converting the preprocessed user features of the plurality of second users into second numerical features; and The second numerical feature is normalized to obtain a set of the second user samples.

16. The system according to claim 13, wherein the user characteristics of the first user and the user characteristics of the second user each include one or more of the following: user natural attributes, user transaction attributes, and user active attributes.

17. The system according to claim 13, wherein the execution of the computer program results in constructing a multidimensional space data structure based on the set of the second user samples, comprising: Taking each second user sample in the set of the second user samples as a data point in the multidimensional space to obtain a data set; Determining a segmentation plane of multiple dimensions based on the data points and dividing the data set based on the segmentation plane dividing the multi-dimensional space; as well as The multidimensional space data structure is constructed based on the partition result of the data set.

18. The system according to claim 13, wherein the multidimensional space data structure is a multidimensional space partitioning tree structure.

19. The system according to claim 13, wherein the execution of the computer program causes determining the second user sample closest to each first user sample in the multidimensional space data structure to generate the first control group of the experimental group comprises: For each of the first user samples, searching in parallel in the multidimensional space data structure for a second user sample that is closest to each of the first user samples; as well as A first control group of the experimental group is generated based on a second user sample that is most adjacent to each of the first user samples.

20. The system of claim 19, wherein the execution of the computer program causes the parallel searching of the multi-dimensional space data structure for the second user samples that are closest to each of the first user samples to comprise: Determine a search path for each of the first user samples and search in parallel in the multidimensional space data structure for a candidate second user sample that is closest to each of the first user samples according to the search path; Determine a distance between each of the first user samples and a candidate second user sample that is most adjacent to each of the first user samples; as well as The search path is selectively updated based on the distance and second user samples that are closest to each of the first user samples are searched in parallel in the multi-dimensional space data structure according to the updated search path.

21. The system of claim 13, wherein the execution of the computer program causes determining in the multidimensional space data structure a plurality of second user samples associated with each of the first user samples to include: For each of the first user samples, searching in parallel in the multidimensional space data structure for a predetermined number of second user samples adjacent to each of the first user samples; as well as A predetermined number of searched second user samples adjacent to each first user sample are determined as a plurality of second user samples associated with each first user sample.

22. The system of claim 13, wherein the execution of the computer program causes the distance information and sample mapping information between each first user sample and the plurality of second user samples to be determined by: Determining the distance between each of the first user samples and each of the second user samples in the plurality of second user samples in the multidimensional space data structure and determining a distance matrix generated based on the distance as the distance information; and A mapping relationship between a first user sample and a second user sample corresponding to each element in the distance matrix is ​​determined, and a sample mapping matrix generated based on the mapping relationship is determined as the sample mapping information.

23. The system according to claim 22, wherein the execution of the computer program causes the second control group of the experimental group to be determined based on the distance information and sample mapping information between each of the first user samples and the plurality of second user samples, comprising cyclically executing the following steps until the distance matrix is ​​an empty matrix: Determine the element with the smallest element value in the distance matrix; Determine, based on the sample mapping matrix, a first target user sample and a second target user sample corresponding to the element with the smallest element value; Determine a second control group of the experimental group based on the first target user sample and the second target user sample corresponding to the element with the smallest element value; Deleting rows associated with the first target user sample in the distance matrix and the sample mapping matrix respectively; and The value of the element associated with the second target user sample in the distance matrix is ​​set to infinity.

24. The system of claim 13, wherein execution of the computer program further results in: constructing a first loss function based on the mean of the distances between each first user sample in the set of the first user samples and each second user sample in the first control group; constructing a second loss function based on the mean distance between each first user sample in the set of the first user samples and each second user sample in the second control group; and Based on the value of the first loss function, an identical distribution test is performed on each first user sample in the set of the first user samples and each second user sample in the first control group, and based on the value of the second loss function, an identical distribution test is performed on each first user sample in the set of the first user samples and each second user sample in the second control group.

25. A computer-readable storage medium, characterized in that: The computer storage medium comprises instructions which, when executed, perform the method for determining experimental data according to any one of claims 1-12.

26. A computer program product, characterized in that The computer program product comprises instructions, which, when executed by a processor, implement the method for determining experimental data according to any one of claims 1-12.