User sample data processing method and device, electronic equipment and medium

By sorting and similarity screening the characteristics importance of sample users on e-commerce platform, the sample user group was processed using a random forest algorithm model, which solved the problem of inaccurate experimental results caused by random sampling, and achieved a more reliable control experiment.

CN120372310APending Publication Date: 2025-07-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510459001.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, randomly sampled sample users have not been reliable enough in the control experiment of e-commerce platform due to differences in other factors other than independent variables.

Method used

By sorting the feature importance of sample users of e-commerce platforms using a pre-trained random forest algorithm model, a sample user group with feature similarity higher than the threshold was selected, and a controlled experiment was conducted.

Benefits of technology

The results reliability of the control experiments are improved and the impact of other factors besides independent variables on the experimental dependent variables is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372310A_ABST
    Figure CN120372310A_ABST
Patent Text Reader

Abstract

The invention provides a user sample data processing method and device, electronic equipment and a medium, and relates to the technical field of data processing, in particular to the technical field of data screening and data clustering. According to the implementation scheme, an initial sample user group of the e-commerce platform is obtained, and each sample user in the initial sample user group comprises multiple to-be-processed user features associated with independent variables of a control experiment of the e-commerce platform and result features associated with dependent variables of the control experiment; processing the initial sample user group by using a pre-trained random forest algorithm model to obtain a plurality of feature importance sorting results; screening out a plurality of target user features from the plurality of to-be-processed user features according to a plurality of feature importance sorting results; based on the multiple target user features, multiple sample users with the feature similarity higher than a preset threshold value are screened out from the initial sample user group and added into a target sample user group; and performing a control experiment based on the target sample user group.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technologies, and more particularly to the fields of data screening and data clustering technologies. Specifically, the present disclosure relates to a method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for processing user sample data. Background Art

[0002] Currently, e-commerce platforms extract sample users through random sampling, and use the sampled sample users for control experiments related to the e-commerce platforms.

[0003] The methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, no method described in this section should be considered prior art merely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0004] The present disclosure provides a method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for processing user sample data.

[0005] According to one aspect of the present disclosure, there is provided a method for processing user sample data, including: obtaining an initial sample user group of an e-commerce platform, where each sample user in the initial sample user group includes a plurality of to-be-processed user features associated with an independent variable of a control experiment of the e-commerce platform and a result feature associated with a dependent variable of the control experiment, and each to-be-processed user feature in the plurality of to-be-processed user features is associated with a behavior feature or a user portrait feature of the corresponding sample user on the e-commerce platform; processing the initial sample user group using a pre-trained random forest algorithm model to obtain a plurality of feature importance ranking results, where each feature importance ranking result in the plurality of feature importance ranking results indicates the influence degree of each to-be-processed user feature on the control experiment result; screening out a plurality of target user features from the plurality of to-be-processed user features according to the plurality of feature importance ranking results; screening out a plurality of sample users with a feature similarity higher than a preset threshold from the initial sample user group based on the plurality of target user features and adding them to a target sample user group; and performing the control experiment based on the target sample user group.

[0006] According to another aspect of the present disclosure, a device for processing user sample data is provided, comprising: an acquisition module configured to acquire an initial sample user group of an e-commerce platform, wherein each sample user in the initial sample user group includes a plurality of to-be-processed user features associated with an independent variable of a control experiment of the e-commerce platform and a result feature associated with a dependent variable of the control experiment, and each of the plurality of to-be-processed user features is associated with a behavioral feature or a user portrait feature of the corresponding sample user on the e-commerce platform; a processing module configured to process the initial sample user group using a pre-trained random forest algorithm model to obtain a plurality of feature importance ranking results, wherein each of the plurality of feature importance ranking results indicates the degree of influence of each to-be-processed user feature on the control experiment result; a first screening module configured to screen out a plurality of target user features from the plurality of to-be-processed user features based on the plurality of feature importance ranking results; a second screening module configured to screen out a plurality of sample users whose feature similarity is higher than a preset threshold from the initial sample user group based on the plurality of target user features and add them to the target sample user group; and an experiment module configured to conduct the control experiment based on the target sample user group.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above method.

[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the above method when executed by a processor.

[0010] According to one or more embodiments of the present disclosure, a method for processing user sample data is provided, in which an initial sample user group is processed by a random forest algorithm to rank the feature importance of user behavior characteristics and user portrait characteristics on an e-commerce platform, thereby screening out sample users with high similarity as an experimental group and a control group of a control experiment of the e-commerce platform, thereby being able to exclude as much as possible the influence of irrelevant factors other than the experimental independent variable on the experimental dependent variable in the control experiment, so that the results of the control experiment based on the screened sample users are more accurate.

[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood from the following description. Description of the Drawings

[0012] The drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0013] Figure 1 is a schematic diagram illustrating an example system in which various methods described herein can be implemented according to an exemplary embodiment;

[0014] Figure 2 shows a flowchart of a method for processing user sample data according to an embodiment of the present disclosure;

[0015] Figure 3 shows a partial flowchart of a method for processing another user sample data according to an embodiment of the present disclosure;

[0016] Figure 4 shows a partial flowchart of a method for processing another user sample data according to an embodiment of the present disclosure;

[0017] Figure 5 shows a partial flowchart of a method for processing another user sample data according to an embodiment of the present disclosure;

[0018] Figure 6 shows a partial flowchart of a method for processing another user sample data according to an embodiment of the present disclosure;

[0019] Figure 7 shows a block diagram of the structure of a processing device for user sample data according to an embodiment of the present disclosure; and

[0020] Figure 8 shows a block diagram of the structure of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. Detailed Embodiments

[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0022] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the description of the context, they may also refer to different instances.

[0023] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.

[0024] In the related art, a method of randomly sampling sample users of an e-commerce platform is proposed to use the sampled sample users for control experiments related to the e-commerce platform (for example, AB tests). However, there are usually large differences among the sample users obtained by random sampling in other factors except the independent variable of the control experiment. Therefore, it is impossible to determine whether these factors will affect the experimental results at this time, resulting in the results of the control experiment being unreliable.

[0025] To solve the above problems, the present disclosure provides a method for processing user sample data. The initial sample user group is processed by a random forest algorithm to rank the importance of the behavioral characteristics and user portrait characteristics of users on the e-commerce platform, so as to screen out sample users with higher similarity as the experimental group and the control group of the control experiment on the e-commerce platform respectively. Thus, it is possible to effectively control the influence of other irrelevant factors except the experimental independent variable in the control experiment on the experimental dependent variable, making the results of the control experiment based on the screened sample users more reliable.

[0026] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0027] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to an embodiment of the present disclosure. Referring Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more applications.

[0028] In an embodiment of the present disclosure, the server 120 may run one or more services or software applications that enable a processing method for user sample data to be performed.

[0029] In some embodiments, the server 120 may also provide other services or software applications that may include a non-virtual environment and a virtual environment. In some embodiments, these services may be provided as web-based services or cloud services, such as provided to users of client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.

[0030] In Figure 1 the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may be different from the system 100. Thus, Figure 1 is an example of a system for implementing the data quantization method of the optimizer described herein and is not intended to be limiting.

[0031] Users may use the client devices 101, 102, 103, 104, 105, and / or 106 to perform the processing method for user sample data. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.

[0032] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.

[0033] Network 110 may be any type of network known to those skilled in the art, which may support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, virtual network, virtual private network (VPN), intranet, extranet, public switched telephone network (PSTN), infrared network, wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0034] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.

[0035] The computing units in server 120 can run one or more operating systems including any of the above-mentioned operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0036] In some embodiments, server 120 can include one or more applications to analyze and combine data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0037] In some embodiments, server 120 can be a server of a distributed system, or a server combined with a blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.

[0038] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store relevant information and other information of sample users. For example, one or more of databases 130 can be used to store such as behavioral characteristics and user portrait characteristics of sample users. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120, or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be, for example, relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.

[0039] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.

[0040] Figure 1System 100 can be configured and operated in various ways to enable the application of the various methods and devices described in this disclosure.

[0041] Figure 2 The flowchart of the processing method of user sample data according to an embodiment of the present disclosure is shown.

[0042] As Figure 2 shown, the processing method 200 of user sample data includes:

[0043] Step 210, obtain an initial sample user group of the e-commerce platform, where each sample user in the initial sample user group includes a plurality of to-be-processed user features associated with the independent variable of the controlled experiment on the e-commerce platform and a result feature associated with the dependent variable of the controlled experiment, and each to-be-processed user feature in the plurality of to-be-processed user features is associated with the behavior feature or user portrait feature of the corresponding sample user on the e-commerce platform;

[0044] Step 220, process the initial sample user group using a pre-trained random forest algorithm model to obtain a plurality of feature importance ranking results, where each feature importance ranking result in the plurality of feature importance ranking results indicates the influence degree of each to-be-processed user feature on the controlled experiment result;

[0045] Step 230, screen out a plurality of target user features from the plurality of to-be-processed user features according to the plurality of feature importance ranking results;

[0046] Step 240, based on the plurality of target user features, screen out a plurality of sample users with a feature similarity higher than a preset threshold from the initial sample user group and add them to the target sample user group; and

[0047] Step 250, conduct a controlled experiment based on the target sample user group.

[0048] Thus, by processing the initial sample user group through the random forest algorithm to rank the feature importance of the behavior features and user portrait features of users on the e-commerce platform, so as to screen out sample users with relatively high similarity as the experimental group and the control group of the controlled experiment on the e-commerce platform respectively. Thus, it is possible to effectively control the influence of other irrelevant factors except the experimental independent variable on the experimental dependent variable in the controlled experiment, making the result of the controlled experiment based on the screened sample users more reliable.

[0049] In step 210, the initial sample user group can be randomly sampled from the users of the e-commerce platform.

[0050] In step 210, the control experiment for the e-commerce platform can be, for example, a test experiment conducted in aspects such as experience optimization, advertisement optimization, and conversion rate optimization of the e-commerce platform. In one example, the control experiment is used to test which version among different versions of button colors, layout methods, or animation effects can attract users to perform consumption behaviors more effectively. In another example, the control experiment is used to test which version of the advertisement strategy among different versions of advertisement strategies can attract users and improve the conversion rate more effectively. In yet another example, the control experiment is used to test the response differences of user groups with different consumption levels to different marketing strategies.

[0051] In step 210, according to some embodiments, the behavioral characteristics of the sample users include the consumption behavioral characteristics and non-consumption behavioral characteristics of the corresponding sample users with respect to the e-commerce platform. Exemplarily, the consumption behavioral characteristics can be, for example, the number of orders placed by the user, the amount of orders placed, and the types of goods ordered, etc., to characterize the user's consumption level and product preferences. Exemplarily, the non-consumption behavioral characteristics can be, for example, the daily activity level of the user, the types of goods browsed, and the types of goods clicked to view, etc., to characterize the user's product preferences and the likelihood of placing an order (for example, users with a high activity level are more likely to generate purchase behaviors).

[0052] Thus, it is possible to take into account as many relevant factors as possible that may be involved in the control experiment of the e-commerce platform, so as to be generally applicable to more types of control experiments related to the e-commerce platform.

[0053] In step 210, the user portrait characteristics of the sample users can be, for example, characteristics such as the age, gender, region, occupation, and consumption level of the users, which have a relatively strong relevance to the control experiment of the e-commerce platform.

[0054] In one example, based on the category information of the goods provided by the e-commerce platform, feature engineering can be performed on goods of different categories to extract the associated content of multiple to-be-processed user characteristics of the sample users, so that the finally generated multiple to-be-processed user characteristics can better characterize the tendencies of the sample users.

[0055] Exemplarily, transaction records, user evaluations, product description information, and external resources of the goods associated with the sample users (such as purchased or browsed) can be obtained from the e-commerce platform as the original data for extracting the to-be-processed user characteristics.

[0056] Specifically, the transaction records can be, for example, the purchase frequency, consumption amount, and purchase time of a sample user for a specific commodity; the user evaluations can be, for example, the ratings and evaluation contents of a sample user for a specific commodity, so as to be able to characterize the views of the user on the specific commodity; the commodity description information can be, for example, the brand, model, specifications, color, and material of a specific commodity, etc.; and, the external resources can be, for example, unstructured data sources such as social media comment contents, news reports, and industry reports related to a specific commodity, etc.

[0057] Exemplarily, content associated with the independent variable of the control experiment can be extracted from the above-mentioned original data to generate multiple to-be-processed user characteristics of the sample user.

[0058] Specifically, the to-be-processed user characteristics can include, for example: (1) commodity attribute information directly scraped from the commodity details page, such as price, inventory status, and shipping location, etc.; (2) behavioral characteristics generated based on the interaction behaviors of the sample user, such as the number of views, and actions of adding to the shopping cart or favorites, etc.; (3) context characteristics comprehensively considering time and location factors, such as performance during holiday promotions and demand changes under specific geographical locations, etc.; and (4) other implicit characteristics, such as the sentiment tendency in user comments determined through natural language understanding, and potential correlations discovered through collaborative algorithms (which commodities are often purchased together), etc.

[0059] In step 210, the dependent variable of the control experiment can be, for example, commodity sales volume and total commodity transaction volume (GMV, Gross Merchandise Volume), etc. Correspondingly, the result characteristics can be, for example, the purchase quantity and transaction amount of the sample user for different categories of commodities.

[0060] In step 220, not every to-be-processed user characteristic among the multiple to-be-processed user characteristics is effective, and some to-be-processed user characteristics may be interference items. The random forest algorithm can rank the feature importance according to the influence degree of the multiple to-be-processed user characteristics on the dependent variable of the control experiment, so as to exclude redundant to-be-processed user characteristics with relatively low contribution or high correlation with other to-be-processed user characteristics.

[0061] Figure 3 Shows a partial flowchart of another method for processing user sample data according to an embodiment of the present disclosure.

[0062] According to some embodiments, as Figure 3 shown, step 220 includes:

[0063] Step 310, obtain multiple initial random forest algorithm models; and

[0064] Step 320: Use multiple reference training data sets to train multiple initial random forest algorithm models respectively to obtain corresponding multiple pre-trained random forest algorithm models;

[0065] Step 330: Use multiple pre-trained random forest algorithm models to process the initial sample user group respectively to obtain multiple feature importance ranking results.

[0066] By using multiple random forest algorithm models trained with different reference training data sets to process the initial sample user group, it is possible to determine the ranking stability degree of each user feature to be processed based on multiple feature importance ranking results obtained from multiple models, so as to more effectively screen out the user features to be processed with a relatively high degree of association with the independent variable of the control experiment, and improve the accuracy of the determined target sample user group.

[0067] In step 310, an initial random forest algorithm model can be constructed using, for example, a decision tree model, a deep learning model, and a combination of the two.

[0068] It is understandable that the specific implementation process of the random forest algorithm belongs to the prior art and will not be elaborated here.

[0069] Figure 4 The figure shows a partial flowchart of another method for processing user sample data according to an embodiment of the present disclosure.

[0070] According to some embodiments, as Figure 4 shown, step 220 includes:

[0071] Step 410: Obtain a pre-trained random forest algorithm model;

[0072] Step 420: Randomly set initial parameters for the pre-trained random forest algorithm model;

[0073] Step 430: Use the pre-trained random forest algorithm model with randomly set initial parameters to process the initial sample user group to obtain corresponding feature importance ranking results;

[0074] Step 440: In response to the number of obtained feature importance ranking results not reaching the target number, return to the step of randomly setting initial parameters for the pre-trained random forest algorithm model, where the target number is an integer greater than 1; and

[0075] Step 450: In response to the number of obtained feature importance ranking results reaching the target number, obtain multiple feature importance ranking results.

[0076] By setting random initial parameters for a pre-trained random forest algorithm model and using the model to process the initial sample user group after each parameter setting, multiple feature importance ranking results can be obtained, thereby enabling the determination of the ranking stability degree of each user feature to be processed, so as to more effectively screen out the user features to be processed with a relatively high degree of association with the independent variables of the control experiment and improve the accuracy of the determined target sample user group.

[0077] In step 440, the target quantity can be determined according to the specific type of the control experiment and the data volume of the initial sample user group.

[0078] According to some embodiments, step 230 includes: for each user feature to be processed among the multiple user features to be processed, in response to the ranking of the user feature to be processed falling within the target range in each feature importance ranking result among the multiple feature importance ranking results, determining the user feature to be processed as a target user feature.

[0079] Specifically, if the ranking of a certain user feature to be processed fluctuates greatly among the multiple feature importance ranking results, it indicates that the user feature to be processed is more likely to be affected by interfering factors or itself is an interfering factor, and it can be excluded, so that the target sample user group for the control experiment can be more accurately determined based on the screened target user features, and the effectiveness of the results of the control experiment can be improved.

[0080] Figure 5 Shows a partial flowchart of another method for processing user sample data according to an embodiment of the present disclosure.

[0081] According to some embodiments, as Figure 5 shown, step 240 includes:

[0082] Step 510, processing the initial sample user group using the principal component analysis method to transform the multiple target user features into multiple principal components and multiple weights corresponding to the multiple principal components;

[0083] Step 520, processing the initial sample user group using a logistic regression model based on the multiple principal components and the multiple weights to obtain the propensity score of each sample user; and

[0084] Step 530, screening out multiple sample users with a feature similarity greater than the preset threshold from the initial sample user group based on the propensity score of each sample user and adding them to the target sample user group.

[0085] Therefore, by calculating the propensity scores of sample users to screen for sample users with relatively high feature similarity to join the target sample user group for a controlled experiment, the comparison of complex multi-dimensional user features is reduced to a simple and intuitive score comparison, thereby reducing the matching difficulty and saving computing resources.

[0086] In step 510, multiple target user features can be linearly transformed to convert a relatively large number of multiple target user features into a relatively small number of comprehensive variables (principal components), and weights can be assigned based on the contribution rates of the respective principal components to comprehensively consider the influence degrees of the respective target user features on the dependent variable of the controlled experiment.

[0087] In step 520, the specific process of calculating the propensity score using a logistic regression model belongs to the prior art and will not be elaborated here.

[0088] In step 530, the preset thresholds can be, for example, 50%, 70%, and 90%.

[0089] According to some embodiments, in addition to steps 210 to 250, method 200 further includes: excluding sample users whose propensity scores fall outside the target range from the initial sample user group to update the initial sample user group.

[0090] Based on this, some sample users with relatively obvious feature differences in the initial sample user group can be excluded, making the overall similarity of the sample users in the initial sample user group higher, thereby further reducing the subsequent processing difficulty and at the same time avoiding the influence of extreme sample users on the results of the controlled experiment and improving the reliability of the results of the controlled experiment.

[0091] In one example, sample users whose feature scores fall outside the target range can be directly deleted from the initial sample user group. In another example, sample users whose feature scores fall outside the target range can be replaced with sample users whose feature scores are closest to them and fall within the target range.

[0092] According to some embodiments, step 240 includes: for each sample user in the initial sample user group, determining whether the sample user can join the target sample user group based on the first data associated with multiple target user features of the sample user in the first time period.

[0093] Therefore, by obtaining the first data associated with multiple target user features of the sample user in the same time period as the judgment basis, the influence of a specific time period on the sample user can be avoided (for example, the differences in the consumption behaviors of users during shopping festivals and non-shopping festival time periods, etc.), making the selected target sample user group more effective.

[0094] Exemplarily, the first time period can be, for example, one month to better reflect the overall consumption behavior of the user over a period of time.

[0095] However, in some specific cases where the sample users, for example, belong to inactive users with a low consumption frequency, the user may not generate any behavior within the first time period, and it is necessary to further expand the time range to obtain the characteristic data of a sufficient number of sample users.

[0096] Figure 6 FIG. shows a partial flowchart of another method for processing user sample data according to an embodiment of the present disclosure.

[0097] To solve the above problems, according to some embodiments, as Figure 6 described, in addition to steps 210 to 250, method 200 further includes:

[0098] For each sample user in the initial sample user group,

[0099] Step 610, in response to determining that the sample user has not generated any behavior on the e-commerce platform within the first time period, obtain second data associated with multiple target user characteristics of the sample user within a second time period, where the duration of the second time period is longer than that of the first time period; and

[0100] Step 620, based on the second data, determine whether the sample user can be added to the target sample user group.

[0101] Thereby, it is possible to further improve the completeness of the second data associated with the sample user and multiple target user characteristics, making the target sample user group selected based on the second data more effective.

[0102] According to some embodiments, method 200 further includes: performing a feature similarity test on multiple sample users in the target sample user group to determine whether all pairs of the multiple sample users in the target sample user group satisfy that the feature similarity is higher than a preset threshold.

[0103] Through the above posterior operation, it is possible to better ensure the effectiveness of the selected target sample user group, thereby ensuring the reliability of the experimental results of the control experiment based on the target sample user group.

[0104] In one example, the target sample user group can be randomly divided into a control group and an experimental group for a goodness-of-fit test. Specifically, by a visualization method, kernel density curves of the feature scores of the sample users in the control group and the experimental group are respectively plotted to observe the distribution of the feature scores of the two groups of sample users. For example, if the shapes of the kernel density plots of the control group and the experimental group are similar, it indicates that the distribution of the feature similarities of the two groups of sample users is similar.

[0105] In another example, it is also possible to verify whether multiple sample users in the target sample user group all meet the preset conditions based on the distribution of characteristic scores of the experimental group and the control group observed through histograms, scatter plots, etc.

[0106] In yet another example, it is also possible to verify whether multiple sample users in the target sample user group all meet the preset conditions based on statistical methods such as standardized difference statistics.

[0107] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved all complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0108] Figure 7 The structural block diagram of a processing device for user sample data according to an embodiment of the present disclosure is shown.

[0109] According to another aspect of the present disclosure, a processing device for user sample data is provided. As Figure 7 shown, the data quantization device 700 of the optimizer includes: an acquisition module 710 configured to acquire an initial sample user group of an e-commerce platform, where each sample user in the initial sample user group includes a plurality of to-be-processed user characteristics associated with an independent variable of a control experiment of the e-commerce platform and a result characteristic associated with a dependent variable of the control experiment, and each to-be-processed user characteristic in the plurality of to-be-processed user characteristics is associated with a behavior characteristic or a user portrait characteristic of the corresponding sample user on the e-commerce platform; a processing module 720 configured to process the initial sample user group using a pre-trained random forest algorithm model to obtain a plurality of feature importance ranking results, where each feature importance ranking result in the plurality of feature importance ranking results indicates the influence degree of each to-be-processed user characteristic on the control experiment result; a first screening module 730 configured to screen out a plurality of target user characteristics from the plurality of to-be-processed user characteristics according to the plurality of feature importance ranking results; a second screening module 740 configured to screen out a plurality of sample users with a feature similarity higher than a preset threshold from the initial sample user group based on the plurality of target user characteristics and add them to the target sample user group; and an experiment module 750 configured to perform the control experiment based on the target sample user group.

[0110] According to another aspect of the present disclosure, an electronic device is further provided, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the foregoing method.

[0111] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the foregoing method.

[0112] According to another aspect of the present disclosure, there is also provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the foregoing method.

[0113] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0114] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 808 can be any type of device capable of inputting information into the electronic device 800. The input unit 806 can receive input digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 807 can be any type of device capable of presenting information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include but are not limited to a magnetic disk, an optical disk. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth TM device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0115] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the GPU-based matrix calculation method. For example, in some embodiments, the GPU-based matrix calculation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the GPU-based matrix calculation method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the GPU-based matrix calculation method in any other suitable way (e.g., by means of firmware).

[0116] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0117] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0118] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0119] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0120] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0121] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0122] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0123] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after the present disclosure.

Claims

1. A method for processing user sample data, comprising: Obtaining an initial sample user group of an e-commerce platform, wherein each sample user in the initial sample user group includes a plurality of to-be-processed user features associated with an independent variable of a controlled experiment on the e-commerce platform and a result feature associated with a dependent variable of the controlled experiment, and each to-be-processed user feature in the plurality of to-be-processed user features is associated with a behavior feature or a user portrait feature of the corresponding sample user on the e-commerce platform; Processing the initial sample user group using a pre-trained random forest algorithm model to obtain a plurality of feature importance ranking results, wherein each feature importance ranking result in the plurality of feature importance ranking results indicates the influence degree of each to-be-processed user feature on the result of the controlled experiment; Screening out a plurality of target user features from the plurality of to-be-processed user features according to the plurality of feature importance ranking results; Based on the plurality of target user features, screening out a plurality of sample users with a feature similarity higher than a preset threshold from the initial sample user group and adding them to a target sample user group; and Performing the controlled experiment based on the target sample user group.

2. The method according to claim 1, wherein, The processing the initial sample user group using a pre-trained random forest algorithm model to obtain a plurality of feature importance ranking results includes: Obtaining a plurality of initial random forest algorithm models; Training the plurality of initial random forest algorithm models respectively using a plurality of reference training data sets to obtain the corresponding plurality of pre-trained random forest algorithm models; and Processing the initial sample user group respectively using the plurality of pre-trained random forest algorithm models to obtain the plurality of feature importance ranking results.

3. The method according to claim 1, wherein The processing the initial sample user group using a pre-trained random forest algorithm model to obtain a plurality of feature importance ranking results includes: Obtaining the pre-trained random forest algorithm model; Randomly setting initial parameters for the pre-trained random forest algorithm model; Processing the initial sample user group using the pre-trained random forest algorithm model with randomly set initial parameters to obtain a corresponding feature importance ranking result; In response to the number of obtained feature importance ranking results not reaching a target number, returning to execute the step of randomly setting initial parameters for the pre-trained random forest algorithm model, wherein the target number is an integer greater than 1; and In response to the number of obtained feature importance ranking results reaching the target number, obtaining the plurality of feature importance ranking results.

4. The method according to any one of claims 1-3, wherein, The screening out a plurality of target user features from the plurality of to-be-processed user features according to the plurality of feature importance ranking results includes: For each to-be-processed user feature in the plurality of to-be-processed user features, in response to the ranking of the to-be-processed user feature in each feature importance ranking result in the plurality of feature importance ranking results falling within a target range, determining the to-be-processed user feature as a target user feature.

5. The method according to any one of claims 1-4, wherein, The screening out a plurality of sample users with a feature similarity higher than a preset threshold from the initial sample user group and adding them to a target sample user group based on the plurality of target user features includes: Process the initial sample user group using the principal component analysis method to transform the multiple target user features into multiple principal components and multiple weights corresponding to the multiple principal components; Based on the multiple principal components and the multiple weights, process the initial sample user group using a logistic regression model to obtain the propensity score of each sample user; and Based on the propensity score of each sample user, screen out multiple sample users with feature similarity greater than the preset threshold from the initial sample user group and add them to the target sample user group.

6. The method according to claim 5, further comprising: Exclude sample users whose propensity scores fall outside the target range from the initial sample user group to update the initial sample user group.

7. The method according to any one of claims 1-6, wherein, The screening out multiple sample users with feature similarity greater than a preset threshold from the initial sample user group based on the multiple target user features and adding them to the target sample user group includes: For each sample user in the initial sample user group, determine whether the sample user can be added to the target sample user group based on the first data associated with the multiple target user features by the sample user within the first time period.

8. The method according to claim 7, further comprising: For each sample user in the initial sample user group, In response to determining that the sample user has not generated behavior on the e-commerce platform within the first time period, obtain second data associated with the multiple target user features by the sample user within the second time period, wherein the duration of the second time period is longer than the duration of the first time period; and Based on the second data, determine whether the sample user can be added to the target sample user group.

9. The method according to any one of claims 1-8, further comprising: Conduct a feature similarity test on the multiple sample users in the target sample user group to determine whether all pairs of the multiple sample users in the target sample user group satisfy the feature similarity higher than the preset threshold.

10. A processing device for user sample data, comprising: An acquisition module configured to acquire an initial sample user group of an e-commerce platform, wherein each sample user in the initial sample user group includes multiple to-be-processed user features associated with the independent variable of a controlled experiment on the e-commerce platform and a result feature associated with the dependent variable of the controlled experiment, and each to-be-processed user feature among the multiple to-be-processed user features is associated with the behavior feature or user portrait feature of the corresponding sample user on the e-commerce platform; A processing module configured to process the initial sample user group using a pre-trained random forest algorithm model to obtain multiple feature importance ranking results, wherein each feature importance ranking result among the multiple feature importance ranking results indicates the influence degree of each to-be-processed user feature on the result of the controlled experiment; A first screening module configured to screen out multiple target user features from the multiple to-be-processed user features according to the multiple feature importance ranking results; A second screening module, configured to screen out a plurality of sample users with a feature similarity higher than a preset threshold from the initial sample user group based on the plurality of target user features, and add them to the target sample user group; and An experiment module, configured to conduct the controlled experiment based on the target sample user group.

11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.

13. A computer program product, comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-9.

Citation Information

Cited By

  • Chip batch detection system and method

    CN120807517A