Data processing method and device, electronic equipment, medium and chip

CN116050543BActive Publication Date: 2026-09-29BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310134393.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-09
Publication Date
2026-09-29
Estimated Expiration
2043-02-09

AI Technical Summary

Benefits of technology

[0011]根据本公开的一个或多个实施例,提供了一种数据处理方法,将预训练得到的树结构中的节点作为用来确定用户特征向量的规则,从而实现了规则的自动化生成和对用户的特征编码。进一步基于由规则得到的特征向量对用户群中的用户进行聚类,实现了对用户的自动化分类。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116050543B_ABST
    Figure CN116050543B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method and device, electronic equipment and medium, relates to the technical field of artificial intelligence, in particular to the technical field of deep learning. The implementation scheme is: obtaining a tree structure generated based on pre-training of a first sample data set, wherein the first sample data in the first sample data set has a plurality of attributes; generating a rule set based on the nodes of the tree structure, wherein the nodes of the tree structure represent the value interval corresponding to one of the plurality of attributes, and the rule set is a collection of at least one node; determining a feature vector of second sample data in a second sample data set based on the rule set, wherein each dimension of the feature vector corresponds to each rule in the rule set; and dividing the second sample data set into at least one second sample data subset based on the feature vector of the second sample data in the second sample data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the field of deep learning technology, specifically to a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] Depending on the application scenario, users can be categorized based on their different attributes. This facilitates personalized recommendations and content sharing to users in different categories. At the same time, more accurate user categorization can also improve the effectiveness of content recommendations.

[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0005] This disclosure provides a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product.

[0006] According to one aspect of this disclosure, a data processing method is provided, comprising: obtaining a tree structure generated based on pre-training of a first sample dataset, wherein the first sample data in the first sample dataset has multiple attributes; generating a rule set based on the nodes of the tree structure, wherein the nodes of the tree structure represent the value range corresponding to one of the multiple attributes, and the rule set is a set of at least one node; determining a feature vector of a second sample data in a second sample dataset based on the rule set, wherein each dimension of the feature vector corresponds to a rule in the rule set; and dividing the second sample dataset into at least one subset of second sample data based on the feature vector of the second sample data in the second sample dataset.

[0007] According to another aspect of this disclosure, a data processing apparatus is provided, comprising: an acquisition module configured to acquire a tree structure generated based on pre-training of a first sample dataset, wherein first sample data in the first sample dataset has multiple attributes; a generation module configured to generate a rule set based on nodes of the tree structure, wherein nodes of the tree structure represent value ranges corresponding to one of the multiple attributes, and the rule set is a set of at least one node; a first determination module configured to determine feature vectors of second sample data in a second sample dataset based on the rule set, wherein each dimension of the feature vector corresponds to a rule in the rule set; and a first partitioning module configured to partition the second sample dataset into at least one subset of second sample data based on the feature vectors of the second sample data in the second sample dataset.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods described above.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the above-described method.

[0010] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the above-described method.

[0011] According to one or more embodiments of this disclosure, a data processing method is provided that uses nodes in a pre-trained tree structure as rules to determine user feature vectors, thereby achieving automated rule generation and user feature encoding. Furthermore, based on the feature vectors obtained from the rules, users in the user group are clustered, achieving automated user classification.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0014] Figure 1 This is a schematic diagram illustrating example systems in which the various methods described herein can be implemented according to exemplary embodiments.

[0015] Figure 2 A flowchart of a data processing method according to an embodiment of the present disclosure is shown;

[0016] Figure 3 A schematic diagram of a decision tree according to an embodiment of the present disclosure is shown;

[0017] Figure 4 A flowchart of a portion of a data processing method according to an embodiment of the present disclosure is shown;

[0018] Figure 5 A structural block diagram of a data processing apparatus according to embodiments of the present disclosure is shown; and

[0019] Figure 6 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0021] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0022] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0023] In related technologies, the classification of user groups relies on manual annotation. Manual annotation methods are labor-intensive and easily affected by the subjective factors of the annotators, leading to sample bias and thus poor classification results.

[0024] To address the aforementioned issues, this disclosure provides a data processing method that uses nodes in a pre-trained tree structure as rules to determine user feature vectors, thereby achieving automated rule generation and user feature encoding. Furthermore, based on the feature vectors obtained from the rules, users within the user group are clustered, achieving automated user classification.

[0025] It should be noted that the acquisition, storage, and application of user personal information (such as historical behavior information and geographical location information) involved in the technical solution disclosed herein all comply with the provisions of relevant laws and regulations and do not violate public order and good morals. Furthermore, user personal information has undergone de-identification processing (i.e., anonymization processing) during the acquisition, storage, and application processes.

[0026] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0027] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0028] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of data processing methods.

[0029] In some embodiments, server 120 may also provide other services or software applications that may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.

[0030] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the methods described herein, and is not intended to be limiting.

[0031] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to execute data processing methods. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0032] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0033] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0034] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0035] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0036] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0037] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0038] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0039] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0040] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0041] Figure 2 A flowchart of a data processing method according to an embodiment of the present disclosure is shown.

[0042] like Figure 2 As shown, the data processing method 200 includes: step S201, obtaining a tree structure generated based on pre-training of a first sample dataset, wherein the first sample data in the first sample dataset has multiple attributes;

[0043] Step S202: Based on the nodes of the tree structure, generate a rule set, wherein the nodes of the tree structure represent the value range corresponding to one of the multiple attributes, and the rule set is a set of at least one node;

[0044] Step S203: Based on the rule set, determine the feature vector of the second sample data in the second sample dataset, wherein each dimension of the feature vector corresponds to a rule in the rule set; and

[0045] Step S204: Based on the feature vectors of the second sample data in the second sample dataset, divide the second sample dataset into at least one subset of second sample data.

[0046] Understandably, the first sample dataset serves as the training set to train a tree structure for partitioning the second sample dataset. The second sample dataset can be the target user group to be classified. Data processing method 200 uses the tree structure trained on the first sample dataset to perform feature representation and classification of the target users in the second sample dataset. The tree structure obtained in step S201 can be selected from the appropriate first sample dataset for training based on the specific application scenario. Furthermore, the multiple attributes of the first sample data can be determined based on the application scenario and corresponding user behavior.

[0047] For example, when classifying user groups for the purpose of promoting or recommending a certain type of product or service, the first sample dataset can be a user group consisting of users who have purchased the product or service and users who have not purchased the product or service. Accordingly, in this example, multiple attributes of the first sample data in step S201 can include attributes related to the target behavior, such as the user's income, age, gender, and monthly spending, which are related to whether or not they have purchased the product or service. Thus, the pre-trained tree structure can be used to determine the attributes that are strongly related to the target behavior and their corresponding value ranges.

[0048] In step S202, each node in the tree structure represents a value range corresponding to one of the multiple attributes. For example, when the first sample data has attributes such as monthly income, age, and monthly spending, the nodes in the obtained tree structure can represent the specific value space of these attributes. For instance, node 1 represents a monthly income greater than 5000, node 2 represents an age less than 30, and node 3 represents a monthly spending greater than 3000. For example, each node can be a single rule, or multiple nodes can be combined into a single rule. A rule set can be obtained by combining multiple rules.

[0049] The nodes in the tree structure obtained through pre-training are usually more related to whether the sample users perform the target behavior. For example, users with a monthly income of no more than 5,000 tend not to buy a certain type of product. Therefore, in step S203, these nodes themselves can be used as rules to vectorize the features of the users to be classified in order to obtain the feature vector of the target users. In step S204, the user group is further divided based on the feature vector to achieve the classification of users.

[0050] Therefore, by using nodes in the pre-trained tree structure as rules to determine user feature vectors, the automatic generation of rules and the encoding of user features are achieved. Furthermore, based on the feature vectors obtained from the rules, users within the user group are clustered, achieving automatic user classification and improving the accuracy of user classification.

[0051] According to some embodiments, the tree structure is obtained by pre-training the ensemble tree model using the first sample dataset. The ensemble tree model can be one or a combination of algorithms including, but not limited to, Gradient Boosting Decision Tree (GBDT) and eXtreme Gradient Boosting (XGBoost), and this disclosure does not limit the type of model selected.

[0052] According to some embodiments, the tree structure generated based on the first sample dataset pre-training includes at least one decision tree. Step S202 includes: for each leaf node in each decision tree in the at least one decision tree, combining each non-leaf node on the path from the root node of the decision tree to the leaf node to obtain the rule corresponding to the leaf node; and aggregating the rules corresponding to each leaf node in each decision tree in the at least one decision tree to obtain the rule set.

[0053] Figure 3 A schematic diagram of a decision tree according to an embodiment of the present disclosure is shown. It will be understood that the tree structure generated based on the first sample dataset pre-trained includes at least one such... Figure 3The decision tree shown.

[0054] like Figure 3 As shown, for the first leaf node in the decision tree, based on the path from the root node to the first leaf node, combining the non-leaf nodes on the path yields the rule {age < 30, monthly income < 5000, monthly consumption < 2000}. Similarly, the rule corresponding to each leaf node in the decision tree can be obtained. The rule set is obtained by combining the rules corresponding to each leaf node of each decision tree in the at least one decision tree.

[0055] Therefore, by traversing the path from the root node to the leaf node of each decision tree, new attribute combinations are automatically generated based on the tree structure as rules for determining feature vectors, without needing to formulate rules based on business experience, thus improving the accuracy of vectorization of the second sample data. Furthermore, rules generated in this way enable the model to be transferred and reused between different samples in the same scenario.

[0056] According to some embodiments, the second sample data has multiple attribute values ​​corresponding to the multiple attributes. Step S203 includes: for at least one rule in the rule set, obtaining the attribute value of the second sample data corresponding to the rule; and in response to determining that the attribute value of the second sample data conforms to the rule, determining that the feature vector of the second sample data takes a first value on the dimension corresponding to the rule, or in response to determining that the attribute value of the second sample data does not conform to the rule, determining that the feature vector of the second sample data takes a second value on the dimension corresponding to the rule, so as to determine the feature vector of the second sample data.

[0057] Taking the rule {age < 30, monthly income < 5000, monthly spending < 2000} as an example, we obtain the attribute values ​​corresponding to the second sample data, i.e., the second sample data has an age of 25, a monthly income of 6000 yuan, and a monthly spending of 3000 yuan. Since the monthly income and monthly spending of this second sample data do not meet the rule, we can determine that the dimension corresponding to the rule in the feature vector of this second sample data takes the second value. In one example, the first value is 1, and the second value is 0, that is, when the rule is met, the corresponding dimension takes the value of 1, otherwise the corresponding dimension takes the value of 0. Only when all attribute values ​​of the second sample data, i.e., age, monthly income, and monthly spending, meet the rule {age < 30, monthly income < 5000, monthly spending < 2000}, the dimension corresponding to the rule in the feature vector of this second sample data takes the value of 1. Therefore, the feature vector of each second sample data in the second sample dataset can be determined by traversing each rule in the rule set.

[0058] According to some embodiments, step S204 includes: using Principal Component Analysis (PCA) to reduce the dimensionality of the feature vectors of the second sample data to obtain the dimensionality-reduced vectors corresponding to the second sample data; and clustering the second sample data in the second sample dataset based on the dimensionality-reduced vectors corresponding to the second sample data to divide the second sample dataset into at least one subset of second sample data.

[0059] Since tree structures typically contain many decision trees, the resulting rule set also contains a large number of rules, leading to feature vectors with high dimensionality and computational complexity. Dimensionality reduction of the feature vectors from the second sample data using PCA can eliminate data from dimensions with low importance or minimal impact on classification results, thereby reducing the difficulty of clustering and improving its efficiency and accuracy.

[0060] For example, a Gaussian Mixed Model (GMM) can be used to cluster the second sample data in the second sample dataset based on the dimensionality reduction vector corresponding to each second sample data, so as to divide the second sample dataset into at least one subset of second sample data, thereby achieving the classification of the second sample data.

[0061] Figure 4 A flowchart of a portion of a data processing method according to an embodiment of the present disclosure is shown.

[0062] like Figure 4 As shown, based on steps S201-S204, the data processing method 200 further includes: step S401, determining the feature vector of the first sample data in the first sample dataset based on the rule set;

[0063] Step S402: Based on the feature vector of the first sample data, divide the first sample dataset into at least one subset of first sample data;

[0064] Step S403: For the first sample data in the first sample dataset, based on the subset of the first sample data where the first sample data is located, determine the classification features corresponding to the first sample data; and

[0065] Step S404: Based on the classification features corresponding to the first sample data, retrain the ensemble tree model to obtain an updated tree structure.

[0066] Understandably, the same operation can be repeated on the first sample data used as the training set to obtain the classification result of the first sample data. This classification result can be added as a new one-dimensional feature to the original feature list used to train the tree model, thereby retraining the tree model and improving its training performance. Furthermore, in subsequent processes, classifying the target user based on the retrained tree structure can further improve the classification effect.

[0067] In one example, weight of evidence (WOE) encoding can be used to encode the classification result of the first sample data, so that the classification result is reflected as a new one-dimensional feature in the pre-training process.

[0068] According to another aspect of this disclosure, a data processing apparatus is provided. For example... Figure 5 As shown, the data processing device 500 includes: an acquisition module 501 configured to acquire a tree structure generated based on pre-training of a first sample dataset, wherein the first sample data in the first sample dataset has multiple attributes; a generation module 502 configured to generate a rule set based on the nodes of the tree structure, wherein the nodes of the tree structure represent the value range corresponding to one of the multiple attributes, and the rule set is a set of at least one node; a first determination module 503 configured to determine the feature vector of the second sample data in the second sample dataset based on the rule set, wherein each dimension of the feature vector corresponds to each rule in the rule set; and a first partitioning module 504 configured to partition the second sample dataset into at least one subset of second sample data based on the feature vector of the second sample data in the second sample dataset.

[0069] Understandably, the first sample dataset serves as the training set to train a tree structure for partitioning the second sample dataset. The second sample dataset can be the target user group to be classified. The data processing device 500 uses the tree structure trained on the first sample dataset to perform feature representation and classification of the target users in the second sample dataset. The tree structure acquired by the acquisition module 501 can be selected from the appropriate first sample dataset for training based on the specific application scenario. Furthermore, the multiple attributes of the first sample data can be determined based on the application scenario and corresponding user behavior.

[0070] For example, when classifying a user group for the purpose of promoting or recommending a certain type of product or service, the first sample dataset can be a user group consisting of users who have purchased the product or service and users who have not purchased the product or service. Accordingly, in this example, the multiple attributes of the first sample data acquired by the acquisition module 501 can include attributes related to the target behavior, such as the user's income, age, gender, and monthly spending, which are related to whether or not they have purchased the product or service. Thus, the pre-trained tree structure can be used to determine the attributes that are strongly related to the target behavior and their corresponding value ranges.

[0071] Each node in the tree structure represents a value range corresponding to one of the multiple attributes. For example, when the first sample data has attributes such as monthly income, age, and monthly spending, the nodes in the tree structure acquired by the acquisition module 501 can represent the specific value space of these attributes. For instance, node 1 represents a monthly income greater than 5000, node 2 represents an age less than 30, and node 3 represents a monthly spending greater than 3000. For example, the generation module 502 can treat each node as a separate rule, or combine multiple nodes into a single rule. By aggregating multiple rules, a rule set can be obtained.

[0072] The nodes in the tree structure obtained through pre-training are usually more related to whether the sample users perform the target behavior. For example, users with a monthly income of no more than 5,000 tend not to buy a certain type of product. Therefore, the first determination module 503 can use these nodes themselves as rules to vectorize the features of the users to be classified in order to obtain the feature vector of the target user. The first partitioning module 504 further partitions the user group based on the feature vector to achieve user classification.

[0073] Therefore, the generation module 502 uses the nodes in the pre-trained tree structure as rules to determine user feature vectors, thus achieving automated rule generation and user feature encoding. The first segmentation module 504 further clusters users in the user group based on the feature vectors obtained from the rules, achieving automated user classification and improving the accuracy of user classification.

[0074] According to some embodiments, the tree structure is obtained by pre-training the ensemble tree model using the first sample dataset. The ensemble tree model can be one or a combination of algorithms including, but not limited to, Gradient Boosting Decision Tree (GBDT) and eXtreme Gradient Boosting (XGBoost), and this disclosure does not limit the type of model selected.

[0075] According to some embodiments, the tree structure obtained by the acquisition module 501 based on the pre-trained first sample dataset includes at least one decision tree, and the generation module 502 includes: a combination unit configured to combine each non-leaf node on the path from the root node to the leaf node of each decision tree in the at least one decision tree to obtain the rule corresponding to the leaf node; and a set unit configured to set the rules corresponding to each leaf node of each decision tree in the at least one decision tree to obtain the rule set.

[0076] Therefore, by traversing the path from the root node to the leaf node of each decision tree, new attribute combinations are automatically generated based on the tree structure as rules for determining feature vectors, without needing to formulate rules based on business experience, thus improving the accuracy of vectorization of the second sample data. Furthermore, rules generated in this way enable the model to be transferred and reused between different samples in the same scenario.

[0077] According to some embodiments, the second sample data has multiple attribute values ​​corresponding to the multiple attributes. The first determining module 503 includes: an acquisition unit configured to acquire the attribute value of the second sample data corresponding to the rule for at least one rule in the rule set; and a determining unit configured to, in response to determining that the attribute value of the second sample data conforms to the rule, determine that the feature vector of the second sample data takes a first value in the dimension corresponding to the rule, or in response to determining that the attribute value of the second sample data does not conform to the rule, determine that the feature vector of the second sample data takes a second value in the dimension corresponding to the rule, thereby determining the feature vector of the second sample data.

[0078] Taking the rule {age < 30, monthly income < 5000, monthly consumption < 2000} as an example, the acquisition unit obtains the attribute values ​​corresponding to the second sample data, i.e., the age of the second sample data is 25 years old, the monthly income is 6000 yuan, and the monthly consumption is 3000 yuan. Since the monthly income and monthly consumption of the second sample data do not meet the rule, the determination unit can determine that the dimension corresponding to the rule in the feature vector of the second sample data takes the second value. In one example, the first value is 1, and the second value is 0, that is, when the rule is met, the corresponding dimension takes the value of 1, otherwise the corresponding dimension takes the value of 0. Only when all attribute values ​​of the second sample data, i.e., age, monthly income, and monthly consumption, meet the rule {age < 30, monthly income < 5000, monthly consumption < 2000}, the dimension corresponding to the rule in the feature vector of the second sample data takes the value of 1. Therefore, the first determination module 503 can determine the feature vector of each second sample data in the second sample dataset by traversing each rule in the rule set.

[0079] According to some embodiments, the first partitioning module 504 includes: a dimensionality reduction unit configured to use principal component analysis to reduce the dimensionality of the feature vector of the second sample data to obtain the dimensionality reduction vector corresponding to the second sample data; and a clustering unit configured to cluster the second sample data in the second sample dataset based on the dimensionality reduction vector corresponding to the second sample data to divide the second sample dataset into at least one subset of second sample data.

[0080] Since tree structures typically include many decision trees, the resulting rule set also contains a large number of rules, leading to feature vectors with high dimensionality and computational complexity. The dimensionality reduction unit, using PCA to reduce the dimensionality of the feature vectors from the second sample data, can eliminate data from dimensions with low importance or minimal impact on the classification results, thereby reducing the difficulty of clustering and improving its efficiency and accuracy.

[0081] For example, the clustering unit can use a Gaussian Mixed Model (GMM) to cluster the second sample data in the second sample dataset based on the dimensionality reduction vector corresponding to each second sample data, so as to divide the second sample dataset into at least one subset of second sample data, thereby realizing the classification of the second sample data.

[0082] According to some embodiments, the data processing apparatus 500 further includes: a second determining module configured to determine a feature vector of a first sample data in the first sample dataset based on the rule set; a second partitioning module configured to partition the first sample dataset into at least one subset of first sample data based on the feature vector of the first sample data; a third determining module configured to determine a classification feature corresponding to the first sample data in the first sample dataset based on the subset of first sample data in which the first sample data is located; and a training module configured to retrain the ensemble tree model based on the classification feature corresponding to the first sample data to obtain an updated tree structure.

[0083] Understandably, the data processing device 500 can repeat the same operation on the first sample data used as the training set to obtain the classification result of the first sample data. This classification result can be added as a new one-dimensional feature to the original feature list used to train the tree model, so that the training module can retrain the tree model to improve the training effect of the tree model. Furthermore, in subsequent processes, classifying the target user based on the retrained tree structure can further improve the classification effect.

[0084] In one example, the third determination module can use weight of evidence (WOE) encoding to encode the classification result of the first sample data, so as to incorporate the classification result as a new one-dimensional feature in the pre-training process.

[0085] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a data processing method.

[0086] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to cause the computer to perform a data processing method.

[0087] According to another aspect of this disclosure, a computer program product is also provided, including a computer program, wherein the computer program implements a data processing method when executed by a processor.

[0088] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0089] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, output unit 607, storage unit 608, and communication unit 609. Input unit 606 can be any type of device capable of inputting information to electronic device 600. Input unit 606 can receive input digital or character information and generate key signal input related to user settings and / or function control of electronic device, and can include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 607 can be any type of device capable of presenting information, and can include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 608 can include, but is not limited to, disk and optical disk. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and can include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth. TM Devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.

[0090] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform data processing methods by any other suitable means (e.g., by means of firmware).

[0091] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0092] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0093] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0094] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0095] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0096] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0097] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0098] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A data processing method, comprising: Obtain a tree structure generated based on pre-training on a first sample dataset, wherein the first sample data in the first sample dataset has multiple attributes, wherein the multiple attributes include at least one of the sample user's income, age, gender, and monthly consumption, and wherein the first sample dataset is a user group composed of users who have purchased the target product or service and users who have not purchased the target product or service, and the multiple attributes include attributes related to whether or not the target product or service has been purchased; Based on the nodes of the tree structure, a rule set is generated, wherein each node of the tree structure represents a value range corresponding to one of the multiple attributes, the rule set is a set of at least one node, and the tree structure generated based on the first sample dataset pre-training includes at least one decision tree. The generation of the rule set based on the nodes of the tree structure includes: For each leaf node in each of the at least one decision trees, each non-leaf node on the path from the root node to that leaf node is combined to obtain the rule corresponding to that leaf node; and The rules corresponding to each leaf node of each decision tree in the at least one decision tree are set together to obtain the rule set; Based on the rule set, feature vectors of the second sample data in the second sample dataset are determined, wherein each dimension of the feature vector corresponds to a rule in the rule set; and Based on the feature vectors of the second sample data in the second sample dataset, the second sample dataset is divided into at least one subset of second sample data for use in recommending the target products or services to target users under different categories.

2. The method according to claim 1, wherein, The second sample data has multiple attribute values ​​corresponding to the multiple attributes, and determining the feature vector of the second sample data in the second sample dataset based on the rule set includes: For at least one rule in the rule set, obtain the attribute value of the second sample data corresponding to that rule; and In response to determining that the attribute value of the second sample data conforms to the rule, the feature vector of the second sample data is determined to have a first value in the dimension corresponding to the rule, or in response to determining that the attribute value of the second sample data does not conform to the rule, the feature vector of the second sample data is determined to have a second value in the dimension corresponding to the rule, so as to determine the feature vector of the second sample data.

3. The method according to claim 1 or 2, wherein, The step of dividing the second sample dataset into at least one subset of second sample data based on the feature vectors of the second sample data in the second sample dataset includes: Principal component analysis (PCA) is used to reduce the dimensionality of the feature vectors of the second sample data, resulting in a dimensionality-reduced vector corresponding to the second sample data; and Clustering is performed on the second sample data in the second sample dataset based on the dimensionality reduction vector corresponding to the second sample data, so as to divide the second sample dataset into at least one subset of second sample data.

4. The method according to claim 1 or 2, wherein, The tree structure is obtained by pre-training the ensemble tree model using the first sample dataset.

5. The method according to claim 4, further comprising: Based on the set of rules, determine the feature vector of the first sample data in the first sample dataset; Based on the feature vector of the first sample data, the first sample dataset is divided into at least one subset of the first sample data; For the first sample data in the first sample dataset, the classification features corresponding to the first sample data are determined based on the subset of the first sample data to which the first sample data is located; as well as Based on the classification features corresponding to the first sample data, the ensemble tree model is retrained to obtain an updated tree structure.

6. A data processing apparatus, comprising: The acquisition module is configured to acquire a tree structure generated based on a pre-trained first sample dataset, wherein the first sample data in the first sample dataset has multiple attributes, wherein the multiple attributes include at least one of the sample user's income, age, gender, and monthly consumption, and wherein the first sample dataset is a user group composed of users who have purchased the target product or service and users who have not purchased the target product or service, and the multiple attributes include attributes related to whether or not the target product or service has been purchased; A generation module is configured to generate a rule set based on the nodes of the tree structure, wherein each node of the tree structure represents a value range corresponding to one of the plurality of attributes, the rule set is a set of at least one node, and wherein the tree structure obtained by the acquisition module based on the pre-trained first sample dataset includes at least one decision tree. The generation module includes: The combination unit is configured to, for each leaf node in each of the at least one decision trees, combine each non-leaf node on the path from the root node of the decision tree to the leaf node to obtain the rule corresponding to the leaf node; and The set unit is configured to set the rules corresponding to each leaf node of each decision tree in the at least one decision tree to obtain the rule set; The first determining module is configured to determine, based on the rule set, the feature vectors of the second sample data in the second sample dataset, wherein each dimension of the feature vector corresponds to a rule in the rule set; and The first segmentation module is configured to divide the second sample dataset into at least one subset of second sample data based on the feature vector of the second sample data in the second sample dataset, for use in recommending the target product or service to target users under different categories.

7. The apparatus according to claim 6, wherein, The second sample data has multiple attribute values ​​corresponding to the multiple attributes, and the first determining module includes: The acquisition unit is configured to acquire, for at least one rule in the rule set, the attribute value of the second sample data corresponding to that rule; and The determining unit is configured to determine the feature vector of the second sample data in response to determining that the attribute value of the second sample data conforms to the rule, and to determine that the feature vector of the second sample data in response to determining that the attribute value of the second sample data does not conform ... does not conform to the rule.

8. The apparatus according to claim 6 or 7, wherein, The first partitioning module includes: The dimensionality reduction unit is configured to use principal component analysis to reduce the dimensionality of the feature vectors of the second sample data, so as to obtain the dimensionality-reduced vectors corresponding to the second sample data; and The clustering unit is configured to cluster the second sample data in the second sample dataset based on the dimensionality reduction vector corresponding to the second sample data, so as to divide the second sample dataset into at least one subset of the second sample data.

9. The apparatus according to claim 6 or 7, wherein, The tree structure is obtained by pre-training the ensemble tree model using the first sample dataset.

10. The apparatus according to claim 9, further comprising: The second determining module is configured to determine the feature vector of the first sample data in the first sample dataset based on the rule set. The second partitioning module is configured to divide the first sample dataset into at least one subset of first sample data based on the feature vector of the first sample data. The third determining module is configured to determine the classification features corresponding to the first sample data based on the first sample data subset in the first sample dataset. as well as The training module is configured to retrain the ensemble tree model based on the classification features corresponding to the first sample data to obtain an updated tree structure.

11. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

13. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment, and medium

    CN114021650A