Method and device for processing data

By extracting and screening key features from user historical data and building a portrait decision tree, we solve the problems of low precision and efficiency of user portrait labels under large-scale historical data, and achieve more refined and efficient user portrait management.

CN120670805APending Publication Date: 2025-09-19BEIJING JINGDONG QIANSHITECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410308984.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-18
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies have low precision and efficiency when determining user portrait labels based on large-scale historical data.

Method used

Extract multiple user features from user historical data, screen out key features, build a portrait decision tree, and use the branch path and feature value of the decision tree to determine the portrait label of the target user.

Benefits of technology

The sophistication and processing efficiency of user portrait labels have been improved, and the analysis and management flexibility of user portraits have been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670805A_ABST
    Figure CN120670805A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, and relates to the technical field of intelligent supply chains. According to one specific embodiment, the method comprises the steps that multiple user features can be extracted from historical data of a user, key features are screened out from the multiple user features, the key features and feature values, corresponding to the key features, included in the historical data are utilized to construct a portrait decision tree, and the constructed portrait decision tree comprises multiple branch paths; according to the key feature and the feature value on each branch path, determining a portrait label of a target user on a leaf node corresponding to the branch path; according to the embodiment of the invention, the key features are determined, the portrait decision tree is constructed, and the corresponding relationship between the portrait label and the feature value is visually reflected through the portrait decision tree, so that the precision degree of determining the portrait label is improved, and the efficiency of processing data and determining the portrait label is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent supply chain, and in particular to a method and device for processing data. Background Art

[0002] With the widespread application of smart supply chains, managers of smart supply chains usually need to manage the business information of the smart supply chains, and also need to manage the business status of each item provider in the smart supply chain (such as user portraits).

[0003] Currently, feature extraction is usually performed on the historical data of the item provider, and user portrait labels are determined manually based on the extracted features according to historical experience, or basic user features are selected, and classification processing is performed on the selected basic features using a model, and the classification results are directly used as user portrait labels. Existing methods have the problem of low precision and low efficiency in determining user portrait labels when the historical data is large in scale, the number of features contained is far greater than the basic features, and the order of magnitude is large. Summary of the Invention

[0004] In view of this, an embodiment of the present invention provides a method and device for processing data, which can extract multiple user features from the user's historical data, and screen out key features from the multiple user features, and use the key features and the feature values ​​corresponding to the key features included in the historical data to construct a portrait decision tree, wherein the constructed portrait decision tree includes multiple branch paths, and according to the key features and feature values ​​on each of the branch paths, the portrait label of the target user on the leaf node corresponding to the branch path is determined; the embodiment of the present invention improves the precision of determining portrait labels and improves the efficiency of processing data and determining portrait labels by determining key features and constructing a portrait decision tree, and intuitively reflects the correspondence between portrait labels and feature values ​​through the portrait decision tree.

[0005] To achieve the above-mentioned purpose, according to one aspect of an embodiment of the present invention, a method for processing data is provided, which is characterized by comprising: obtaining historical data of multiple users; extracting multiple user features associated with user portraits from the historical data, and screening out a first set number of key features from the multiple user features; the importance of the key features associated with the user features of the user portrait; constructing a portrait decision tree using the first set number of key features and the feature values ​​corresponding to the key features included in the historical data, wherein the constructed portrait decision tree includes multiple branch paths, and the end of each branch path has a leaf node for determining the target user; the root node and each child node included in the branch path correspond to one key feature, and the different branches connected to the root node and each child node are used to distinguish different feature values ​​of the same key feature; based on the key features and feature values ​​on each branch path included in the portrait decision tree, the portrait label of the target user on the leaf node corresponding to the branch path is determined.

[0006] Optionally, the extracting of multiple user features associated with user portraits from the historical data includes: extracting one or more of the following user features from the historical data: trend features, periodic features, volatility features, special time features, statistical value features, time series features, classification features, and custom features; wherein the time series features are determined by a time series data feature mining model; and the classification features are determined by a clustering model.

[0007] Optionally, the screening out of a first set number of key features from the multiple user features includes: using a random forest model to obtain primary key features exceeding the first set number from the multiple user features; using the Gini index and / or the out-of-bag data error rate as evaluation indicators to determine the order of importance of the primary key features; and screening out a first set number of primary key features ranked at the top as key features according to the order of importance of the primary key features.

[0008] Optionally, the method for processing data further includes: selecting a second set number of other user features other than the key features; for each of the other user features, executing: determining the proportion of the characteristic value of the other user feature in the characteristic value of the user feature of the target user; when the proportion exceeds a set threshold, using the other user features and their characteristic values ​​as new portrait labels for the user.

[0009] Optionally, the selecting a second set number of other user features other than the key feature includes: obtaining the order of importance of the user features according to a preset strategy; and selecting a second set number of other user features other than the key feature that are ranked higher according to the order of importance of the user features.

[0010] Optionally, the method for processing data further includes: configuring a corresponding first label for the key feature and the feature value corresponding to the key feature, and / or configuring a corresponding second label for the other user features and the feature value corresponding to the other user features; assigning a portrait label to the user according to the first label corresponding to the key feature and / or the second label corresponding to the other user features.

[0011] Optionally, the classification features are determined through a clustering model, including: obtaining multiple classification indicators corresponding to the user features; for each classification indicator, executing: using a clustering model to perform a clustering operation on the user features based on the classification indicator, and determining the classification indicator label corresponding to each cluster cluster according to the multiple cluster clusters contained in the clustering results of the clustering operation; and constructing an association relationship between the classification indicator label and the classification feature of the user.

[0012] To achieve the above-mentioned purpose, according to a second aspect of an embodiment of the present invention, a device for processing data is provided, characterized in that it includes: a feature extraction module, a decision tree construction module and a label determination module; wherein,

[0013] The feature extraction module is used to obtain historical data of multiple users; extract multiple user features associated with user profiles from the historical data, and screen a first set number of key features from the multiple user features; the importance of the key features associated with the user features of the user profiles;

[0014] The decision tree construction module is used to construct a portrait decision tree using the first set number of key features and the feature values ​​corresponding to the key features included in the historical data, wherein the constructed portrait decision tree includes multiple branch paths, and the end of each branch path has a leaf node for determining the target user; the root node and each child node included in the branch path correspond to one of the key features, and the different branches connected to the root node and each of the child nodes are used to distinguish different feature values ​​of the same key feature;

[0015] The label determination module is used to determine the portrait label of the target user on the leaf node corresponding to the branch path based on the key features and feature values ​​on each branch path included in the portrait decision tree.

[0016] Optionally, the data processing device is used to extract multiple user features associated with user portraits from the historical data, including: extracting one or more of the following user features from the historical data: trend features, periodic features, volatility features, special time features, statistical value features, time series features, classification features, and custom features; wherein the time series features are determined by a time series data feature mining model; and the classification features are determined by a clustering model.

[0017] Optionally, the data processing device is used to screen out a first set number of key features from the multiple user features, including: using a random forest model to obtain primary key features exceeding the first set number from the multiple user features; using the Gini index and / or the out-of-bag data error rate as evaluation indicators to determine the order of importance of the primary key features; and screening out a first set number of primary key features ranked at the top as key features according to the order of importance of the primary key features.

[0018] Optionally, the data processing device is further used to select a second set number of other user features in addition to the key features; for each of the other user features, perform: determining the proportion of the characteristic value of the other user feature in the characteristic value of the user feature of the target user; if the proportion exceeds a set threshold, use the other user feature and its characteristic value as a new portrait label for the user.

[0019] Optionally, the data processing device is used to select a second set number of other user features other than the key feature, including: obtaining the order of importance of the user features according to a preset strategy; and selecting a second set number of other user features other than the key feature that are ranked higher according to the order of importance of the user features.

[0020] Optionally, the data processing device is also used to configure a corresponding first label for the key feature and the feature value corresponding to the key feature, and / or, to configure a corresponding second label for the other user features and the feature value corresponding to the other user features; and assign a portrait label to the user according to the first label corresponding to the key feature and / or the second label corresponding to the other user features.

[0021] Optionally, the classification features in the data processing device are determined by a clustering model, including: obtaining multiple classification indicators corresponding to the user features; for each classification indicator, executing: using a clustering model to perform a clustering operation on the user features based on the classification indicator, and determining the classification indicator label corresponding to each cluster cluster according to the multiple cluster clusters contained in the clustering results of the clustering operation; and constructing an association relationship between the classification indicator label and the classification feature of the user.

[0022] To achieve the above-mentioned purpose, according to the third aspect of an embodiment of the present invention, there is provided an electronic device for processing data, characterized in that it includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement any of the methods described above for processing data.

[0023] To achieve the above-mentioned purpose, according to a fourth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, characterized in that when the program is executed by a processor, any method as described in the above-mentioned method for processing data is implemented.

[0024] An embodiment of the above invention has the following advantages or beneficial effects: it can extract multiple user features from the user's historical data, and screen out key features from the multiple user features, and use the key features and the feature values ​​corresponding to the key features included in the historical data to construct a portrait decision tree, the constructed portrait decision tree contains multiple branch paths, and according to the key features and feature values ​​on each of the branch paths, the portrait label of the target user on the leaf node corresponding to the branch path is determined; the embodiment of the present invention improves the precision of determining portrait labels and improves the efficiency of processing data and determining portrait labels by determining key features and constructing a portrait decision tree, and intuitively reflects the correspondence between portrait labels and feature values ​​through the portrait decision tree.

[0025] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings are provided for a better understanding of the present invention and are not intended to limit the present invention.

[0027] Figure 1 This is a flow chart of a method for processing data provided by one embodiment of the present invention;

[0028] Figure 2 is a flowchart of another method for processing data provided by one embodiment of the present invention;

[0029] Figure 3 1 is a schematic structural diagram of a data processing device provided by an embodiment of the present invention;

[0030] Figure 4 is an exemplary system architecture diagram in which embodiments of the present invention may be applied;

[0031] Figure 5 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, in which various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0033] It should be noted that the collection, collection, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solutions disclosed herein all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken with respect to user personal information to prevent unauthorized access to user personal information data and to safeguard the security of user personal information, network security, and national security.

[0034] It should be noted that in the technical solution of the present invention, the collection, use, storage, sharing and transfer of user personal information involved are in compliance with the provisions of relevant laws and regulations, and it is necessary to inform the user and obtain the user's consent or authorization. When applicable, the user's personal information is de-identified and / or anonymized and / or encrypted.

[0035] When collecting historical data or displaying portrait tags, we will use technical means to de-identify the data.

[0036] like Figure 1 As shown, an embodiment of the present invention provides a method for processing data, which may include the following steps:

[0037] Step S101: Acquire historical data of multiple users; extract multiple user features associated with user portraits from the historical data, and screen out a first set number of key features from the multiple user features; the key features are associated with the importance of the user features of the user portraits.

[0038] Specifically, in various application scenarios, in order to improve the data support of the manager and the personalized management of users, there is a need to determine user portraits for users; for example, the manager is the manager of the goods supply chain or the goods provider; it is understandable that the manager of the goods supply chain has the need to determine store portraits for each store it manages, and the goods provider (such as the store) has the need to determine the buyer portrait of the buyer (consumer) who purchases the goods; for the manager of the goods supply chain, the users it manages are stores, and the corresponding user portraits are portraits for stores; for the goods provider, the users it manages are buyers (consumers) who purchase the goods, and the corresponding user portraits are portraits for buyers.

[0039] In embodiments of the present invention, user historical data refers to actual data associated with existing users. For example, if the user is a store, the historical data includes basic store information, various sales data, and information about items sold. It is understood that a user profile label can be determined from the historical data by extracting and processing features.

[0040] Furthermore, the extraction of multiple user features associated with user profiles from the historical data includes: extracting one or more of the following user features from the historical data: trend features, cycle features, volatility features, special time features, statistical value features, time series features, classification features, and custom features. Specifically, for store users, multiple features can be obtained based on the processing of multiple sales data contained in the historical data. For example, a linear trend test model or a Mann-Kendall trend test can be used to extract trend features (e.g., year-on-year growth trend, month-on-month growth trend, overall growth trend, etc.) from sales data; special time features can be extracted based on the impact of special time periods (e.g., holidays, etc.) on sales; statistical value features (e.g., total amount, total amount, etc.) can be determined using one or more statistical methods (e.g., calculating mean, variance, median, percentile, etc.); classification features of historical data can be extracted using clustering models and classification models (e.g., multiple types divided according to store addresses, classification based on sales data, etc.); that is, the classification features are determined by clustering models; and custom features (other features extracted based on users included in the application scenario for determining user profile labels) can also be extracted.

[0041] Autoregressive forecasting can be used to analyze time series features in sales data and extract temporal, cyclical, and volatility features. Furthermore, the tsfresh model can be used as a time series data feature mining model to extract multiple time series features from historical data. Specifically, the time series features are determined using the time series data feature mining model. The tsfresh model can automatically calculate a large number of time series features, such as basic time series features like peak count, average, or maximum, or more complex features like time-reversal symmetry statistics. Simultaneously, hypothesis testing is used to reduce features to those that best explain trends, a process known as decorrelation. Because the tsfresh model can extract features in batches, it improves both efficiency and accuracy.

[0042] Furthermore, the classification features are determined by a clustering model, including: obtaining a plurality of classification indicators corresponding to the user features; for each classification indicator, executing: using a clustering model to perform a clustering operation on the user features based on the classification indicator, and determining the classification indicator label corresponding to each cluster cluster according to the multiple cluster clusters contained in the clustering result of the clustering operation; and constructing a correlation relationship between the classification indicator label and the classification feature of the user. Specifically, a plurality of classification indicators can be determined according to the application scenario and the actual user characteristics of the user, such as: sales statistics indicators, growth trend indicators, etc., and the data points corresponding to the data corresponding to the classification indicators contained in the historical data are clustered for different indicators to obtain a plurality of cluster clusters. For example, for the sales statistics indicator, a plurality of sales statistics intervals and type labels corresponding to each sales statistics interval can be divided through a plurality of cluster clusters, such as: low sales, medium sales, high sales, etc. It can be understood that through clustering and classification results, features can be extracted from multiple dimensions, thereby improving the richness of features.

[0043] Furthermore, in an embodiment of the present invention, a set number of key features are screened out from a plurality of extracted features; the key features are associated with the importance of the user features of the user portrait; it is understandable that the number of features extracted from historical data may be in the thousands or hundreds; preferably, key features are further screened out from a larger number of features. Specifically, the method for screening out key features can use a random forest model to screen out primary key features exceeding a first set number (for example, 5, 10, etc.) from the extracted user features; then use the Gini index (Giniindex) and / or the out-of-bag data error rate as evaluation indicators to determine the importance ranking of the primary key features exceeding the first set number, and screen out the first set number of primary key features ranked at the top based on the first set number of features (for example, 5, 10, etc.) ranked at the top as key features; that is, screening out the first set number of key features from a plurality of user features includes: using a random forest model to obtain primary key features exceeding the first set number screened out from a plurality of user features; using the Gini index and / or the out-of-bag data error rate as evaluation indicators to determine the order of importance of the primary key features; according to the order of importance of the primary key features, screening out the first set number of primary key features ranked at the top as key features.

[0044] It can be seen that the embodiments of the present invention improve the accuracy and richness of features associated with user portraits by extracting one or more of the following user features from historical data: trend features, periodic features, volatility features, special time features, statistical value features, time series features, classification features, and custom features, and screen out key features that reflect importance (i.e., the importance of key features to user features of user portraits), thereby improving the accuracy of user portrait labels and portrait effects.

[0045] Step S102: Construct a portrait decision tree using the first set number of key features and the feature values ​​corresponding to the key features included in the historical data, wherein the constructed portrait decision tree includes multiple branch paths, and the end of each branch path has a leaf node for determining the target user; the root node and each child node included in the branch path correspond to one of the key features, and the different branches connected to the root node and each of the child nodes are used to distinguish different feature values ​​of the same key feature.

[0046] Specifically, after determining the first set number of key features through the method described in step S101, a portrait decision tree is constructed using the first set number of key features and the characteristic values ​​corresponding to the key features included in the historical data; it can be understood that the portrait decision tree can be a binary tree or a multi-branch tree, and the number of its branches is related to the number of characteristic value types of the key features; for example, the key features are: statistical values ​​of item sales, and the characteristic values ​​include, for example, sales statistics less than a first threshold, between the first threshold and the second threshold, and higher than the second threshold; that is, there are three types of characteristic values ​​of the key features, corresponding to the three branches derived from the root node or child node of the portrait decision tree. It can be seen that the portrait decision tree is constructed by combining the multiple characteristic values ​​of each key feature in the first set number of key features as nodes and branches.

[0047] Furthermore, the constructed portrait decision tree includes multiple branch paths, and the end of each branch path has a leaf node for determining the target user; the root node and each child node included in the branch path correspond to one of the key features, and the different branches connected to the root node and each of the child nodes are used to distinguish different feature values ​​of the same key feature; that is, the feature value of the user feature of the target user satisfies the feature value of each branch included in the branch path where the leaf node is located. It can be understood that there can be correlation between key features, and the correlation can be reflected by the branch path of the portrait decision tree; for example, the number of key features is 5, key feature 1 has two eigenvalues ​​A and B; key feature 2 has two eigenvalues ​​C and D, and so on; the following uses two key eigenvalues ​​as an example, for example: user A has eigenvalue A and eigenvalue C, user B has eigenvalue B and eigenvalue D, then user A and user B are divided into leaf nodes located on different branch paths of the portrait decision tree, that is, the different branches connected to the root node and each of the child nodes are used to distinguish different eigenvalues ​​of the same key feature. It can be understood that a leaf node can contain multiple users, for example, users 1-user 10 all meet the same eigenvalues ​​and are located in the same leaf node, for example, users 1-user 10 are all users with eigenvalue A, eigenvalue C, and eigenvalues ​​1...N.

[0048] It can be seen that by constructing a portrait decision tree, the specific feature values ​​of each user for key features can be obtained through the leaf nodes of the portrait decision tree, and the target user information on different branch paths can be obtained. By utilizing the correlation between the feature values ​​indicated by the branches of the portrait decision tree and the user features, the user can be located, thereby improving the efficiency of executing user portraits, improving the intuitiveness of the correlation between users and user feature values, and thus improving the analysis and management flexibility of user portraits.

[0049] Step S103: Determine the portrait label of the target user on the leaf node corresponding to the branch path according to the key features and feature values ​​on each branch path included in the portrait decision tree.

[0050] Specifically, according to the portrait decision tree determined in step S102, the portrait label of the target user on the leaf node corresponding to the branch path is assigned by the key features and feature values ​​on each branch path; for example, a node on a branch path in the portrait decision tree is a sales statistical value feature, and is divided into three branches according to different feature values ​​of the sales statistical value feature, the three branches being the sales statistical value less than a first threshold, the sales statistical value between the first threshold and the second threshold, and the sales statistical value greater than the second threshold; the user's portrait label can be set for different feature values, for example, the portrait label set for users with sales statistical values ​​less than the first threshold is low sales, the portrait label set for users with sales statistical values ​​between the first threshold and the second threshold is medium sales, and the portrait label set for users with sales statistical values ​​greater than the second threshold is high sales, etc. Similarly, for the branch path of the portrait decision tree to which each leaf node belongs, the portrait label of each user is determined by the key features and feature values ​​on its branch path. That is, the key features and the feature values ​​corresponding to the key features are configured with corresponding first labels, and the user is assigned a portrait label according to the first labels corresponding to the key features.

[0051] like Figure 2 As shown, an embodiment of the present invention provides another method for processing data, which may include the following steps:

[0052] Step S201: Acquire historical data of multiple users; extract multiple user features associated with user portraits from the historical data, and screen out a first set number of key features from the multiple user features; the key features are associated with the importance of the user features of the user portraits.

[0053] Step S202: Construct a portrait decision tree using the first set number of key features and the feature values ​​corresponding to the key features included in the historical data, wherein the constructed portrait decision tree includes multiple branch paths, and the end of each branch path has a leaf node for determining the target user; the root node and each child node included in the branch path correspond to one of the key features, and the different branches connected to the root node and each of the child nodes are used to distinguish different feature values ​​of the same key feature.

[0054] Step S203: Determine the portrait label of the target user on the leaf node corresponding to the branch path according to the key features and feature values ​​on each branch path included in the portrait decision tree.

[0055] Specifically, the description of steps S201 to S203 is consistent with the description of steps S101 to S103 and will not be repeated here.

[0056] Step S204: Select a second set number of other user features other than the key features; select a second set number of other user features other than the key features; for each of the other user features, perform: determine the proportion of the characteristic value of the other user feature in the characteristic value of the user feature of the target user; if the proportion exceeds a set threshold, use the other user feature and its characteristic value as a new portrait label for the user.

[0057] Specifically, in addition to the first set number of key features indicated by the portrait decision tree and the feature values ​​corresponding to the key features included in the historical data, more user features for determining the user portrait can be selected for the target user, that is, a second set number of other user features in addition to the first set number of key features.

[0058] The method for determining the second set number of other user features can be: for each of the other user features, executing: determining the proportion of the characteristic value of the other user feature in the characteristic value of the user feature of the target user; for example: other user feature 1 is the user's address. It can be understood that the user's address is correlated with the sales feature. If it is determined that more than 70% (that is, the proportion exceeds the set threshold) of all target users have addresses within the set range (for example: city center, commercial district, etc.), the specific address information corresponding to the user's address is used as the user's newly added portrait label (for example: urban area, suburbs, etc.), that is, when the proportion exceeds the set threshold, the other user features and their characteristic values ​​are used as the user's new portrait label.

[0059] Furthermore, when selecting a second set number of other user features in addition to the key features, the importance order of the user features can be determined using a random forest model in combination with the Gini index and / or out-of-bag data error rate as evaluation indicators; that is, the importance order of the user features is obtained according to a preset strategy; further, based on the order of importance of the user features, a second set number of other user features ranked higher in addition to the key features are selected. Embodiments of the present invention improve the accuracy of determining user profiles using user features by selecting other user features based on the order of importance of the user features.

[0060] It can be seen that by selecting a second set number of other user features in addition to the key features as new user portraits, the determined user portrait labels are enriched and supplemented, thereby improving the accuracy and feature richness of the determined user portraits.

[0061] Furthermore, the method further includes configuring a corresponding first tag for the key feature and the feature value corresponding to the key feature, and / or configuring a corresponding second tag for the other user feature and the feature value corresponding to the other user feature; and assigning a user profile tag based on the first tag corresponding to the key feature and / or the second tag corresponding to the other user feature. Specifically, by assigning a user profile tag based on the tag configured in combination with the feature value of the key feature and the other user feature, the efficiency, accuracy, and effectiveness of determining the user profile are improved, thereby improving the efficiency of the management in processing user data.

[0062] like Figure 3 As shown, an embodiment of the present invention provides a device 300 for processing data, including: a feature extraction module 301, a decision tree construction module 302 and a label determination module 303; wherein,

[0063] The feature extraction module 301 is configured to obtain historical data of multiple users; extract multiple user features associated with user profiles from the historical data; and select a first set number of key features from the multiple user features; the key features are associated with the importance of the user features of the user profiles;

[0064] The decision tree construction module 302 is configured to construct a portrait decision tree using the first set number of key features and the feature values ​​corresponding to the key features included in the historical data, wherein the constructed portrait decision tree includes multiple branch paths, each of which ends with a leaf node for determining the target user; the root node and each child node included in the branch path correspond to one of the key features, and the different branches connected to the root node and each of the child nodes are used to distinguish different feature values ​​of the same key feature;

[0065] The label determination module 303 is used to determine the portrait label of the target user on the leaf node corresponding to the branch path according to the key features and feature values ​​on each branch path included in the portrait decision tree.

[0066] An embodiment of the present invention also provides an electronic device for processing data, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in any of the above embodiments.

[0067] An embodiment of the present invention further provides a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, the method provided in any of the above embodiments is implemented.

[0068] Figure 4 An exemplary system architecture 400 is shown to which the method for processing data or the apparatus for processing data according to the embodiment of the present invention can be applied.

[0069] like Figure 4 As shown, system architecture 400 may include terminal devices 401, 402, 403, a network 404, and a server 405. Network 404 is used to provide a medium for communication links between terminal devices 401, 402, 403 and server 405. Network 404 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0070] Users can use terminal devices 401, 402, and 403 to interact with server 405 via network 404 to receive or send messages, etc. Various client applications can be installed on terminal devices 401, 402, and 403, such as e-commerce client applications, web browser applications, search applications, instant messaging tools, and email clients.

[0071] The terminal devices 401 , 402 , and 403 may be various electronic devices having a display screen and supporting various client applications, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0072] Server 405 may be a server that provides various services, such as a background management server that provides support for client applications used by users using terminal devices 401, 402, and 403. The background management server may process the received request to determine the user portrait tag and feed back the determined user portrait tag result to the terminal device.

[0073] It should be noted that the method for processing data provided in the embodiment of the present invention is generally executed by the server 405 , and accordingly, the device for processing data is generally set in the server 405 .

[0074] It should be understood that Figure 4 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0075] Reference below Figure 5 , which shows a schematic structural diagram of a computer system 500 of a terminal device suitable for implementing an embodiment of the present invention. Figure 5The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0076] like Figure 5 As shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the system 500 are also stored in the RAM 503. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0077] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, and the like; an output section 507 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 508 including a hard disk; and a communication section 509 including a network interface card such as a LAN card or a modem. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 510 as needed, so that computer programs read therefrom can be installed into the storage section 508 as needed.

[0078] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from a removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above-mentioned functions defined in the system of the present invention are performed.

[0079] It should be noted that the computer-readable medium described in the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.

[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0081] The modules and / or units involved in the embodiments of the present invention may be implemented in software or hardware. The modules and / or units described may also be provided in a processor. For example, they may be described as follows: a processor includes a feature extraction module, a decision tree construction module, and a label determination module. The names of these modules do not, in some cases, constitute a limitation on the modules themselves. For example, the feature extraction module may also be described as a "module for selecting key features from a variety of user features."

[0082] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes: obtaining historical data of multiple users; extracting multiple user features associated with user portraits from the historical data, and screening out a first set number of key features from the multiple user features; constructing a portrait decision tree using the first set number of key features and the feature values ​​corresponding to the key features included in the historical data, wherein the constructed portrait decision tree includes multiple branch paths, and the end of each branch path has a leaf node for determining the target user; the root node and each child node included in the branch path correspond to one key feature, and the different branches connected to the root node and each child node are used to distinguish different feature values ​​of the same key feature; based on the key features and feature values ​​on each branch path included in the portrait decision tree, the portrait label of the target user on the leaf node corresponding to the branch path is determined.

[0083] An embodiment of the present invention extracts multiple user features from the user's historical data, and screens out key features from the multiple user features, and uses the key features and the feature values ​​corresponding to the key features included in the historical data to construct a portrait decision tree, wherein the constructed portrait decision tree includes multiple branch paths, and according to the key features and feature values ​​on each of the branch paths, the portrait label of the target user on the leaf node corresponding to the branch path is determined; the embodiment of the present invention improves the precision of determining the portrait label and improves the efficiency of processing data and determining the portrait label by determining the key features and constructing the portrait decision tree, and intuitively reflects the correspondence between the portrait label and the feature value through the portrait decision tree.

[0084] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for processing data, characterized in that include: Acquire historical data of multiple users; extract multiple user features associated with user profiles from the historical data, and screen a first set number of key features from the multiple user features; determine the importance of the key features to the user features associated with the user profiles; Constructing a portrait decision tree using the first set number of key features and the feature values ​​corresponding to the key features included in the historical data, wherein the constructed portrait decision tree includes multiple branch paths, each of which ends with a leaf node for determining a target user; a root node and each child node included in the branch path correspond to one of the key features, and different branches connected to the root node and each of the child nodes are used to distinguish different feature values ​​of the same key feature; According to the key features and feature values ​​on each branch path included in the portrait decision tree, the portrait label of the target user on the leaf node corresponding to the branch path is determined.

2. The method according to claim 1, characterized in that The extracting of multiple user features associated with the user profile from the historical data includes: Extracting one or more of the following user features from the historical data: trend features, period features, volatility features, special time features, statistical value features, time series features, classification features, and custom features; The time series features are determined by a time series data feature mining model; and the classification features are determined by a clustering model.

3. The method according to claim 1, characterized in that The step of selecting a first set number of key features from the plurality of user features includes: Using a random forest model to obtain primary key features exceeding the first set number from the plurality of user features; Determine the order of importance of the primary key features using the Gini index and / or the out-of-bag data error rate as evaluation indicators; According to the order of importance of the primary key features, a first set number of primary key features ranked at the top are screened out as key features.

4. The method according to claim 1, wherein Further including: selecting a second set number of other user features other than the key features; For each of the other user characteristics, execute: Determine the proportion of the characteristic values ​​of the other user characteristics in the characteristic values ​​of the user characteristics of the target user; if the proportion exceeds a set threshold, use the other user characteristics and their characteristic values ​​as new portrait tags for the user.

5. The method according to claim 4, characterized in that The selecting of a second set number of other user features other than the key features includes: Obtaining the order of importance of the user features according to a preset strategy; According to the order of importance of the user features, a second set number of other user features ranked higher than the key features are selected.

6. The method according to claim 4, characterized in that Also includes: Configuring a corresponding first tag for the key feature and the feature value corresponding to the key feature, and / or configuring a corresponding second tag for the other user feature and the feature value corresponding to the other user feature; A portrait label is assigned to the user according to the first label corresponding to the key feature and / or the second label corresponding to the other user features.

7. The method according to claim 2, characterized in that The classification features are determined by a clustering model and include: Obtaining multiple classification indicators corresponding to the user characteristics; For each classification indicator, execute: A clustering operation is performed on the user features based on the classification index using a clustering model, and a classification index label corresponding to each cluster is determined according to a plurality of clusters included in the clustering result of the clustering operation; and an association relationship between the classification index label and the classification feature of the user is established.

8. A device for processing data, characterized in that: include: Extract feature module, build decision tree module and determine label module; among them, The feature extraction module is used to obtain historical data of multiple users; extract multiple user features associated with user profiles from the historical data, and screen a first set number of key features from the multiple user features; the importance of the key features associated with the user features of the user profiles; The decision tree construction module is used to construct a portrait decision tree using the first set number of key features and the feature values ​​corresponding to the key features included in the historical data, wherein the constructed portrait decision tree includes multiple branch paths, and the end of each branch path has a leaf node for determining the target user; the root node and each child node included in the branch path correspond to one of the key features, and the different branches connected to the root node and each of the child nodes are used to distinguish different feature values ​​of the same key feature; The label determination module is used to determine the portrait label of the target user on the leaf node corresponding to the branch path based on the key features and feature values ​​on each branch path included in the portrait decision tree.

9. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Precision marketing oriented construction method of telecom user consumption portrait

    CN107358269A

  • Method and device for obtaining customer group feature association degree, storage medium and electronic device

    CN113538020A

  • User portrait generation method and device and computer equipment

    CN115329909A