User Classification Method, Device, Equipment and Storage Medium Based on Random Forest
The classification tree is constructed through the clustered sample subset to form a random forest for user classification, which solves the problems of large amount of user classification calculation and inaccurate classification in the existing technology, and achieves more efficient and accurate user classification.
Patent Information
- Application Number
- CN202111270906.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-10-29
AI Technical Summary
In the prior art, the user classification method has a large amount of calculation and inaccurate classification, resulting in low classification accuracy of the system user and affecting the efficiency of users to find the required products.
By obtaining the training sample set, clustering is performed to determine the sample subset, and then building a classification tree based on the sample subset to form a random forest for user classification.
The construction of similar decision trees is reduced, the fitting speed and accuracy of the classification tree is improved, and the accuracy and efficiency of user classification are improved.
Smart Images

Figure CN113920374B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence, and particularly relates to a user classification method, device, equipment and storage medium based on random forest. Background Art
[0002] When users use a wealth management platform or product purchase webpage, there are many types of products on the platform or page, and the target population for each product is also different. In order to enable users to efficiently find products that match them, the category of users is usually identified based on user characteristics, and then corresponding products are recommended according to the identified category.
[0003] Currently, the commonly used identification methods include forming a training set by sampling with replacement through the random forest algorithm, determining multiple decision trees based on the formed training set, and making decisions through the random forest composed of multiple decision trees. However, since the training set is formed by sampling with replacement in the random forest algorithm, duplicate data may be drawn, and there may be data that is not drawn. While the calculation amount is large, similar decision trees are likely to be generated, masking the true classification results, resulting in inaccurate classification results, low accuracy of the system's recommended products, and being unfavorable for improving the efficiency of users to find the products they need. Summary of the Invention
[0004] In view of this, embodiments of this application provide a user classification method, device, equipment and storage medium based on random forest to solve the problems in the existing user classification method that the classification calculation amount for users is large, the classification is inaccurate, resulting in low accuracy of the system's user classification and being unfavorable for improving the efficiency of users to find the products they need.
[0005] The first aspect of the embodiments of this application provides a user classification method based on random forest, and the method includes:
[0006] Obtain a training sample set, where the training samples in the training sample set include user characteristics and user classifications;
[0007] Cluster the training samples according to the user characteristics in the training sample set, and determine sample subsets according to the clustering results;
[0008] Construct corresponding classification trees respectively according to the determined sample subsets;
[0009] Determine a random forest according to the classification trees, and classify the users to be classified according to the random forest.
[0010] Combined with the first aspect, in the first possible implementation manner of the first aspect, clustering the training samples according to the user characteristics in the training sample set includes:
[0011] Determine the cluster centers in the training sample set;
[0012] Calculate the distance between each training sample and the cluster centers, and determine the cluster to which the training sample belongs according to the distance;
[0013] Redetermine the cluster centers according to the average coordinates of the determined clusters, and redetermine the clusters to which the training samples belong according to the determined cluster centers until the change range of the cluster centers is less than a predetermined change threshold or the number of clustering times reaches a predetermined number of times.
[0014] Combined with the first aspect, in the second possible implementation manner of the first aspect, determining the cluster centers in the training sample set includes:
[0015] Determine the number of cluster centers in the training sample set according to the number of classifications included in the training sample set.
[0016] Combined with the first aspect, in the third possible implementation manner of the first aspect, classifying the user to be classified according to the random forest includes:
[0017] Determine the cluster to which the user belongs according to the user characteristics of the user to be classified;
[0018] Search for the classification tree used to classify the user to be classified according to the determined cluster;
[0019] Calculate the classification of the user to be classified according to the searched classification tree.
[0020] Combined with the first aspect, in the fourth possible implementation manner of the first aspect, classifying the user to be classified according to the random forest includes:
[0021] Determine the first weight according to the similarity between the user to be classified and the cluster determined according to the user characteristics of the user to be classified;
[0022] Determine the classification to which the user to be classified belongs according to the classification tree of the random forest;
[0023] Fuse to obtain the classification result of the user to be classified according to the first weight and the classification result, and determine the classification of the user to be classified according to the fused classification result.
[0024] Combined with the first aspect, in the fifth possible implementation manner of the first aspect, fusing to obtain the classification result of the user to be classified according to the first weight and the classification result, and determining the classification of the user to be classified according to the fused classification result includes:
[0025] Determine the weight of the classification tree according to the first weight and the classification result;
[0026] Sum the weights of the same classification results in different classification trees to determine the weight of the classification result output by the random forest;
[0027] Determine the classification result with the largest weight as the classification of the user to be classified.
[0028] Combined with the first aspect, in the sixth possible implementation manner of the first aspect, the user features include one or more of user income information, the degree of volatility tolerated, investment preferences, investment experience, and investment term.
[0029] The second aspect of the embodiments of the present application provides a user classification device based on a random forest. The device includes:
[0030] A training sample set acquisition unit, configured to acquire a training sample set, where the training samples in the training sample set include user features and user classifications;
[0031] A clustering unit, configured to cluster the training samples according to the user features in the training sample set, and determine a sample subset according to the clustering result;
[0032] A classification tree determination unit, configured to construct corresponding classification trees according to the determined sample subsets respectively;
[0033] A user classification unit, configured to determine a random forest according to the classification tree, and classify the user to be classified according to the random forest.
[0034] The third aspect of the embodiments of the present application provides a user classification device based on a random forest, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to any one of the first aspect are implemented.
[0035] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of the first aspect are implemented.
[0036] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: The classification trees constructed from the clustered sample subsets can reduce the probability of generating similar sample subsets, reduce the construction of similar decision trees, and reduce the influence of similar decision trees on the true classification results. In addition, the purity of the clustered sample subsets is relatively high, which can effectively improve the fitting speed of the classification trees and reduce the fitting calculation amount. Moreover, constructing classification trees based on the clustered sample subsets can avoid missing some training samples when sampling to generate training subsets, thereby effectively improving the accuracy of the constructed classification trees, obtaining more accurate classification results, and improving the efficiency of users to find the products they need. Brief Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 It is a schematic diagram of the implementation process of a user classification method based on random forest provided by an embodiment of the present application;
[0039] Figure 2 It is a schematic diagram of the implementation process of a clustering method for training samples provided by an embodiment of the present application;
[0040] Figure 3 It is a schematic diagram of the user classification process based on random forest provided by an embodiment of the present application;
[0041] Figure 4 It is a schematic diagram of a user classification device based on random forest provided by an embodiment of the present application;
[0042] Figure 5 It is a schematic diagram of a user classification device based on random forest provided by an embodiment of the present application. Detailed Embodiments
[0043] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented in order to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0044] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments.
[0045] The random forest algorithm is a classifier that uses multiple trees to train and predict samples. The general process of the random forest algorithm is as follows:
[0046] 1. Determine the training samples (or also referred to as training cases). Determine that the number of training samples is N, and the total number of sample features (for the user classification scenario, the sample features are user features) is M.
[0047] 2. Input the number of features m to determine the decision result of a node on the decision tree. Among them, m is much smaller than the total number of sample features M.
[0048] 3. Sample N times with replacement from N training samples (the number of sampling times can also be different from the number of game card samples) to form a training subset. The remaining training samples that are not drawn can be used for the prediction calculation of the decision tree to evaluate the error of the decision tree.
[0049] 4. For each node, randomly select m features and determine the nodes on the decision tree through these m features. Determine the best splitting method according to the selected m features.
[0050] 5. Do not perform pruning during the decision tree generation process to allow each tree to grow completely. A completed normal tree-shaped classifier can be used for the prediction calculation of the classification result.
[0051] Determining the sample subset through sampling with replacement can improve the generalization ability of the model. However, for a small amount of data or low-dimensional data, through sampling with replacement, it may be possible to sample multiple similar sample subsets, generate multiple similar decision trees from multiple similar sample subsets, and the results calculated based on multiple similar decision trees may mask the true results. Also, through sampling with replacement, there may be unsampled samples, which will also affect the classification accuracy. In addition, the purity of the samples obtained based on sampling with replacement is not high, and the convergence speed of the fitting calculation of the decision tree is slow, which is not conducive to improving the calculation efficiency.
[0052] Based on this, an embodiment of the present application proposes a user classification method based on a random forest, as Figure 1 shown. The method includes:
[0053] In S101, obtain a training sample set, where the training samples in the training sample set include user features and user classifications.
[0054] Specifically, the training sample set in the embodiment of the present application includes training samples for training a random forest model. The random forest model can be determined by selecting a part of the training samples in the training sample set, and the other training samples can be used for the prediction calculation of the determined random forest model. If the classification result of the prediction calculation is consistent with the classification result in the training sample, it indicates that the accuracy of the random forest model meets the set requirements. If the classification result calculated by the random forest model is not consistent with the classification result in the training sample, it is necessary to further optimize the decision tree of the random forest model until the result calculated by the random forest model is consistent with the classification result of the training sample.
[0055] In a possible implementation, for users who need to purchase wealth management products, their training samples include user characteristics and classification results. The user characteristics may include one or more of the following data: the user's income information, the degree of volatility that the user can tolerate, investment preferences, investment experience, investment term, etc. The above user characteristics can be obtained through the form data filled in by the user.
[0056] The income information of the user may include the user's total income information or the income information available for investment by the user. The specific category of the income information of the user characteristics in the training sample can be determined as needed. For example, the total income amount of user A is 1 million yuan, and the amount available for investment is 100,000 yuan. The total income amount of user B is 1 million yuan per year, and the amount available for investment per year is 500,000 yuan, etc. Through the income characteristics, it can be used to match the amount characteristics of the products required by the user. For example, the minimum purchase amount of wealth management product X1 is 50,000 yuan, and the minimum purchase amount of wealth management product X2 is 1,000 yuan, etc.
[0057] The degree of volatility that can be tolerated can be represented by the maximum drawdown of the principal that can be tolerated, or it can also be determined in combination with the stop-profit data set by the user. According to the user's risk tolerance and combined with the historical volatility data of the wealth management product, determine the products within the acceptable volatility range of the user. For example, the maximum drawdown set by users A and B is 20%. According to the historical data of the product, products with a maximum drawdown of less than 20% can be selected for users A and B. When the user also sets stop-profit data, personalized recommendations can be further made according to the stop-profit data. For example, the stop-profit data of user A is 120%, and the stop-profit data of user B is 130%. Then, product X1 with better stability can be determined as the matching category for user A, and product X2 with higher income volatility can be determined as the matching category for user B.
[0058] The investment preferences may include aggressive, active, balanced, prudent, conservative, etc., and can be determined according to the questionnaire data filled in by the user.
[0059] In a possible implementation, if the user does not fill in relevant information and it may not be possible to obtain the user characteristics of the user through the form data, the user characteristics of the user can be determined according to the data of the products purchased by the user.
[0060] Specifically, according to the set of products purchased by the unknown user in the historical purchase record, find users with the same or similar purchase sets, and determine the user characteristics of the unknown user according to the found users.
[0061] For example, the historical products purchased by user A include X1, X2, X3, X4, and X5. Through similarity matching based on the user's purchase records, it is found that in the historical purchase records of user B, the purchased products are the same as those of user A, and the user characteristics of user B can be determined according to the form data filled in by user B. Then, the user characteristics of user A can be determined based on the matched user characteristics of user B.
[0062] In a possible implementation, if multiple matched users are included, the common user characteristics can be obtained from the user characteristics of the multiple matched users as the user characteristics of user A. The user characteristics with a proportion exceeding 50% of the maximum number of occurrences can be used as the user characteristics of user A. For example, among the 5 users with the same purchased products found, the user characteristic Y1 appears 3 times, the user characteristic Y2 appears 5 times, the user characteristic Y3 appears 2 times, and the user characteristic Y4 appears 1 time. The maximum number of occurrences of a single user characteristic is 5. Then, the proportion of user characteristic Y1 is 60%, the proportion of user characteristic Y2 is 100%, the proportion of user characteristic Y3 is 40%, and the proportion of user characteristic Y4 is 20%. Since the proportions of user characteristic Y1 and user characteristic Y2 are both greater than 50%, Y1 and Y2 can be determined as the user characteristics of user A.
[0063] Considering that the products purchased by the user may change over time, when determining similar users, the selected historical purchase records can be the products purchased within a predetermined period closest to the current time, so as to more accurately obtain the user characteristics of the user in the current period.
[0064] In a possible implementation, the corresponding relationship between products and user characteristics can also be set to determine the user characteristics of users with unknown characteristics. The user characteristics corresponding to the products can be found according to the products purchased by the user, or according to the products purchased by the user within a predetermined period closest to the current time, and the user characteristics of the user with unknown characteristics can be determined according to the repetition rate of the found user characteristics.
[0065] In the embodiments of the present application, the classification result of the training sample may include user types. For example, the classification result may include different types of customers. After determining the user type, the products corresponding to different types of classification results can be determined according to the pre-set corresponding relationship between customer types and products, so as to facilitate matching the required products for the user.
[0066] Alternatively, in a possible implementation, the classification result may also be the matched product. For example, the classification result of user A is product a, and the classification result of user B is product b. Of course, it is not limited to this. The classification result may also include the matched product type. For example, the product types in the system can be divided into different risk levels, including R1, R2, R3, R4, and R5, etc. According to the pre-set correspondence between product types and products, determine the products corresponding to the classification results of different product types.
[0067] The training sample set is a data set that has been correctly classified in advance. This training sample set can be determined based on historical data. For example, the satisfaction of the products purchased by users can be determined through follow-up data, and the data with a satisfaction greater than a predetermined value is extracted from the historical data as the training sample. Alternatively, according to the number of repeated purchases of users, the historical data with the number of repeated purchases greater than a predetermined number can be used as the training sample.
[0068] In S102, cluster the training samples according to the user characteristics in the training sample set, and determine the sample subsets according to the clustering results.
[0069] Clustering users according to user characteristics can be understood as clustering users with a user similarity greater than a certain value into the same sample subset. When performing clustering calculations, the similarity between two users can be calculated based on the user characteristics of the users. For example, the characteristics of the users can be quantified, and each dimension of the characteristics has a certain quantified value. For a user including N user characteristics, the position of the user in the N-dimensional feature space can be determined through the feature values of the N user characteristics. According to the distance between two users in the N-dimensional feature space, the similarity between the two users can be determined.
[0070] When clustering the training samples in the training sample set in the embodiments of the present application, it can be as Figure 2 shown, including:
[0071] In S201, determine the clustering centers in the training sample set.
[0072] When determining the cluster centers in the training sample set, the number of cluster centers can be determined according to the number of classifications included in the training sample set. Among them, the value corresponding to each cluster center can be an array composed of multiple user features. When initializing the cluster centers, multiple arrays of data can be randomly generated as the cluster centers, or any user in the training sample set can be specified as the initial cluster center. For example, assuming that the users in the training sample set include k types of users, or the training samples in the training sample set include k types of products, the number of cluster centers can be determined to be k. The positions of the k cluster centers can be specified by any k training samples in the training sample set as the cluster centers, or k cluster centers can be randomly initialized.
[0073] S22: Calculate the distance between each training sample and the cluster center, and determine the cluster to which the training sample belongs according to the distance.
[0074] For each training sample, it includes the user features of a single user and the classification to which the single user belongs. To facilitate the calculation of the distance between the training sample and the cluster center, the user features of the user can be scalarized so that the user features correspond to different values. For example, data such as the user's income information, the degree of volatility that can be tolerated, investment preferences, investment experience, investment term, etc. can determine the value corresponding to the user features according to the corresponding relationship between the preset scalar value and the user features. For example, when a user includes five user features, through quantization processing, a five-dimensional array corresponding to the user can be obtained. This five-dimensional data can determine a unique position in the five-dimensional feature space, and this position is the position of the user in the five-dimensional feature space.
[0075] Converting the user features of all training samples in the training sample set into corresponding values can determine the positions of the users of the training samples in the feature space. According to the positions of the users and the positions of the cluster centers, calculate the distances between the users and each cluster center, and according to the proximity of the distances. For example, the Euclidean distance calculation formula can be used to calculate the distances between the training samples and each cluster center.
[0076] After calculating the distances between the training samples and each cluster center, the training samples can be divided into different clusters according to the comparison results of the distances between the training samples and each cluster center. For example, the training samples are divided into the cluster center closest to the training samples, so as to obtain the clusters corresponding to multiple cluster centers.
[0077] S23: Re-determine the cluster centers according to the average coordinates of the determined clusters, and re-determine the clusters to which the training samples belong according to the determined cluster centers until the change range of the cluster centers is less than a predetermined change threshold, or the number of clustering times reaches a predetermined number.
[0078] After all users are assigned to the clusters corresponding to the cluster centers, recalculate the cluster centers. The average value of the user features of the training samples in each divided cluster can be used as the new cluster center in that cluster. After all new cluster centers are re-determined, calculate the distances between all training samples and the re-determined cluster centers, and based on the calculated distances, perform the clustering operation on the training samples again.
[0079] Through multiple clustering iterations, when the positions of the cluster centers no longer change significantly, such as the moving distance of the cluster centers is less than a predetermined distance threshold, or the number of iterations exceeds a predetermined number threshold, a sample subset composed of multiple clusters can be obtained. Each sample subset includes several training samples with relatively similar features (corresponding to the user features and classification results of users). Through clustering, the initial division operation of users can be completed, so that users with similar features are assigned to the same cluster, and a sample subset with higher purity is obtained, that is, the user features of the users in the sample subset are relatively close.
[0080] As Figure 3 shown, through the clustering operation, the training samples in the training sample set can be classified to obtain three sample subsets. Among the three obtained sample subsets, the features of the training samples in each sample subset are relatively similar (in the figure, the shape of the training sample represents the user features of the training sample, and if the shapes are relatively similar, it means that the user features of the users are relatively similar).
[0081] In S103, corresponding classification trees are constructed according to the determined sample subsets.
[0082] When constructing a classification tree based on a sample subset, for each node, the splitting feature can be selected according to the information gain criterion or the information gain ratio, and the most classification-capable user feature is selected among multiple user features. The classification threshold of the user feature can be determined by the classification results of multiple training samples. By gradually determining the user features of the classification nodes, that is, the node variables, and generating to the maximum extent without any pruning, the decision tree (or also called the classification tree) corresponding to the sample subset can be obtained.
[0083] Since the training samples included in the sample subset obtained by clustering will not be missed due to sampling operations, nor will similar sample subsets be generated due to sampling with replacement (the samples of two or more samplings are similar, and the obtained sample subsets are also similar), the classification tree constructed based on the sample subset generated by clustering can effectively avoid generating similar sample subsets and will not miss training samples, thus effectively avoiding the generation of similar classification trees for sample subsets and avoiding missing the construction of important classification trees. Therefore, a more accurate classification result can be obtained based on the random forest constructed by a more comprehensive and accurate classification tree.
[0084] In addition, the training samples of the training subset generated in the embodiments of the present application have a high similarity, that is, the purity of the training samples is high. When constructing a classification tree based on a sample subset with a high purity, the depth and complexity of the classification tree can be effectively reduced, thereby improving the fitting speed of the classification tree and the construction efficiency of the classification tree.
[0085] In S104, a random forest is determined according to the classification tree, and the user to be classified is classified according to the random forest.
[0086] Based on the multiple generated classification trees, a random forest can be formed. Based on the constructed random forest, it can be used to perform classification calculations on unclassified users, that is, users to be classified, so as to determine the user type to which the user belongs, including, for example, determining the user type of the user to be classified or determining the product type required by the user. If the determined classification result is the user type of the user to be classified, the product to be matched to the user can be determined according to the pre-set correspondence between the user type and the product. If the determined classification result is the product required by the user, the determined product can be directly provided for the user to select.
[0087] In a possible implementation, the user to be classified can be first clustered to obtain the cluster to which the user to be classified belongs. Since the classification trees in the random forest in the present application are constructed based on the clustered sample subset, when classifying the sample to be classified, the correspondence between the cluster and the classification tree can be determined according to the clustering of the sample subset when constructing the classification tree. After determining the cluster to which the user to be classified belongs, the classification tree corresponding to the cluster to which the user to be classified belongs can be selected for classification calculation according to the correspondence between the cluster and the classification tree, and the classification to which the user to be classified belongs can be obtained through the classification calculation of the classification tree. By performing classification calculations through the classification tree corresponding to the cluster to which the user to be classified belongs, the classification calculation efficiency can be effectively improved, and since the user to be classified is similar to the sample subset of the constructed classification tree, the accuracy of the classification can be effectively guaranteed.
[0088] Alternatively, the users to be classified can also be clustered to obtain the similarities between the users to be classified and different clusters, and the first weight is determined according to the similarities. Then, all classification trees of the random forest are used to perform classification calculations on the users to be classified, and the classification results calculated by different classification trees for the users to be classified are obtained. According to the classification results and the first weight, a fusion calculation is performed to obtain the classification to which the users to be classified belong. When determining the similarities between the users to be classified and the clusters, the distances between the users to be classified and the cluster centers can be calculated, and according to the corresponding relationship between the distances and the similarities, the similarities between the users to be classified and different clusters are determined. The higher the similarity, the greater the first weight of the user to be classified and the cluster, and the lower the similarity, the smaller the first weight of the user to be classified and the cluster. According to each classification tree, the classification results of the users to be classified can be obtained, and the first weights with the same classification results are fused, for example, by summing, to obtain the fusion value of the classification result. For example, according to the magnitudes of the fusion values of different classification results, the result with the largest fusion value is selected as the classification result of the random forest.
[0089] As Figure 3 shown in the figure is a schematic diagram of the implementation of a user classification method based on a random forest provided by an embodiment of the present application. Among them, for the training samples in the training sample set, after determining the sample subsets after clustering through a clustering algorithm, classification trees are respectively constructed through the sample subsets after clustering, a random forest is constructed based on the constructed classification trees, and the random forest is used to classify the data to be classified.
[0090] In a possible classification implementation method, the users to be classified can be clustered through a clustering method to obtain the clustering results of the users to be classified. A classification tree corresponding to the clustering results can be used to perform classification calculations on them. By selecting the classification tree corresponding to the users to be classified for classification calculations, the classification to which the users to be classified belong can be quickly obtained.
[0091] In a possible implementation method, the similarities between the users to be classified and each cluster can be determined. For example, the similarities between the users to be classified and the clusters can be determined by the distances between the user features of the users to be classified and the cluster centers of each cluster. The first weight is determined according to the similarities. Then, classification calculations are performed on the users to be classified according to each classification tree to determine the classification results calculated by different classification trees. Then, a fusion calculation is performed according to the first weight and the classification results to obtain the classification result of the users to be classified calculated by the random forest. For example, as Figure 3As shown in the figure, the random forest includes 3 classification trees and 3 clusters. The similarities between the user to be classified and each cluster are 40%, 35%, and 25% respectively. The classifications calculated by the three classification trees are A, B, and B respectively. Through the fusion calculation, the fusion value of the classification result of A is 40%, and the fusion value of the classification result of B is 60%. According to the magnitude of the fusion value, the classification result with the larger fusion value can be selected as the classification result output by the random forest.
[0092] In a possible implementation manner, after determining the first weight according to the similarity between the user to be classified and the cluster, the probability of the user to be classified having different classification results can also be output by the classification tree. The second weight can be determined according to the probability of the user to be classified having different classification results, and the fusion calculation is performed according to the first weight and the second weight to determine the classification result output by the random forest.
[0093] For example, Figure 3 As shown in the figure, the random forest includes 3 classification trees, and each classification tree calculates 3 classification results. Through the classification calculation of the random forest for the user to be classified, 9 classification results can be obtained. Among the results obtained by any classification tree, the types are the same as those of other classification trees. For example, the classification results all include three types: A, B, and C. After determining the first weight according to the similarity between the user to be classified and each cluster, the second weight is determined in combination with the possibility of the classification results obtained by the classification tree. The fusion process can be performed according to the first weight, the second weight, and the classification results of each classification tree to obtain the classification result of the random forest. For example, the first weights determined by the user to be classified and the three clusters are 40%, 35%, and 25% respectively. The probabilities of the three classification results of A, B, and C obtained by the first classification tree, that is, the second weights, are 60%, 20%, and 20% respectively. The probabilities of the three classification results of A, B, and C obtained by the second classification tree, that is, the second weights, are 30%, 50%, and 20% respectively. The probabilities of the three classification results of A, B, and C obtained by the third classification tree, that is, the second weights, are 80%, 9%, and 11% respectively. Through the fusion calculation, the fusion value of the classification result of A is: (60% + 30% + 80%) * 40% = 68%. The fusion value of the classification result of B is: (20% + 50% + 9%) * 35% = 27.65%. The fusion value of the classification result of C is: (20% + 20% + 11%) * 25% = 12.75%. Since the fusion value of the classification result of A is the largest, the classification result of the user to be classified determined by the random forest is of type A.
[0094] In the embodiments of the present application, a method of clustering users is used to replace the existing sampling with replacement method to generate a sample subset, so that the purity of the sample subset is higher, resulting in a lower depth and faster fitting speed when constructing a classification tree. And the clustered sample subset does not need to be resampled with replacement, so there will be no similarity problem of the classification tree, effectively improving the classification accuracy of the random forest. In addition, constructing a classification tree based on the clustered sample subset can avoid missing some training samples when sampling to generate a training subset, thus effectively improving the accuracy of the constructed classification tree.
[0095] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0096] Figure 4 The following is a schematic diagram of a user classification device based on a random forest provided by an embodiment of the present application, as Figure 4 shown, the device includes:
[0097] A training sample set acquisition unit 401, configured to acquire a training sample set, where the training samples in the training sample set include user features and user classifications;
[0098] A clustering unit 402, configured to cluster the training samples according to the user features in the training sample set, and determine a sample subset according to the clustering result;
[0099] A classification tree determination unit 403, configured to construct corresponding classification trees according to the determined sample subsets respectively;
[0100] A user classification unit 404, configured to determine a random forest according to the classification tree, and classify a user to be classified according to the random forest.
[0101] In a possible implementation manner, the clustering unit 402 includes:
[0102] A clustering center determination subunit, configured to determine a clustering center in the training sample set;
[0103] A distance clustering subunit, configured to calculate the distance between each training sample and the clustering center, and determine the cluster to which the training sample belongs according to the distance;
[0104] A loop update subunit, configured to re-determine the clustering center according to the average coordinates of the determined clusters, and re-determine the clusters to which the training samples belong according to the determined clustering center until the change range of the clustering center is less than a predetermined change threshold, or the number of clustering times reaches a predetermined number of times.
[0105] In a possible implementation, the clustering center determination subunit includes:
[0106] A quantity determination module, configured to determine the quantity of clustering centers in the training sample set according to the quantity of classifications included in the training sample set.
[0107] In a possible implementation, the user classification unit 404 includes:
[0108] A clustering subunit, configured to determine the cluster to which a user to be classified belongs according to the user characteristics of the user to be classified;
[0109] A classification tree determination subunit, configured to search for a classification tree for classifying the user to be classified according to the determined cluster;
[0110] A first classification subunit, configured to calculate the classification of the user to be classified according to the searched classification tree.
[0111] In a possible implementation, the user classification unit 404 includes:
[0112] A weight determination subunit, configured to determine a first weight for the similarity between the user to be classified and a cluster according to the user characteristics of the user to be classified;
[0113] A second classification subunit, configured to determine the classification to which the user to be classified belongs according to the classification tree of the random forest;
[0114] A third classification subunit, configured to fuse to obtain a classification result of the user to be classified according to the first weight and the classification result, and determine the classification of the user to be classified according to the fused classification result.
[0115] In a possible implementation, the third classification subunit includes:
[0116] A classification tree weight calculation module, configured to determine the weight of a classification tree according to the first weight and the classification result;
[0117] A classification result weight calculation module, configured to sum the weights with the same classification result in different classification trees to determine the weight of the classification result output by the random forest;
[0118] A classification module, configured to determine the classification result with the largest weight as the classification of the user to be classified.
[0119] In a possible implementation, the user characteristics include one or more of user income information, the degree of volatility tolerated, investment preferences, investment experience, and investment term.
[0120] Figure 4 The shown user classification device based on a random forest, and Figure 1Corresponding to the user classification method based on random forest shown.
[0121] Figure 5 It is a schematic diagram of a user classification device based on random forest provided by an embodiment of the present application. As Figure 5 shown, the user classification device 5 based on random forest in this embodiment includes: a processor 50, a memory 51, and a computer program 52 stored in the memory 51 and executable on the processor 50, such as a user classification program based on random forest. When the processor 50 executes the computer program 52, the steps in the above-mentioned various embodiments of the user classification method based on random forest are implemented. Alternatively, when the processor 50 executes the computer program 52, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0122] Exemplarily, the computer program 52 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 51 and executed by the processor 50 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 52 in the user classification device 5 based on random forest.
[0123] The user classification device 5 based on random forest can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The user classification device based on random forest may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art can understand that Figure 5 merely examples of the user classification device 5 based on random forest, which do not constitute a limitation on the user classification device 5 based on random forest. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the user classification device based on random forest may further include input / output devices, network access devices, buses, etc.
[0124] The so-called processor 50 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0125] The memory 51 may be an internal storage unit of the random forest-based user classification device 5, such as a hard disk or memory of the random forest-based user classification device 5. The memory 51 may also be an external storage device of the random forest-based user classification device 5, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the random forest-based user classification device 5. Further, the memory 51 may also include both an internal storage unit and an external storage device of the random forest-based user classification device 5. The memory 51 is used to store the computer program and other programs and data required by the random forest-based user classification device. The memory 51 may also be used to temporarily store data that has been output or will be output.
[0126] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0127] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0128] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0129] In the embodiments provided in the present application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0130] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0131] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0132] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, it can also be completed by hardware related to computer program instructions. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0133] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A user classification method based on random forest, characterized in that, the method includes: Obtain a training sample set, where the training samples in the training sample set include user features and user classifications; Cluster the training samples according to the user features in the training sample set, and determine a sample subset according to the clustering result; Construct corresponding classification trees respectively according to the determined sample subsets; Determine a random forest according to the classification trees, and classify the users to be classified according to the random forest; Among them, the classifying the users to be classified according to the random forest includes: According to the distances between the user features of the users to be classified and each clustering center, determine the similarities between the users to be classified and different clusters according to the distances, and determine the first weight according to the similarities; Determine the classification to which the users to be classified belong according to all the classification trees of the random forest; Fuse according to the first weight and the classification result to obtain the classification result of the users to be classified, and determine the classification of the users to be classified according to the fused classification result.
2. The method according to claim 1, characterized in that, Clustering the training samples according to the user features in the training sample set includes: Determine the clustering centers in the training sample set; Calculate the distances between each training sample and the clustering centers, and determine the clusters to which the training samples belong according to the distances; Redetermine the clustering centers according to the average coordinates of the determined clusters, and redetermine the clusters to which the training samples belong according to the determined clustering centers until the change range of the clustering centers is less than a predetermined change threshold or the number of clustering times reaches a predetermined number of times.
3. The method according to claim 2, characterized in that, Determining the clustering centers in the training sample set includes: Determine the number of clustering centers in the training sample set according to the number of classifications included in the training sample set.
4. The method according to claim 1, characterized in that, Classifying the users to be classified according to the random forest includes: Determine the cluster to which the users to be classified belong according to the user features of the users to be classified; Find the classification tree used to classify the users to be classified according to the determined cluster; Calculate the classification of the users to be classified according to the found classification tree.
5. The method according to claim 1, characterized in that, Fusing according to the first weight and the classification result to obtain the classification result of the users to be classified, and determining the classification of the users to be classified according to the fused classification result includes: Determine the weights of the classification trees according to the first weight and the classification results; Sum the weights with the same classification results in different classification trees to determine the weight of the classification result output by the random forest; Determine the classification result with the largest weight as the classification of the users to be classified.
6. The method according to claim 1, characterized in that, The user features include one or more of user income information, degree of volatility tolerated, investment preferences, investment experience, and investment term.
7. A user classification device based on random forest, characterized in that, the device includes: A training sample set acquisition unit for acquiring a training sample set, where the training samples in the training sample set include user features and user classifications; A clustering unit for clustering the training samples according to the user features in the training sample set and determining sample subsets according to the clustering results; A classification tree determination unit for respectively constructing corresponding classification trees according to the determined sample subsets; A user classification unit for determining a random forest according to the classification trees and classifying users to be classified according to the random forest; Among them, the user classification unit includes: A weight determination subunit for determining the similarity between the user to be classified and different clusters according to the distance between the user features of the user to be classified and each cluster center point, and determining a first weight according to the similarity; A second classification subunit for determining the classification to which the user to be classified belongs according to the classification trees of the random forest; A third classification subunit for fusing to obtain the classification result of the user to be classified according to the first weight and the classification result, and determining the classification of the user to be classified according to the fused classification result.
8. A user classification device based on a random forest, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.