User classification method and device, electronic equipment and storage medium
By obtaining the data of bank card holders and using the random forest model or logistic regression model to obtain target demand parameters, the problems of low accuracy of user classification and large amount of calculation in the prior art are solved, and more efficient user classification is achieved.
Patent Information
- Application Number
- CN202311607032.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the user classification method has problems such as low classification accuracy and large calculation amount.
By obtaining the data of the target cardholder holding a bank card, obtaining the target demand parameters based on the data of the target cardholder, and classifying the target cardholder using a random forest model or a logistic regression model.
It improves the accuracy of user classification and reduces the amount of calculation required for user classification, and is suitable for push scenarios for bank card application services.
Smart Images

Figure CN120067778A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of user classification, and particularly to a user classification method, apparatus, electronic device, and storage medium. Background Art
[0002] Currently, with the development of network technology, various applications and web pages have enriched people's lives. In related technologies, in order to provide better services to users, it is necessary to classify users and perform services such as information pushing according to the user categories. However, the user classification methods in related technologies have problems of low classification accuracy and large computational complexity. Summary of the Invention
[0003] The present disclosure provides a user classification method, apparatus, electronic device, computer-readable storage medium, and computer program product to at least solve the problems of low classification accuracy and large computational complexity in the user classification methods in related technologies. The technical solutions of the present disclosure are as follows:
[0004] According to the first aspect of the embodiments of the present disclosure, a user classification method is provided, including: obtaining data of a target cardholder who holds a bank card; based on the data of the target cardholder, obtaining target demand parameters for the target cardholder to apply for a bank card again; and classifying the target cardholder based on the target demand parameters.
[0005] In an embodiment of the present disclosure, the obtaining target demand parameters for the target cardholder to apply for a bank card again based on the data of the target cardholder includes: inputting the data of the target cardholder into a target model, and obtaining the target demand parameters through the target model based on the data of the target cardholder.
[0006] In an embodiment of the present disclosure, the target model is a random forest model, and the obtaining the target demand parameters through the target model based on the data of the target cardholder includes: obtaining an i-th candidate demand parameter through the i-th decision tree in the random forest model based on the data of the target cardholder, where i is a positive integer not greater than N, and N is a positive integer; and voting on the N candidate demand parameters to screen out the target demand parameters from the N candidate demand parameters.
[0007] In an embodiment of the present disclosure, the target model is a logistic regression model, and the obtaining the target demand parameters through the target model based on the data of the target cardholder includes: obtaining the probability of the target cardholder in a candidate category through the logistic regression model based on the data of the target cardholder; and obtaining the target demand parameters based on the probability of the target cardholder in the candidate category.
[0008] In one embodiment of the present disclosure, classifying the target cardholder based on the target demand parameter includes: obtaining the mapping relationship between candidate demand parameters and candidate categories; and determining the target category of the target cardholder based on the target demand parameter and the mapping relationship.
[0009] In one embodiment of the present disclosure, the target demand parameter includes the probability of the target cardholder under a candidate category. Classifying the target cardholder based on the target demand parameter includes: if the probability of the target cardholder under the candidate category is greater than a set threshold, taking the candidate category as the target category of the target cardholder.
[0010] In one embodiment of the present disclosure, the candidate categories include a first category and a second category. Classifying the target cardholder based on the target demand parameter includes: if the probability of the target cardholder under the first category is less than the set threshold, taking the second category as the target category of the target cardholder.
[0011] In one embodiment of the present disclosure, the method further includes: obtaining training samples, where the training samples include the data of sample cardholders and the sample categories of the sample cardholders; training multiple candidate models based on the training samples; and screening out the target model from the multiple candidate models.
[0012] In one embodiment of the present disclosure, obtaining the data of the target cardholder holding a bank card includes: obtaining the correlation parameter between the candidate data categories of the data of the sample cardholder and the sample category of the sample cardholder; screening out the target data category from multiple candidate data categories based on the correlation parameter; and screening out the data of the target data category from the dataset of the target cardholder as the data of the target cardholder.
[0013] In one embodiment of the present disclosure, the method further includes: if the target category of the target cardholder is a set category, taking the target cardholder as the push object for the bank card application service.
[0014] In one embodiment of the present disclosure, the number of bank cards held by the target cardholder is 1.
[0015] According to the second aspect of the embodiments of the present disclosure, a user classification device is provided, including: a first obtaining module configured to obtain the data of the target cardholder holding a bank card; a second obtaining module configured to obtain the target demand parameter of the target cardholder for re-applying for a bank card based on the data of the target cardholder; and a classification module configured to classify the target cardholder based on the target demand parameter.
[0016] In one embodiment of the present disclosure, the second acquisition module is further configured to perform: input the data of the target cardholder into a target model, and obtain the target demand parameter through the target model based on the data of the target cardholder.
[0017] In one embodiment of the present disclosure, the target model is a random forest model, and the second acquisition module is further configured to perform: obtain the i-th candidate demand parameter through the i-th decision tree in the random forest model based on the data of the target cardholder, where i is a positive integer not greater than N, and N is a positive integer; vote on the N candidate demand parameters to screen out the target demand parameter from the N candidate demand parameters.
[0018] In one embodiment of the present disclosure, the target model is a logistic regression model, and the second acquisition module is further configured to perform: obtain the probability of the target cardholder under a candidate category through the logistic regression model based on the data of the target cardholder; obtain the target demand parameter based on the probability of the target cardholder under the candidate category.
[0019] In one embodiment of the present disclosure, the classification module is further configured to perform: obtain the mapping relationship between the candidate demand parameter and the candidate category; determine the target category of the target cardholder based on the target demand parameter and the mapping relationship.
[0020] In one embodiment of the present disclosure, the target demand parameter includes the probability of the target cardholder under the candidate category, and the classification module is further configured to perform: if the probability of the target cardholder under the candidate category is greater than a set threshold, use the candidate category as the target category of the target cardholder.
[0021] In one embodiment of the present disclosure, the candidate categories include a first category and a second category, and the classification module is further configured to perform: if the probability of the target cardholder under the first category is less than the set threshold, use the second category as the target category of the target cardholder.
[0022] In one embodiment of the present disclosure, the second acquisition module is further configured to perform: obtain training samples, where the training samples include the data of the sample cardholder and the sample category of the sample cardholder; train multiple candidate models based on the training samples; screen out the target model from the multiple candidate models.
[0023] In one embodiment of the present disclosure, the first acquisition module is further configured to perform: obtaining a correlation parameter between a candidate data category of the data of the sample cardholder and the sample category of the sample cardholder; screening out a target data category from a plurality of the candidate data categories based on the correlation parameter; and screening out the data of the target data category from the dataset of the target cardholder as the data of the target cardholder.
[0024] In one embodiment of the present disclosure, the classification module is further configured to perform: if the target category of the target cardholder is a set category, taking the target cardholder as a push object for a bank card application service.
[0025] In one embodiment of the present disclosure, the number of bank cards held by the target cardholder is 1.
[0026] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the steps of the method according to the first aspect of the embodiments of the present disclosure.
[0027] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method according to the first aspect of the embodiments of the present disclosure are implemented.
[0028] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, characterized in that when the computer program is executed by a processor of an electronic device, the steps of the method according to the first aspect of the embodiments of the present disclosure are implemented.
[0029] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: obtaining the data of the target cardholder who holds a bank card, obtaining the target demand parameter for the target cardholder to apply for a bank card again based on the data of the target cardholder, and classifying the target cardholder based on the target demand parameter. Thus, the data of the target cardholder who holds a bank card can be considered, the target demand parameter for the target cardholder to apply for a bank card again can be obtained, and the target cardholder can be classified considering the target demand parameter, improving the accuracy of user classification, and only the users who hold bank cards need to be classified, greatly reducing the calculation amount required for user classification, and being applicable to the push scenario of bank card application services.
[0030] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0032] Figure 1 is a flowchart of a user classification method shown according to an exemplary embodiment.
[0033] Figure 2 is a flowchart of a user classification method shown according to another exemplary embodiment.
[0034] Figure 3 is a flowchart of a user classification method shown according to another exemplary embodiment.
[0035] Figure 4 is a flowchart of obtaining a target model in a user classification method shown according to an exemplary embodiment.
[0036] Figure 5 is a block diagram of a user classification device shown according to an exemplary embodiment.
[0037] Figure 6 is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0038] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order different from those illustrated or described here. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0040] In the technical solutions of the present disclosure, the acquisition, storage, use, processing, etc. of data all comply with the provisions of relevant laws and regulations.
[0041] Figure 1 is a flowchart of a user classification method shown according to an exemplary embodiment. As Figure 1 shown, the user classification method of the embodiments of the present disclosure includes the following steps.
[0042] S101. Obtain data of a target cardholder who holds a bank card.
[0043] It should be noted that the execution subject of the user classification method in the embodiments of the present disclosure is an electronic device, and the electronic device includes a mobile phone, a notebook, a desktop computer, a vehicle-mounted terminal, a smart home appliance, a wearable device, etc. Among them, the wearable device may include a wrist-worn device (such as a smart watch, a smart bracelet), a head-mounted device, a foot-worn device, etc. The user classification method in the embodiments of the present disclosure may be executed by the user classification device in the embodiments of the present disclosure, and the user classification device in the embodiments of the present disclosure may be configured in any electronic device to execute the user classification method in the embodiments of the present disclosure. For example, the execution subject of the user classification method in the embodiments of the present disclosure may include an advertising system.
[0044] It should be noted that there are no excessive limitations on the bank card. For example, it may include a debit card, a credit card, a savings card, a prepaid card, etc. For example, the bank card may include a bank card issued by a financial institution (such as a bank). The target cardholder refers to the user to be classified. The target cardholder holds a bank card, and there are no excessive limitations on the target cardholder. For example, it may include an individual, an enterprise, etc.
[0045] In one implementation, users who hold bank cards issued by a set financial institution may be screened out from the user set as the target cardholders.
[0046] In one implementation, the number of bank cards held by the target cardholder is 1. It can be understood that if a user only holds 1 bank card, the user's demand for applying for a bank card again is relatively large, and the user can be used as the target cardholder to further predict the target demand parameters of the user to classify the user. On the contrary, if a user holds multiple bank cards, the user's demand for applying for a bank card again is relatively small, and there is no need to classify the user.
[0047] In one implementation, users who hold 1 bank card may be screened out from the user set as the target cardholders.
[0048] It should be noted that there are no excessive limitations on the data of the target cardholder. For example, it may include the static data, social data, behavior data, and interest data of the target cardholder.
[0049] Among them, the static data may include age, gender, location, user category (such as a new user, an old user), the bank card activation method, etc. It should be noted that the user category in the data of the target cardholder is not obtained based on the target demand parameters.
[0050] Among them, social data may include the number of contacts, the number of objects followed by the target cardholder on the target APP (Application), the number of followed objects, etc. The target APP may include an APP for pushing information to the target object. For example, it may include a video APP, a shopping APP, etc.
[0051] Among them, behavioral data may include data such as the login, browsing, transactions, collections, rewards, card binding status, and bound device status of the target cardholder. Transaction data may include the number of monthly transactions, monthly transaction amount, number of annual transactions, annual transaction amount, transaction channels, etc. The card binding status may include whether the bank card is bound to an APP (such as a payment APP), and the bound device status may include whether the bank card is bound to a device (such as a mobile phone, a computer). Transaction channels may include bank card counter transactions, ATM (Automated Teller Machine) transactions, online banking transactions, POS (Point of sales terminal) transactions, etc.
[0052] Among them, interest data may include short-term interests, long-term interests, long-tail interests, and no interests, etc.
[0053] In one implementation, obtaining data of a target cardholder holding a bank card includes obtaining the original data of the target cardholder and performing desensitization processing on the original data to obtain the data of the target cardholder. It should be noted that the desensitization processing of the original data can be implemented by any data desensitization method in related technologies, and no excessive limitation is made here.
[0054] In one implementation, obtaining data of a target cardholder holding a bank card includes obtaining the correlation parameter between the candidate data categories of the data of the sample cardholder and the sample categories of the sample cardholder. Based on the correlation parameter, screening out the target data category from multiple candidate data categories, and screening out the data of the target data category from the dataset of the target cardholder as the data of the target cardholder. Thus, the target data category can be screened out considering the correlation parameter, and the data of the target data category of the target cardholder can be used as the data of the target cardholder, which can ensure that the data of the target cardholder has a relatively high correlation with the user classification task and helps to improve the user classification accuracy.
[0055] It should be noted that the sample cardholder holds at least 1 bank card, and the sample categories may include the first category and the second category. Among them, the first category means that the user holds multiple bank cards, and the second category means that the user only holds 1 bank card.
[0056] It should be noted that the correlation parameter is used to characterize the correlation (also called the correlation relationship) between the candidate data category and the sample category. The correlation can include the degree of correlation, positive correlation, negative correlation, non-correlation, etc. There are no excessive restrictions on the correlation parameter. For example, it can include the F value of linear regression, the correlation coefficient, the non-linear correlation coefficient, the multiple correlation coefficient, etc. Among them, the F value of linear regression is used to characterize whether the linear combination of independent variables has a significant impact on the explanation of the dependent variable. Regarding the relevant content of obtaining the correlation coefficient, any correlation analysis method in the relevant technology can be used to achieve it, and there are no excessive restrictions here.
[0057] In some examples, the correlation parameter is positively correlated with the degree of correlation. Among them, the degree of correlation refers to the degree of correlation between the candidate data category and the sample category. Based on the correlation parameter, the target data category is selected from multiple candidate data categories, including sorting the multiple candidate data categories in descending order according to the correlation parameter, and taking the first M candidate data categories as the target data category. Among them, M is a positive integer.
[0058] For example, the candidate data category and the target data category are shown in Tables 1 and 2 respectively. The correlation parameter is positively correlated with the degree of correlation. Based on the correlation parameter, the 9 candidate data categories shown in Table 1 can be sorted in descending order according to the correlation parameter, and the first 5 candidate data categories are taken as the target data category.
[0059] Table 1 Candidate Data Categories
[0060]
[0061]
[0062] Table 2 Target Data Categories
[0063] Serial number Target data category Relevance parameter 1 User category 93.30 2 Card binding status 19.13 3 Bank card activation method 6.5 4 Monthly transaction times 2.18 5 Monthly transaction amount 2.03
[0064] In some examples, the correlation parameter is positively correlated with the degree of correlation. Among them, the degree of correlation refers to the degree of correlation between the candidate data category and the sample category. Based on the correlation parameter, the target data category is selected from multiple candidate data categories, including if the correlation parameter corresponding to the candidate data category is greater than the set threshold, taking the candidate data category as the target data category.
[0065] S102. Based on the data of the target cardholder, obtain the target demand parameter for the target cardholder to apply for a bank card again.
[0066] It should be noted that the target demand parameter is used to characterize the demand of the target cardholder to apply for a bank card again. There are no excessive restrictions on the target demand parameter.
[0067] In one embodiment, the target demand parameter may include the probability that the target cardholder applies for a bank card again. It can be understood that the probability that the target cardholder applies for a bank card again is positively correlated with the demand of the target cardholder to apply for a bank card again.
[0068] In one embodiment, the target demand parameter may include the target category of the target cardholder. For example, if the target category of the target cardholder is the first category, the demand of the target cardholder to apply for a bank card again is relatively large; if the target category of the target cardholder is the second category, the demand of the target cardholder to apply for a bank card again is relatively small.
[0069] In one embodiment, the target demand parameter may include the probability of the target cardholder under the candidate category. For example, the candidate categories include the first category and the second category. If the probability of the target cardholder under the first category is relatively large, the demand of the target cardholder to apply for a bank card again is relatively large; if the probability of the target cardholder under the second category is relatively large, the demand of the target cardholder to apply for a bank card again is relatively small.
[0070] In one embodiment, based on the data of the target cardholder, obtaining the target demand parameter for the target cardholder to apply for a bank card again includes obtaining the mapping relationship between the candidate interval and the candidate demand parameter, identifying the target interval in which the data of the target cardholder is located from multiple candidate intervals, and obtaining the target demand parameter based on the target interval and the mapping relationship. For example, taking the candidate demand parameter mapped by the target interval as the target demand parameter.
[0071] In one embodiment, based on the data of the target cardholder, obtaining the target demand parameter for the target cardholder to apply for a bank card again includes inputting the data of the target cardholder into the target model, and obtaining the target demand parameter through the target model based on the data of the target cardholder. Thus, the target model can be used to process the data of the target cardholder to obtain the target demand parameter. It should be noted that the target model is not overly limited. For example, it may include an RF (Random Forest) model, a logistic regression model, a deep learning model, an Xgboost (eXtreme gradient boosting) model, etc.
[0072] S103. Classify the target cardholder based on the target demand parameter.
[0073] In the embodiments of the present disclosure, classifying the target cardholder based on the target demand parameter may include the following possible implementation manners:
[0074] Method 1: Obtain the mapping relationship between candidate intervals and candidate categories, identify the target interval in which the probability of the target cardholder applying for a bank card again lies from multiple candidate intervals, and determine the target category of the target cardholder based on the target interval and the mapping relationship.
[0075] In an embodiment of the present disclosure, the target demand parameter includes the probability of the target cardholder applying for a bank card again.
[0076] In one implementation, the candidate intervals include a first interval and a second interval, and the candidate categories include a first category and a second category. Among them, the upper limit value of the first interval and the lower limit value of the second interval are both set probabilities. There is a mapping relationship between the first interval and the second category, and there is a mapping relationship between the second interval and the first category.
[0077] If the probability of the target cardholder applying for a bank card again is greater than or equal to the set probability, that is, the target interval in which the probability of the target cardholder applying for a bank card again lies is the second interval, then take the first category as the target category of the target cardholder.
[0078] If the probability of the target cardholder applying for a bank card again is less than the set probability, that is, the target interval in which the probability of the target cardholder applying for a bank card again lies is the first interval, then take the second category as the target category of the target cardholder.
[0079] Method 2: Obtain the mapping relationship between candidate demand parameters and candidate categories, and determine the target category of the target cardholder based on the target demand parameter and the mapping relationship.
[0080] For example, the candidate demand parameters include C1 and C2, the candidate categories include a first category and a second category. There is a mapping relationship between C1 and the first category, and there is a mapping relationship between C2 and the second category.
[0081] If the target demand parameter is C1, then take the first category as the target category of the target cardholder.
[0082] If the target demand parameter is C2, then take the second category as the target category of the target cardholder.
[0083] Method 3: If the probability of the target cardholder under the candidate category is greater than the set threshold, then take the candidate category as the target category of the target cardholder.
[0084] In an embodiment of the present disclosure, the target demand parameter includes the probability of the target cardholder under the candidate category.
[0085] It should be noted that there are no excessive restrictions on the set threshold. For example, it can be 0.5.
[0086] For example, if the candidate categories include Category A, B, and C, the needs of the target cardholder for re-applying for a bank card characterized by Category A, B, and C decrease in turn. Set the threshold to 0.5. If the probability of the target cardholder under Category A is greater than 0.5, Category A is taken as the target category of the target cardholder.
[0087] In one implementation, the candidate categories include a first category and a second category. Based on the target demand parameter, the target cardholder is classified. If the probability of the target cardholder under the first category is less than the set threshold, the second category is taken as the target category of the cardholder. It should be noted that there are no excessive limitations on the set threshold. For example, it can be 0.5. Thus, when the probability of the target cardholder under the first category is small, the second category can be taken as the target category of the target cardholder, which is applicable to the binary classification scenario.
[0088] Method 4: Extract the target category of the target cardholder from the target demand parameter.
[0089] In the embodiments of the present disclosure, the target demand parameter carries the target category of the target cardholder.
[0090] In one implementation, the method further includes that if the target category of the target cardholder is the set category, the target cardholder is taken as the push object of the bank card application service. The bank card application service can be pushed to the target cardholders of the set category, which helps to improve the push effect of the bank card application service.
[0091] It should be noted that there are no excessive limitations on the set category. For example, the set category can include the first category. Thus, the bank card application service can be pushed to the target cardholders of the first category. Especially when the target cardholder only holds one bank card, the need of the target cardholder for re-applying for a bank card is relatively large, and accurate push of the bank card application service can be achieved.
[0092] The user classification method provided by the embodiments of the present disclosure obtains the data of the target cardholder who holds a bank card, obtains the target demand parameter of the target cardholder for re-applying for a bank card based on the data of the target cardholder, and classifies the target cardholder based on the target demand parameter. Thus, the data of the target cardholder who holds a bank card can be considered, the target demand parameter of the target cardholder for re-applying for a bank card can be obtained, and the target cardholder can be classified considering the target demand parameter, improving the accuracy of user classification. Moreover, only the users who hold bank cards need to be classified, greatly reducing the calculation amount required for user classification, which is applicable to the push scenario of the bank card application service.
[0093] Figure 2 It is a flowchart of a user classification method shown according to another exemplary embodiment. As Figure 2 shown, the user classification method of the embodiments of the present disclosure includes the following steps.
[0094] S201. Obtain the data of the target cardholder who holds a bank card.
[0095] S202. Input the data of the target cardholder into the random forest model.
[0096] For the relevant content of steps S201 - S202, refer to the above embodiments and will not be elaborated here.
[0097] S203. Based on the data of the target cardholder, obtain the i-th candidate demand parameter through the i-th decision tree in the random forest model, where i is a positive integer not greater than N, and N is a positive integer.
[0098] S204. Vote on the N candidate demand parameters to screen out the target demand parameter from the N candidate demand parameters.
[0099] In the embodiments of the present disclosure, the target model is a random forest model, and the random forest model includes N decision trees.
[0100] It should be noted that obtaining the candidate demand parameter through the decision tree and voting on the N candidate demand parameters can be implemented by any random forest model in the related art, and no excessive limitation is made here.
[0101] In one implementation manner, the training process of the random forest model is as follows: Obtain training samples, where the training samples include the data of the sample cardholders and the sample categories of the sample cardholders, and train the random forest model based on the training samples.
[0102] In some examples, an original training set composed of Q training samples can be obtained. Each training sample includes W features of the sample cardholder and the sample category of the sample cardholder. Perform N times of random sampling with replacement on the original training set to obtain N target training sets, and each target training set is a subset of the original training set. For example, the sampling with replacement may include Bagging (BootstrappNd AggrNgation, Bootstrap Aggregation). Where Q and W are both positive integers.
[0103] Based on the i-th target training set, train the i-th decision tree. For example, randomly select E features from the W features, and select the optimal feature for node splitting according to the decision tree generation method, and no pruning operation is performed during the splitting process to generate the decision tree. The training processes of multiple decision trees are parallel.
[0104] Based on the N decision trees, obtain the random forest model.
[0105] It is understandable that the training processes of multiple decision trees in the random forest model are parallel, with a fast training speed. Moreover, the random forest model adopts an ensemble algorithm, which has better accuracy, generalization ability, and stability compared to a single decision tree. Additionally, the random forest is applicable to the processing of data such as high-dimensional data, discrete data, and continuous data, that is, it has good applicability. And the random forest model uses random sampling, which can significantly improve the overfitting problem of decision trees.
[0106] S205. Classify the target cardholders based on the target demand parameters.
[0107] For the relevant content of step S205, reference can be made to the above embodiments, and details will not be elaborated here.
[0108] In the user classification method provided by the embodiments of the present disclosure, the i-th decision tree in the random forest model is used to obtain the i-th candidate demand parameter based on the data of the target cardholders, and votes are cast on the N candidate demand parameters to screen out the target demand parameter from the N candidate demand parameters. The random forest model can be used to process the data of the target cardholders to obtain the target demand parameter. The random forest model has good accuracy, generalization ability, and stability, which helps to improve the accuracy of user classification.
[0109] Figure 3 is a flowchart of a user classification method shown according to another exemplary embodiment. As Figure 3 shown, the user classification method of the embodiments of the present disclosure includes the following steps.
[0110] S301. Obtain the data of the target cardholders who hold bank cards.
[0111] S302. Input the data of the target cardholders into the logistic regression model.
[0112] For the relevant content of steps S301 - S302, reference can be made to the above embodiments, and details will not be elaborated here.
[0113] S303. Based on the data of the target cardholders, obtain the probabilities of the target cardholders under the candidate categories through the logistic regression model.
[0114] In the embodiments of the present disclosure, the target model is the logistic regression model.
[0115] It should be noted that the probabilities of the target cardholders under the candidate categories obtained through the logistic regression model can be implemented by using any logistic regression model in the related technologies, and no further limitations are imposed here.
[0116] In one implementation, the probabilities of the target cardholders under the candidate categories obtained through the logistic regression model can be implemented by using the following formula:
[0117]
[0118]
[0119]
[0120]
[0121] Among them, is the weight vector, is the j-th weight, Z is the vector composed of the data of the target cardholder, and Z j is the j-th data of the target cardholder, is the inner product of and Z (also called dot product, scalar product), P(Y = 1|Z) is the probability of the target cardholder under the first category, P(Y = 0|Z) is the probability of the target cardholder under the second category, j is a positive integer not greater than R, and R is a positive integer.
[0122] In one implementation, the training process of the logistic regression model is as follows: Obtain training samples, where the training samples include the data of the sample cardholders and the sample categories of the sample cardholders, and based on the training samples, train the logistic regression model.
[0123] In some examples, training the logistic regression model based on the training samples includes establishing a log-likelihood function, where the independent variable of the log-likelihood function is the model parameters of the logistic regression model. Based on the training samples, obtain the maximum value of the log-likelihood function, and take the value of the independent variable under the maximum value of the log-likelihood function as the value of the final model parameters of the logistic regression model.
[0124] In some examples, training the logistic regression model based on the training samples includes obtaining the loss function of the logistic regression model based on the training samples, and training the logistic regression model based on the loss function. There is no excessive limitation on the loss function. For example, it can include CE (Cross Entropy), MSE (Mean-Square Error), KL (Kullback-Leibler) divergence, contrast loss function, etc.
[0125] S304. Obtain the target demand parameter based on the probability of the target cardholder under the candidate category.
[0126] In the embodiments of the present disclosure, obtaining the target demand parameter based on the probability of the target cardholder under the candidate category may include the following possible implementation manners:
[0127] Method 1: Take the probability of the target cardholder under the candidate category as the target demand parameter.
[0128] Method 2: Obtain the mapping relationship between candidate intervals and candidate demand parameters, identify the target interval in which the probability of the target cardholder in the candidate category lies from multiple candidate intervals, and determine the target demand parameter based on the target interval and the mapping relationship.
[0129] For example, the candidate demand parameter mapped by the target interval can be used as the target demand parameter.
[0130] For example, the candidate intervals include a third interval and a fourth interval, and the candidate demand parameters include C1 and C2. Among them, the upper limit value of the third interval and the lower limit value of the fourth interval are both set probabilities. There is a mapping relationship between the third interval and C2, and there is a mapping relationship between the fourth interval and C1.
[0131] If the probability of the target cardholder in the first category is greater than or equal to the set probability, that is, the target interval in which the probability of the target cardholder in the first category lies is the fourth interval, then C1 is used as the target demand parameter.
[0132] If the probability of the target cardholder in the first category is less than the set probability, that is, the target interval in which the probability of the target cardholder in the first category lies is the third interval, then C2 is used as the target demand parameter.
[0133] Method 3: Obtain the mapping relationship between candidate intervals and candidate categories, identify the target interval in which the probability of the target cardholder in the candidate category lies from multiple candidate intervals, and determine the target category of the target cardholder based on the target interval and the mapping relationship.
[0134] For example, the candidate category mapped by the target interval can be used as the target category.
[0135] For example, the candidate intervals include a third interval and a fourth interval, and the candidate categories include a first category and a second category. Among them, the upper limit value of the third interval and the lower limit value of the fourth interval are both set probabilities. There is a mapping relationship between the third interval and the second category, and there is a mapping relationship between the fourth interval and the first category.
[0136] If the probability of the target cardholder in the first category is greater than or equal to the set probability, that is, the target interval in which the probability of the target cardholder in the first category lies is the fourth interval, then the first category is used as the target demand parameter.
[0137] If the probability of the target cardholder in the first category is less than the set probability, that is, the target interval in which the probability of the target cardholder in the first category lies is the third interval, then the second category is used as the target demand parameter.
[0138] S305. Classify the target cardholder based on the target demand parameter.
[0139] For the relevant content of step S305, reference can be made to the above embodiments and will not be elaborated here.
[0140] The user classification method provided by the embodiments of the present disclosure uses a logistic regression model to obtain the probability of a target cardholder under candidate categories based on the data of the target cardholder, and obtains a target demand parameter based on the probability of the target cardholder under the candidate categories. The data of the target cardholder can be processed by using the logistic regression model to obtain the target demand parameter.
[0141] Based on any of the above embodiments, as Figure 4 shown, the acquisition of the target model includes the following steps.
[0142] S401. Obtain training samples, where the training samples include the data of sample cardholders and the sample categories of sample cardholders.
[0143] In one implementation, obtaining training samples includes screening out sample users who hold bank cards from a sample user set as sample cardholders, and associating the data of the sample cardholders with the sample categories of the sample cardholders to obtain training samples.
[0144] In one implementation, the method further includes if the sample category of the sample cardholder is the first category, using the training sample corresponding to the sample cardholder as the first type of sample, and if the sample category of the sample cardholder is the second category, using the training sample corresponding to the sample cardholder as the second type of sample. Thus, training samples of multiple sample categories can be collected, improving the comprehensiveness and richness of the training samples and helping to improve the model training effect.
[0145] In one implementation, the method further includes dividing a sample set composed of multiple training samples into a training set and a test set according to a set ratio. It should be noted that there are no excessive limitations on the set ratio. For example, 90% of the sample set can be divided into the training set, and 10% of the sample set can be divided into the test set.
[0146] S402. Train multiple candidate models based on the training samples.
[0147] It should be noted that the training of the candidate models can be implemented by using any model training method in related technologies, and there are no excessive limitations here. The model categories of multiple candidate models may be the same or different.
[0148] In one implementation, before training multiple candidate models based on the training samples, it further includes adjusting the hyperparameters of the candidate models. It can be understood that adjusting the hyperparameters of the candidate models can be implemented by using any hyperparameter adjustment method in related technologies, and there are no excessive limitations here. For example, the hyperparameter adjustment method may include random search, grid search, Bayesian optimization, etc.
[0149] For example, taking the candidate model as a random forest model, the hyperparameters of the random forest model include the number of trees being 86, the depth of the trees being 15, and the minimum number of samples in a leaf node being 2.
[0150] For example, taking the candidate model as a logistic regression model, the hyperparameters of the logistic regression model include the regularization term being 12, and the optimization method for the loss function is selected as L-BFGS (Limited-memory-BFGS, a limited-memory quasi-Newton method).
[0151] S403. Select a target model from multiple candidate models.
[0152] In one implementation, selecting a target model from multiple candidate models includes testing the candidate models based on a test set to obtain test results, and selecting a target model from multiple candidate models based on the test results.
[0153] It should be noted that no excessive restrictions are placed on the test results. For example, it may include indicators such as accuracy (Precision), recall (Recall), F-Measure (also called F-Score), root mean square error, etc. It should be noted that F-Measure is the weighted harmonic mean of accuracy and recall.
[0154] For example, the candidate models include a random forest model and a logistic regression model, and the test results are shown in Table 3.
[0155] Table 3 Test Results
[0156] Candidate model Accuracy rate Recall rate F-Measure Root mean square error Random forest model 0.67 0.7 0.63 0.37 Logistic regression model 0.45 0.6 0.46 0.47
[0157] As shown in Table 3, compared with the logistic regression model, the random forest model has larger accuracy, recall, and F-Measure, and a smaller root mean square error, indicating that the random forest model has better prediction performance and stability, and the random forest model can be used as the target model.
[0158] Thus, in this method, training samples are obtained, where the training samples include data of sample cardholders and sample categories of sample cardholders. Based on the training samples, multiple candidate models are trained, and a target model is selected from multiple candidate models, enabling the acquisition of the target model.
[0159] Figure 5 It is a block diagram of a user classification device shown according to an exemplary embodiment. Refer to Figure 5 , the user classification device 100 of the embodiments of the present disclosure includes: a first acquisition module 110, a second acquisition module 120, and a classification module 130.
[0160] The first acquisition module 110 is configured to acquire data of a target cardholder who holds a bank card;
[0161] The second acquisition module 120 is configured to acquire target demand parameters for the target cardholder to reapply for a bank card based on the data of the target cardholder;
[0162] The classification module 130 is configured to classify the target cardholder based on the target demand parameters.
[0163] In an embodiment of the present disclosure, the second acquisition module 120 is further configured to: input the data of the target cardholder into a target model, and obtain the target demand parameters through the target model based on the data of the target cardholder.
[0164] In an embodiment of the present disclosure, the target model is a random forest model, and the second acquisition module 120 is further configured to: obtain an i-th candidate demand parameter through the i-th decision tree in the random forest model based on the data of the target cardholder, where i is a positive integer not greater than N, and N is a positive integer; vote on the N candidate demand parameters to screen out the target demand parameters from the N candidate demand parameters.
[0165] In an embodiment of the present disclosure, the target model is a logistic regression model, and the second acquisition module 120 is further configured to: obtain the probability of the target cardholder under a candidate category through the logistic regression model based on the data of the target cardholder; obtain the target demand parameters based on the probability of the target cardholder under the candidate category.
[0166] In an embodiment of the present disclosure, the classification module 130 is further configured to: obtain the mapping relationship between the candidate demand parameters and the candidate categories; determine the target category of the target cardholder based on the target demand parameters and the mapping relationship.
[0167] In an embodiment of the present disclosure, the target demand parameters include the probability of the target cardholder under the candidate category, and the classification module 130 is further configured to: if the probability of the target cardholder under the candidate category is greater than a set threshold, use the candidate category as the target category of the target cardholder.
[0168] In an embodiment of the present disclosure, the candidate categories include a first category and a second category, and the classification module 130 is further configured to: if the probability of the target cardholder under the first category is less than the set threshold, use the second category as the target category of the target cardholder.
[0169] In one embodiment of the present disclosure, the second acquisition module 120 is further configured to perform: acquiring training samples, where the training samples include data of a sample cardholder and a sample category of the sample cardholder; training a plurality of candidate models based on the training samples; and screening out the target model from the plurality of candidate models.
[0170] In one embodiment of the present disclosure, the first acquisition module 110 is further configured to perform: acquiring a correlation parameter between a candidate data category of the data of the sample cardholder and the sample category of the sample cardholder; screening out a target data category from the plurality of candidate data categories based on the correlation parameter; and screening out data of the target data category from the dataset of the target cardholder as the data of the target cardholder.
[0171] In one embodiment of the present disclosure, the classification module 130 is further configured to perform: if the target category of the target cardholder is a set category, taking the target cardholder as a push object for a bank card application service.
[0172] In one embodiment of the present disclosure, the number of bank cards held by the target cardholder is 1.
[0173] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0174] The user classification device provided by the embodiments of the present disclosure acquires data of a target cardholder who holds a bank card, obtains a target demand parameter for the target cardholder to apply for a bank card again based on the data of the target cardholder, and classifies the target cardholder based on the target demand parameter. Thus, it is possible to consider the data of the target cardholder who holds a bank card, obtain the target demand parameter for the target cardholder to apply for a bank card again, and classify the target cardholder considering the target demand parameter, improving the accuracy of user classification, and only classifying users who hold bank cards, greatly reducing the computational amount required for user classification, and being applicable to the push scenario of bank card application services.
[0175] Figure 6 It is a block diagram of an electronic device shown according to an exemplary embodiment.
[0176] As Figure 6 shown, the above electronic device 200 includes:
[0177] A memory 210, a processor 220, and a bus 230 connecting different components (including the memory 210 and the processor 220). The memory 210 stores a computer program, and when the processor 220 executes the program, the user classification method described in the embodiments of the present disclosure is implemented.
[0178] The bus 230 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of the various bus architectures. By way of example, and not limitation, these architectures include the Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0179] The electronic device 200 typically includes a variety of electronic device-readable media. These media can be any available media that can be accessed by the electronic device 200, including both volatile and nonvolatile media, removable and non-removable media.
[0180] The memory 210 may also include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 240 and / or cache memory 250. The electronic device 200 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, a storage system 260 may be used for reading from and writing to non-removable, nonvolatile magnetic media ( Figure 6 not shown and typically called a "hard disk drive"). Although Figure 6 not shown in the figures, a disk drive for reading from and writing to a removable nonvolatile disk (e.g., a "floppy disk") and an optical disk drive for reading from and writing to a removable nonvolatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these instances, each drive may be connected to the bus 230 by one or more data media interfaces. The memory 210 may include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of the embodiments of the present disclosure.
[0181] A program / utility 280 having a set (at least one) of program modules 270 may be stored, for example, in the memory 210. Such program modules 270 include - but are not limited to - an operating system, one or more application programs, other program modules, and program data, each of which examples or some combination thereof may include an implementation of a network environment. The program modules 270 typically carry out the functions and / or methods of the embodiments described herein.
[0182] The electronic device 200 can also communicate with one or more external devices 290 (such as a keyboard, a pointing device, a display 291, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 200, and / or communicate with any device that enables the electronic device 200 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 292. Moreover, the electronic device 200 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 293. As Figure 6 shown, the network adapter 293 communicates with other modules of the electronic device 200 through a bus 230. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0183] The processor 220 executes various functional applications and data processing by running programs stored in the memory 210.
[0184] It should be noted that for the implementation process and technical principle of the electronic device in this embodiment, refer to the foregoing explanation of the user classification method of the embodiments of the present disclosure, and details are not described herein again.
[0185] The electronic device provided by the embodiments of the present disclosure can execute the user classification method as described above, obtain data of a target cardholder who holds a bank card, obtain target demand parameters for the target cardholder to apply for a bank card again based on the data of the target cardholder, and classify the target cardholder based on the target demand parameters. Thus, it is possible to consider the data of the target cardholder who holds a bank card, obtain the target demand parameters for the target cardholder to apply for a bank card again, and classify the target cardholder considering the target demand parameters, improving the accuracy of user classification, and only classifying users who hold bank cards, greatly reducing the computational amount required for user classification, and being applicable to the push scenario of bank card application services.
[0186] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the user classification method provided by the present disclosure are implemented.
[0187] Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0188] To implement the above embodiments, the present disclosure also provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor of an electronic device, it implements the user classification method as described above.
[0189] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0190] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A user classification method, characterized in that, comprising: obtaining data of a target cardholder who holds a bank card; based on the data of the target cardholder, obtaining target demand parameters for the target cardholder to apply for a bank card again; classifying the target cardholder based on the target demand parameters.
2. The method according to claim 1, characterized in that, the obtaining target demand parameters for the target cardholder to apply for a bank card again based on the data of the target cardholder includes: inputting the data of the target cardholder into a target model, and obtaining the target demand parameters through the target model based on the data of the target cardholder.
3. The method according to claim 2, characterized in that, the target model is a random forest model, and the obtaining the target demand parameters through the target model based on the data of the target cardholder includes: obtaining an i-th candidate demand parameter through the i-th decision tree in the random forest model based on the data of the target cardholder, where i is a positive integer not greater than N, and N is a positive integer; voting on the N candidate demand parameters to screen out the target demand parameters from the N candidate demand parameters.
4. The method according to claim 2, characterized in that, the target model is a logistic regression model, and the obtaining the target demand parameters through the target model based on the data of the target cardholder includes: obtaining the probability of the target cardholder under a candidate category through the logistic regression model based on the data of the target cardholder; obtaining the target demand parameters based on the probability of the target cardholder under the candidate category.
5. The method according to claim 1, characterized in that, the classifying the target cardholder based on the target demand parameters includes: obtaining a mapping relationship between candidate demand parameters and candidate categories; determining the target category of the target cardholder based on the target demand parameters and the mapping relationship.
6. The method according to claim 1, characterized in that, the target demand parameters include the probability of the target cardholder under a candidate category, and the classifying the target cardholder based on the target demand parameters includes: if the probability of the target cardholder under the candidate category is greater than a set threshold, taking the candidate category as the target category of the target cardholder.
7. The method according to claim 6, characterized in that, the candidate categories include a first category and a second category, and the classifying the target cardholder based on the target demand parameters includes: if the probability of the target cardholder under the first category is less than the set threshold, taking the second category as the target category of the target cardholder.
8. The method according to claim 2, characterized in that, the method further includes: obtaining training samples, where the training samples include data of sample cardholders and sample categories of the sample cardholders; training multiple candidate models based on the training samples; screening out the target model from the multiple candidate models.
9. The method according to any one of claims 1-8, wherein, the obtaining of data of the target cardholder holding a bank card includes: obtaining a correlation parameter between a candidate data category of data of a sample cardholder and the sample category of the sample cardholder; screening out a target data category from a plurality of the candidate data categories based on the correlation parameter; screening out data of the target data category from the dataset of the target cardholder as the data of the target cardholder.
10. The method according to any one of claims 1-8, wherein, the method further includes: if the target category of the target cardholder is a set category, taking the target cardholder as a push object for a bank card application service.
11. The method according to any one of claims 1-8, wherein, the number of bank cards held by the target cardholder is 1.
12. A user classification device, wherein, it includes: a first obtaining module configured to execute obtaining data of a target cardholder holding a bank card; a second obtaining module configured to execute obtaining a target requirement parameter for the target cardholder to re-apply for a bank card based on the data of the target cardholder; a classification module configured to execute classifying the target cardholder based on the target requirement parameter.
13. An electronic device, wherein, it includes: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to: implement the steps of the method according to any one of claims 1-11.
14. A computer-readable storage medium, on which computer program instructions are stored, wherein, when the program instructions are executed by a processor, the steps of the method according to any one of claims 1-11 are implemented.