User classification method and device, electronic equipment and storage medium
By acquiring historical user data and training machine learning models, a secondary default probability prediction model is generated, which solves the problem of inaccurate prediction of secondary default in existing technologies and achieves more accurate risk assessment.
Patent Information
- Application Number
- CN202510993364.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
Existing credit scoring models are unable to effectively predict the probability of secondary default by users in financial debt settlement scenarios, resulting in inaccurate risk assessment.
By acquiring the historical default records and historical user data of the users to be identified, key behaviors are determined, a machine learning model is trained as a secondary default probability prediction model, and the new user data is analyzed based on this model to generate secondary default probabilities and classifications.
It improves the accuracy of predicting the probability of secondary default, enables real-time analysis of user behavior, and enhances the effectiveness of risk assessment.
Smart Images

Figure CN120873808A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a user classification method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of my country's financial market, the scale of personal and corporate loans has continued to increase, leading to a sustained rise in the number of financial debt disputes. Creditors often request installment payments; however, due to the lack of an effective risk assessment mechanism, mediators struggle to determine the creditor's ability to fulfill their obligations, resulting in subsequent defaults. Currently, financial institutions use traditional credit scoring models to predict user defaults, such as FICO scores and logistic regression. These models often rely on static data and are ill-suited for predicting the probability of secondary defaults in current financial debt mediation scenarios. Summary of the Invention
[0003] This invention provides a user classification method, apparatus, electronic device, and storage medium to address the problem of imperfect existing user classification rules. It can predict users' secondary default behavior and improve the accuracy of secondary default probability prediction by analyzing user dynamic data.
[0004] According to one aspect of the present invention, a user classification method is provided, wherein the method includes:
[0005] Obtain the historical default records and historical user data of the user to be identified, and determine the key behaviors within the historical user data based on the historical default records;
[0006] The machine learning model trained based on the key behaviors is a quadratic default probability prediction model.
[0007] The secondary default probability of the new user data of the user to be identified is generated based on the secondary default probability prediction model and the key behaviors.
[0008] The default category of the user to be identified is determined based on the secondary default probability.
[0009] According to another aspect of the present invention, a user classification device is provided, wherein the device comprises:
[0010] The feature recognition module is used to acquire the historical default records and historical user data of the user to be identified, and to determine key behaviors based on the historical default records within the historical user data.
[0011] The model training module is used to train a machine learning model as a quadratic default probability prediction model based on the key behaviors.
[0012] The probability determination module is used to generate the secondary default probability of the new user data of the user to be identified based on the secondary default probability prediction model and the key behaviors.
[0013] The user classification module is used to determine the default classification of the user to be identified based on the secondary default probability.
[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the user classification method according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the user classification method according to any embodiment of the present invention.
[0019] The technical solution of this invention involves acquiring historical default records and historical user data of the user to be identified, determining key behaviors based on the historical default records and historical user data, training a machine learning model into a secondary default probability prediction model according to the key behaviors, processing the new user data of the user to be identified using the secondary default probability prediction model to obtain the secondary default probability, and determining the default category of the user to be identified based on the secondary default probability. This invention identifies key behaviors within historical user data based on historical default records and predicts secondary default behaviors of users based on these key behaviors, which can solve the problem of imperfect existing user secondary default classification rules. It can predict users' secondary default behaviors and improve the accuracy of secondary default probability prediction by analyzing user behavior in real time through dynamic user data.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a user classification method provided according to Embodiment 1 of the present invention;
[0023] Figure 2 This is a flowchart of another user classification method provided according to Embodiment 2 of the present invention;
[0024] Figure 3 This is a flowchart of another user classification method provided according to Embodiment 3 of the present invention;
[0025] Figure 4 This is a system framework diagram of a user classification method provided in Embodiment 4 of the present invention;
[0026] Figure 5 This is an example diagram of a model training process provided in Embodiment 4 of the present invention;
[0027] Figure 6 This is an example diagram of user classification provided in Embodiment 4 of the present invention;
[0028] Figure 7 This is a schematic diagram of the structure of a user classification device according to Embodiment 5 of the present invention;
[0029] Figure 8 This is a schematic diagram of the structure of an electronic device that implements the user classification method of this invention. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] Example 1
[0033] Figure 1 This is a flowchart of a user classification method according to Embodiment 1 of the present invention. This embodiment is applicable to the prediction of the probability of secondary default when a user has already defaulted. The method can be executed by a user classification device, which can be implemented in hardware and / or software, and can be configured in a server or server cluster. Figure 1 As shown, the method includes:
[0034] Step 110: Obtain the historical default records and historical user data of the user to be identified, and determine the key behaviors within the historical user data based on the historical default records.
[0035] The users to be identified are those requiring secondary default prediction. These users have historical default data, which can be records of defaults incurred by the users within a certain period. These historical default records can be obtained from various internal or external platforms and may include at least basic identity information, historical loan records, financial behavior data, credit scores, transaction behavior characteristics, default types, and default amounts. Historical user data can be the financial consumption data of the users to be identified over a certain period. Data sources for historical user data can include local platforms, third-party external platforms, etc. Historical user data can include user-related financial usage data, such as financial lending records, consumption frequency, and financial platform usage records.
[0036] In this embodiment of the invention, historical default records and historical user data can be collected from a user-authorized data source for the user to be identified. Record information related to historical default records can be identified within the historical user data, and this record information can be processed as key behaviors. The degree of correlation between record information within the historical user data and historical default records can be determined through methods such as cross-statistics, correlation matrices, and behavioral sequence pattern mining.
[0037] Step 120: Train a machine learning model based on key behaviors to predict the probability of default.
[0038] Among them, machine learning models can be used to train a quadratic default probability prediction model. Machine learning models can learn through key behaviors and can include XGBoost models, random forest models, and deep neural network models, etc.
[0039] In this embodiment of the invention, a pre-configured machine learning model can be obtained. Historical user behavior data corresponding to key behaviors can be input into the machine learning model as input vectors. The machine learning model can be trained through key behaviors. The trained machine learning model can be used as a secondary default probability prediction model, which can be used to identify the user's secondary default probability.
[0040] Step 130: Generate the secondary default probability of the new user data of the user to be identified based on the secondary default probability prediction model and key behaviors.
[0041] Among them, new user data can be new user data that is constantly generated by the user to be identified over time. This new user data can include financial usage data related to the user's incremental financial usage, and the use of this financial usage data can be authorized by the user.
[0042] In this embodiment of the invention, new user data of the user to be identified can be collected at fixed or variable intervals. The new user data can be statistically analyzed according to key behaviors and then input into the secondary default probability prediction model. The secondary default probability of the new user data can be predicted through the secondary default probability prediction model.
[0043] Step 140: Determine the default category of the user to be identified based on the secondary default probability.
[0044] Specifically, users can be classified according to their secondary default probability. Different default categories can be assigned based on the value of the secondary default probability. For example, when the secondary default probability is below a threshold, the user can be classified as a low-risk defaulting customer; otherwise, the user can be classified as a high-risk defaulting customer.
[0045] This invention, through obtaining historical default records and historical user data of the user to be identified, determines key behaviors based on the historical default records and historical user data. A machine learning model is then trained into a secondary default probability prediction model according to these key behaviors. The secondary default probability prediction model is then used to process the new user data of the user to be identified, yielding a secondary default probability. Based on this secondary default probability, the default category of the user to be identified is determined. This invention, by identifying key behaviors within historical user data based on historical default records and predicting secondary default behaviors based on these key behaviors, solves the problem of imperfect existing user secondary default classification rules. Furthermore, by performing real-time analysis of user behavior through dynamic user data, the accuracy of secondary default probability prediction can be improved.
[0046] Example 2
[0047] Figure 2 This is a flowchart of another user classification method provided in Embodiment 2 of the present invention. This embodiment of the present invention describes the extraction process of key behaviors. See [link to documentation]. Figure 2 The method provided in this embodiment of the invention specifically includes the following steps:
[0048] Step 210: Collect historical default records and historical user data for the user to be identified from at least one data source, wherein the data source includes at least a financial system, a mediation platform, and a consumer platform.
[0049] The data source can be a service platform that stores the data of the user to be identified. The data source can be determined by the authorization of the user to be identified. The data source can include financial systems, mediation platforms, and consumption platforms used by the user to be identified.
[0050] In this embodiment of the invention, historical default records and historical user data of the user to be identified can be collected from one or more authorized data sources. The collected historical default records and historical user data can be structured or unstructured data. Preprocessing such as missing value imputation, outlier detection, and standardization can be performed on the extracted historical default records and historical user data. It is understood that the methods for collecting historical default records and historical user data from the data source can include accessing the data source through an API interface to collect the historical default records and historical user data of the user to be identified; or collecting the historical default records and historical user data of the user to be identified from the data source through a real-time data stream; or collecting the historical default records and historical user data of the user to be identified from the data source through an ETL tool.
[0051] Step 220: Extract historical behavior sequences from historical user data and determine the degree of correlation between each historical behavior sequence and historical default records.
[0052] Among them, the historical behavior sequence can be a set of behaviors arranged in a time sequence from different user behaviors.
[0053] In this embodiment of the invention, historical user data can be extracted and split into different user behavior data according to a threshold time length. Each user behavior data point can be arranged into a historical behavior sequence according to time. Alternatively, each user behavior data point can be processed using a Word2Vec model to obtain behavior feature vectors, which can then be arranged into a historical behavior sequence according to their corresponding time. It is understood that the time range of the feature vectors or user behavior data in each group of historical behavior sequences can be the same. Correlation calculation is performed between each group of historical behavior sequences and historical default records to determine the correlation value between each group of historical behavior sequences and historical default records. Furthermore, this correlation calculation can be implemented using a deep learning embedding model or determined using an approximate nearest neighbor algorithm.
[0054] Step 230: Identify at least one key behavior in each historical behavior sequence according to the degree of correlation.
[0055] Specifically, the relevant values of each historical behavior sequence can be extracted. Historical behavior sequences with relevant values greater than a threshold can be selected. User behavior data or behavior feature vectors within historical behavior sequences can be extracted as key behaviors.
[0056] Step 240: Train a machine learning model based on key behaviors to become a secondary default probability prediction model.
[0057] In this embodiment of the invention, corresponding user behavior data can be extracted from historical user data according to key behaviors, user behavior data can be processed into feature vectors, feature vectors corresponding to each key behavior can be input into a machine learning model for training, and the trained machine learning model can be used as a secondary default probability prediction model.
[0058] Step 250: Generate the secondary default probability of the new user data of the user to be identified based on the secondary default probability prediction model and key behaviors.
[0059] Step 260: Determine the default category of the user to be identified based on the secondary default probability.
[0060] This invention, in its embodiments, extracts historical default records and historical user data from a data source to determine the correlation between historical behavior sequences and historical default records. Based on this correlation, it determines historical behavior sequences and constructs key behaviors according to these sequences. A machine learning model is then trained based on these key behaviors to obtain a secondary default probability prediction model. The secondary default probability of new user data for the target user is generated based on the key behaviors and the secondary default probability prediction model. Finally, the default categories of the target user are classified according to these secondary default probabilities. This invention improves the accuracy of key behavior identification by selecting key behaviors based on the correlation between historical behavior sequences and historical default records. It provides reliable data for training the secondary default probability prediction model based on these key behaviors, enabling the prediction of secondary default behaviors for the target user. Furthermore, real-time analysis of user behavior using dynamic user data further enhances the accuracy of secondary default probability prediction.
[0061] Furthermore, based on the above embodiments of the invention, it also includes: extracting key behaviors from historical user data according to the expert system; wherein, the key behaviors include at least one of the following: number of historical overdue payments, cash flow within a threshold period, intensity of willingness to mediate, installment amount to income ratio, communication frequency, and feedback response time.
[0062] In this embodiment of the invention, an expert system can be constructed to identify key behaviors within historical user data. The expert system uses a financial knowledge base and an inference engine to identify key behaviors that could lead to secondary defaults by the user. The identified key behaviors may include one or more of the following: number of historical overdue payments, cash flow within a threshold period, willingness to mediate, installment amount to income ratio, communication frequency, and feedback response time.
[0063] Example 3
[0064] Figure 3 This is a flowchart of another user classification method provided in Embodiment 3 of the present invention. This embodiment of the present invention describes the training and usage process of the quadratic default probability prediction model. See [link to documentation]. Figure 3 The method provided in this embodiment of the invention specifically includes the following steps:
[0065] Step 310: Obtain the historical default records and historical user data of the user to be identified, and determine the key behaviors within the historical user data based on the historical default records.
[0066] Step 320: Extract behavioral indicator data corresponding to each key behavior from the historical user data.
[0067] Among them, behavioral indicator data can be statistical quantities obtained by statistically analyzing key behaviors within historical user data. Taking key behaviors such as participation in financial mediation as an example, data such as the number of times and duration of participation in financial mediation can be extracted from historical user data as indicator data.
[0068] In this embodiment of the invention, data statistics can be performed on key behaviors within the user's historical data to obtain behavioral indicator data corresponding to each key behavior. This behavioral indicator data can be used to train a secondary default probability prediction model.
[0069] Step 330: Train the machine learning model based on the behavioral indicator data to obtain the secondary default probability prediction model.
[0070] Specifically, a dataset can be constructed based on behavioral indicator data for each key behavior. A machine learning model can be trained on this dataset. The model training process can be achieved through supervised learning training, unsupervised learning training, etc. The trained machine learning model can be used as a secondary default probability prediction model.
[0071] Step 340: Perform cross-validation and hyperparameter tuning on the secondary default probability prediction model.
[0072] Among them, cross-validation and hyperparameter tuning can be used for quadratic probability default prediction models. Cross-validation can be used to test the stability of quadratic probability default prediction models. Cross-validation can be implemented based on K-fold cross-validation or leave-one-out method, while hyperparameter tuning can be implemented through grid search.
[0073] In this embodiment of the invention, cross-validation and hyperparameter tuning can be performed on the trained quadratic default probability prediction model. For example, nested cross-validation and hyperparameter tuning can be used. Multi-layer nested optimization can be performed on the quadratic default probability prediction model. The outer layer of the nested optimization can adopt K-fold cross-validation, while the inner layer is tuned for hyperparameters based on grid search.
[0074] Step 350: Extract new user behaviors from the new user data according to key behaviors.
[0075] Specifically, new user data of users to be identified can be collected within a fixed time window. Behavioral indicator data can be extracted from the new user data based on the key behaviors that have been acquired. The behavioral indicator data extracted for each key behavior can be used as the new user behavior.
[0076] Step 360: Call the secondary default probability prediction model to process the behavior of each new user and obtain the secondary default probability.
[0077] In this embodiment of the invention, the new user behavior corresponding to each key behavior can be input into the secondary default probability prediction model. The secondary default probability prediction model can predict the secondary default probability of the user to be identified based on the new user behavior, thereby obtaining the secondary default probability output by the secondary default probability prediction model.
[0078] Step 370: Obtain the preset risk level configuration and find the corresponding risk level value in the preset risk level configuration table according to the probability value of the second default probability.
[0079] The preset risk level configuration can include risk level values used to classify user defaults and the probability value range corresponding to the risk level values. The preset risk level configuration can exist in the form of a configuration file, which can include multiple risk level values, and each risk level value can correspond to a different probability value range.
[0080] In this embodiment of the invention, a preset risk level configuration can be extracted, and the probability value range to which the probability value of a second default probability belongs can be found within the preset risk level configuration. This allows the acquisition of the risk level value corresponding to the probability value range within the preset risk level configuration. For example, the risk level values in the preset risk level configuration may include high risk, medium risk, and low risk.
[0081] Step 380: Classify the users to be identified into the default categories corresponding to the risk level values.
[0082] This invention, in its embodiments, acquires historical default records and historical user data of the user to be identified. Key behaviors are determined within the historical user data based on these records. Corresponding behavioral indicator data is extracted from the historical user data based on these key behaviors. A machine learning model is trained using these behavioral indicator data to obtain a secondary default probability prediction model. This model undergoes cross-validation and hyperparameter tuning. New user behaviors corresponding to the key behaviors are extracted from the new user data. These new user behaviors are then processed using the secondary default probability prediction model to obtain the secondary default probability. Within a preset risk level configuration, the risk level value corresponding to the secondary default probability is found. The user to be identified is then classified into the corresponding default category based on the risk level value. This invention improves the accuracy of key behavior identification by selecting key behaviors based on the correlation between historical behavior sequences and historical default records. It provides reliable data for training the secondary default probability prediction model based on these key behaviors, enabling the prediction of the user's secondary default behavior. Real-time analysis of user behavior using dynamic user data further enhances the accuracy of secondary default probability prediction.
[0083] Furthermore, based on the above embodiments of the invention, it also includes: adjusting the model parameters of the secondary default probability prediction model according to the new user data.
[0084] In this embodiment of the invention, the secondary default probability prediction model can be retrained using new user data, thereby adjusting the model parameters. When the secondary default probability prediction model is an XGBoost model, its parameters may include general parameters, Booster parameters, and learning objective parameters; when it is a random forest model, its parameters may include decision tree parameters, class weights, etc.; when it includes a deep neural network model, its parameters may include architecture parameters and training parameters.
[0085] Example 4
[0086] Figure 4 This is a system framework diagram of a user classification method according to Embodiment 4 of the present invention. See also... Figure 4 The method provided in this embodiment of the invention includes the following steps:
[0087] Step 1: Data Acquisition and Preprocessing
[0088] The system obtains multi-dimensional feature data such as users' credit information, repayment records, mediation history, and income level from financial systems, mediation platforms, and related databases, and performs preprocessing such as missing value imputation, outlier detection, and standardization on the collected feature data.
[0089] Step 2: Feature Engineering Construction
[0090] Identify key variables and construct feature indicators based on them. These key variables may include, but are not limited to: historical overdue number of times, recent cash flow, strength of willingness to resolve disputes, installment amount to income ratio, and user behavior patterns.
[0091] The aforementioned key variables can be obtained through expert systems or by performing correlation calculations between existing default records and historical user records.
[0092] Step 3: Model Training and Optimization
[0093] This invention can employ machine learning algorithms such as XGBoost, Random Forest, or Deep Neural Networks to suggest a secondary default probability prediction model for users, and improve model performance through cross-validation and hyperparameter tuning. See also Figure 5The system can use multidimensional feature data collected from users as the original dataset. Data cleaning can be performed on the original dataset, and it can be processed according to key variables to select corresponding indicator data. The indicator data can be constructed into feature values, which are then divided into training, validation, and test sets. The training set is input into the secondary default probability prediction model for training. After training, the model is fine-tuned based on the validation set, and finally evaluated using the test set. Evaluation metrics for the secondary default probability prediction model can include accuracy, recall, and F1 score. When the evaluation metrics of the secondary default probability prediction model meet the standards, the optimal model is saved. If the evaluation metrics do not meet the standards, the model can be retrained by adjusting the algorithm or features.
[0094] Step 4: Risk Level Classification and Output
[0095] The probability of a user's second default is predicted using a second default probability model. Based on the prediction results, users are categorized into different risk levels, and a visual report is generated. This report can assist mediators in specifying appropriate risk control measures. Figure 6 The risk level can be configured as high risk, medium risk, and low risk. When the probability of a second default is less than 30%, the user can be classified as low risk, and the corresponding handling strategy is to prioritize the installment plan. When the probability of a second default is greater than 30% but less than 70%, the user can be classified as medium risk, and the corresponding handling strategy is to dynamically monitor and impose installment restrictions. When the probability of a second default is greater than or equal to 70%, the user can be classified as high risk, and the corresponding handling strategy is to refuse installment payments or require a guarantee.
[0096] Step 5: Model Update and Feedback Mechanism
[0097] The secondary default probability prediction model is updated regularly using new data through an online learning mechanism, thereby ensuring the timeliness and accuracy of the secondary default probability prediction model.
[0098] The method disclosed in this invention can collect multi-source data, construct feature engineering adapted to demodulation scenarios, and use machine learning algorithms to build a secondary default probability prediction model, thereby realizing a quantitative assessment of the user's future secondary default probability, improving prediction accuracy, and achieving automated assessment of user default risk.
[0099] Example 5
[0100] Figure 7 This is a schematic diagram of a user classification device according to Embodiment 5 of the present invention. Figure 7As shown, the device includes:
[0101] The feature recognition module 410 is used to acquire the historical default records and historical user data of the user to be identified, and to determine key behaviors based on the historical default records within the historical user data.
[0102] The model training module 420 is used to train a machine learning model as a secondary default probability prediction model based on the key behaviors.
[0103] The probability determination module 430 is used to generate the secondary default probability of the new user data of the user to be identified based on the secondary default probability prediction model and the key behavior.
[0104] User classification module 440 is used to determine the default classification of the user to be identified based on the secondary default probability.
[0105] In this embodiment of the invention, a feature recognition module acquires the historical default records and historical user data of the user to be identified. A model training module determines key behaviors based on the historical default records and historical user data, and trains a machine learning model into a secondary default probability prediction model according to the key behaviors. A probability determination module calls the secondary default probability prediction model to process the new user data of the user to be identified to obtain the secondary default probability. A user classification module determines the default category of the user to be identified based on the secondary default probability. This embodiment of the invention identifies key behaviors in historical user data based on historical default records and predicts the user's secondary default behavior based on the key behaviors. This can solve the problem of imperfect existing user secondary default classification rules, and can predict the user's secondary default behavior. By analyzing user behavior in real time through dynamic user data, the accuracy of secondary default probability prediction can be improved.
[0106] In some embodiments of the invention, the feature recognition module 410 includes:
[0107] The data acquisition unit is used to collect the historical default records and historical user data of the user to be identified from at least one data source, wherein the data source includes at least a financial system, a mediation platform, and a consumer platform.
[0108] The association calculation unit is used to extract historical behavior sequences from the historical user data and determine the degree of association between each historical behavior sequence and the historical default record.
[0109] A behavior recognition unit is used to determine at least one key behavior in each of the historical behavior sequences according to the degree of association.
[0110] In some embodiments of the invention, it further includes: an expert system unit, used to identify the key behaviors within the historical user data based on the expert system; wherein the key behaviors include at least one of the following: number of historical overdue payments, cash flow within a threshold time period, intensity of willingness to mediate, installment amount to income ratio, communication frequency, and feedback response time.
[0111] In some embodiments of the invention, the model training module 420 includes:
[0112] The indicator extraction unit is used to extract behavioral indicator data corresponding to each key behavior from the historical user data.
[0113] The model training unit is used to train the machine learning model according to the behavioral indicator data to obtain the secondary default probability prediction model.
[0114] The model tuning unit is used to perform cross-validation and hyperparameter tuning on the quadratic default probability prediction model.
[0115] In some embodiments of the invention, the probability determination module 430 includes:
[0116] The data extraction unit is used to extract new user behaviors from the new user data according to the key behaviors.
[0117] The probability prediction unit is used to call the secondary default probability prediction model to process the behavior of each new user and obtain the secondary default probability.
[0118] In some embodiments of the invention, the user classification module 440 includes:
[0119] A configuration query unit is used to obtain a preset risk level configuration and look up the corresponding risk level value in the preset risk level configuration table according to the probability value of the second default probability.
[0120] The classification execution unit is used to classify the user to be identified into the default category corresponding to the risk level value.
[0121] Based on the above embodiments of the invention, it further includes: a model update module, used to adjust the model parameters of the secondary default probability prediction model according to the new user data.
[0122] The user classification device provided in this embodiment of the invention can execute the user classification method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0123] Example 6
[0124] Figure 8This is a schematic diagram of the structure of an electronic device implementing the user classification method of this invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0125] like Figure 8 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0126] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0127] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as user classification methods.
[0128] In some embodiments, the user classification method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the user classification method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the user classification method by any other suitable means (e.g., by means of firmware).
[0129] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0130] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0131] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0133] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0134] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0135] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0136] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A user classification method, characterized in that, The method includes: Obtain the historical default records and historical user data of the user to be identified, and determine the key behaviors within the historical user data based on the historical default records; The machine learning model trained based on the key behaviors is a quadratic default probability prediction model. The secondary default probability of the new user data of the user to be identified is generated based on the secondary default probability prediction model and the key behaviors. The default category of the user to be identified is determined based on the secondary default probability.
2. The method according to claim 1, characterized in that, The process of acquiring the historical default records and historical user data of the user to be identified, and determining key behaviors based on the historical default records within the historical user data, includes: The historical default records and historical user data of the user to be identified are collected from at least one data source, wherein the data source includes at least a financial system, a mediation platform, and a consumer platform; Extract historical behavior sequences from the historical user data and determine the degree of correlation between each historical behavior sequence and the historical default record; At least one key behavior is identified in each of the aforementioned historical behavior sequences based on the degree of correlation.
3. The method according to claim 2, characterized in that, Also includes: The key behaviors identified by the expert system within the historical user data include at least one of the following: number of historical overdue payments, cash flow within a threshold period, willingness to mediate, installment amount to income ratio, communication frequency, and feedback response time.
4. The method according to claim 1, characterized in that, The machine learning model trained based on the key behaviors is a quadratic default probability prediction model, including: Extract behavioral indicator data corresponding to each key behavior from the historical user data; The machine learning model is trained using the behavioral indicator data to obtain the secondary default probability prediction model. Cross-validation and hyperparameter tuning were performed on the aforementioned secondary default probability prediction model.
5. The method according to claim 1, characterized in that, The step of generating the secondary default probability of the new user data of the user to be identified based on the secondary default probability prediction model and the key behaviors includes: Extract new user behavior from the new user data according to the key behaviors described above. The secondary default probability prediction model is invoked to process the behavior of each new user, thereby obtaining the secondary default probability.
6. The method according to claim 1, characterized in that, The step of determining the default classification of the user to be identified based on the secondary default probability includes: Obtain the preset risk level configuration, and look up the corresponding risk level value in the preset risk level configuration table according to the probability value of the second default probability; The user to be identified is classified into the default category corresponding to the risk level value.
7. The method according to claim 1, characterized in that, Also includes: Adjust the model parameters of the secondary default probability prediction model based on the new user data.
8. A user classification device, characterized in that, The device includes: The feature recognition module is used to acquire the historical default records and historical user data of the user to be identified, and to determine key behaviors based on the historical default records within the historical user data. The model training module is used to train a machine learning model as a quadratic default probability prediction model based on the key behaviors. The probability determination module is used to generate the secondary default probability of the new user data of the user to be identified based on the secondary default probability prediction model and the key behaviors. The user classification module is used to determine the default classification of the user to be identified based on the secondary default probability.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the user classification method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the user classification method according to any one of claims 1-7.
Citation Information
Patent Citations
Repayment probability prediction model building method and device
CN108256691A
Customer default probability prediction method and device
CN111192140A
Financial risk identification method and device based on incremental learning, equipment and medium
CN116703539A