Credit Risk Identification Method, Device, Equipment and Computer Readable Storage Medium
By generating user behavior sequences, performing natural language processing and sparse data processing, using recurrent neural network model to extract behavior sequence list data, and combining structured data for fusion training, the problems of missing information and sparse data in credit risk identification are solved, and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202310025119.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-01-09
AI Technical Summary
The existing credit risk identification methods have problems such as missing information, sparse data and neglected interactions between behavioral behaviors in dealing with borrowers, resulting in low recognition accuracy.
By obtaining the user's historical behavior data, generating behavior sequences, performing natural language processing and sparse data processing, using the recurrent neural network model to extract behavior sequence list data, combining structured data for fusion training, and identifying credit risks.
It improves the accuracy of credit risk identification, captures more hidden and valuable information, reduces data sparseness, and improves model prediction performance.
Smart Images

Figure CN115983982B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of financial technology (Fintech), and in particular to a credit risk identification method, device, equipment and computer-readable storage medium. Background Art
[0002] In the financial credit industry, investors often need to assess whether a client is at risk of default or fraud, and based on this decide whether to lend to the client.
[0003] At present, in credit scenarios, for borrowers' behavior sequences, such as account opening, loan inquiries, historical repayment records and other information dimensions with time relationships, the most common way to identify credit risks is to use a fixed time window for feature extraction for each information dimension, and process it as structured data for subsequent modeling. In other words, the current credit risk identification method is based on a single-dimensional behavioral time series, which is derived into structured features based on the time window and aggregation function, and then input into the decision tree and ensemble tree models for interaction of different features.
[0004] However, the current credit risk identification method, which processes the borrower's behavior sequence as a structured feature, often faces the following problems:
[0005] The first is information loss. Commonly used aggregation methods, such as average, ratio, standard deviation, etc., are all aggregation processing of original information. These operations will inevitably lose some of the most original data information.
[0006] Second, the data is highly sparse. When processing unstructured behavior sequences into structured information, due to the huge differences in the behavior sequences of different borrowers, the final aggregated features will have a high percentage of missing information.
[0007] Third, it ignores the interaction between behaviors and the order of behaviors. For example, two customers have borrowed and repaid three times in the past six months. One of them borrowed three times and then repaid, while the other needed to borrow again before each repayment. From this perspective, the credit risk of the latter is obviously higher than that of the former. However, if we use the traditional feature derivation method, we will not be able to identify the behavioral differences between these two customers.
[0008] Therefore, how to overcome the above problems and improve the accuracy of credit risk identification has become a technical problem that needs to be solved urgently in the field of financial credit. Summary of the invention
[0009] The main purpose of this application is to provide a credit risk identification method, device, equipment and computer-readable storage medium, aiming to improve the accuracy of credit risk identification.
[0010] To achieve the above object, the present application provides a credit risk identification method, which includes:
[0011] Obtain historical behavior data of a user at different time nodes, sort and splice the historical behavior data at each time node in chronological order to generate a user behavior sequence, where the historical behavior data is unstructured data representing historical credit behaviors, and the user behavior sequence includes behavior features and chronological information of each behavior feature;
[0012] Perform natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector, and perform Embedding sparse data processing on the first behavior sequence encoding vector to obtain a low-dimensional dense second behavior sequence encoding vector;
[0013] Input the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence characterization data;
[0014] Identify the credit risk of the user according to the behavior sequence characterization data.
[0015] In some embodiments, the step of performing natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector includes:
[0016] According to a preset behavior action dictionary, map each behavior feature in the user behavior sequence to a behavior code to obtain a behavior code sequence;
[0017] Perform One-hot encoding on the behavior code sequence to obtain a first behavior sequence encoding vector.
[0018] In some embodiments, the step of identifying the credit risk of the user according to the behavior sequence characterization data includes:
[0019] Obtain the structural characterization data of the user, where the structural characterization data is structured data representing user attribute features;
[0020] Through a preset classification algorithm, fuse the structural characterization data and the behavior sequence characterization data to obtain fused characterization data, and jointly train a credit risk model;
[0021] Input the fused characterization data into the trained credit risk model, predict the default probability of the user, and identify the credit risk of the user according to the default probability.
[0022] In some embodiments, before the step of inputting the fused characterization data into the trained credit risk model, the method further includes:
[0023] Obtain the sample data of the behavior sequence and the sample data of the structure, and fuse the sample data of the behavior sequence and the sample data of the structure to obtain a training set and a validation set, wherein the sample data of the behavior sequence is the sample data representing historical credit behaviors, and the sample data of the structure is the sample data representing user attribute characteristics;
[0024] Based on the training set, iteratively train the credit risk model, and evaluate the effect of the credit risk model according to the validation set to obtain a determination result;
[0025] If the determination result does not meet the preset standard, continue to iteratively train the credit risk model;
[0026] If the determination result meets the preset standard, end the iterative training to obtain a trained credit risk model.
[0027] In some embodiments, the step of fusing the sample data of the behavior sequence and the sample data of the structure to obtain a training set and a validation set includes:
[0028] Fuse the sample data of the behavior sequence and the sample data of the structure through a preset classification algorithm to obtain a fused sample set, wherein the fused sample set includes a plurality of training samples and the default labels associated with each training sample;
[0029] Divide the fused sample set into a training set and a validation set according to a preset ratio.
[0030] In some embodiments, the preset classification algorithm is a logistic regression algorithm.
[0031] In some embodiments, before the step of inputting the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain the behavior sequence representation data, the method further includes:
[0032] Train the recurrent neural network model based on the long short-term memory network (LSTM) algorithm.
[0033] In addition, the present application provides a credit risk identification device, and the credit risk identification device includes:
[0034] A behavior sequence extraction module, configured to obtain the historical behavior data of a user at different time nodes, sort and splice the historical behavior data of each time node in time sequence to generate a user behavior sequence, wherein the historical behavior data is unstructured data representing historical credit behaviors, and the user behavior sequence includes behavior features and the time sequence information of each behavior feature;
[0035] An unstructured data processing module is used to perform natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector, and perform Embedding sparse data processing on the first behavior sequence encoding vector to obtain a low-dimensional dense second behavior sequence encoding vector;
[0036] A behavior sequence characterization module is used to input the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence characterization data;
[0037] A credit risk identification module is used to identify the credit risk of the user according to the behavior sequence characterization data.
[0038] In addition, to achieve the above object, the present application also provides a credit risk identification device. The credit risk identification device includes: a memory, a processor, and a credit risk identification program stored on the memory and executable on the processor. When the credit risk identification program is executed by the processor, the steps of the credit risk identification method as described above are implemented.
[0039] The present application also provides a computer-readable storage medium. A credit risk identification program is stored on the computer-readable storage medium. When the credit risk identification program is executed by the processor, the steps of the credit risk identification method as described above are implemented.
[0040] The technical solution of the present application is to obtain historical behavior data of a user at different time nodes, sort and splice the historical behavior data of each time node in time sequence to generate a user behavior sequence. Among them, the historical behavior data is unstructured data representing historical credit behaviors, and the user behavior sequence includes behavior features and the time sequence information of each behavior feature; perform natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector, and perform Embedding sparse data processing on the first behavior sequence encoding vector to obtain a low-dimensional dense second behavior sequence encoding vector; input the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence characterization data; identify the credit risk of the user according to the behavior sequence characterization data.
[0041] That is to say, the present application extracts the original behavior sequence, and uses technologies such as natural language processing and sparse data processing based on unstructured data to transform the unstructured data as the input of the recurrent neural network model, so as to extract the behavior sequence characterization of the key behavior time sequence, fully capture the behavior sequence relationship information of the user in the past time, make the feature granularity representing the user's historical credit behavior finer and the information expression more accurate, facilitate more accurately predicting the credit default probability of the user, and thus achieve a better prediction effect of the credit risk model, improving the accuracy of credit risk identification.
[0042] The current credit risk identification method is based on a single-dimensional behavioral time series, and structured features are derived according to a time window and an aggregation function and input into models such as decision trees and ensemble trees for interaction of different features.
[0043] Compared with this prior art, in this application, by collecting the original behavioral data of borrowers (i.e., historical behavioral data as unstructured features), and through the use of technologies such as natural language processing and sparse data processing to transform the original behavioral data, thus avoiding the use of conventional aggregation methods (such as average, ratio, or standard deviation, etc.) to aggregate the original behavioral information, avoiding the loss of some of the most original data information, and further making the feature granularity representing the user's historical credit behavior finer and the information expression more accurate, facilitating a more accurate prediction of the user's credit default probability. That is to say, this application transforms from the traditional extraction of structured features from behavioral sequences to directly using unstructured data such as action texts and action flows for modeling. Since more original data is utilized, it is necessary to combine natural language processing technology and specific model application solutions to solve problems such as unstructured data processing, sparse data transformation, and deep sequence feature mining, so as to capture more hidden and valuable information, thereby achieving better performance in predicting the credit default probability.
[0044] In addition, in this application, the Embedding sparse data processing technology is used to reduce the dimension of the first behavioral sequence coding vector as high-dimensional sparse data, obtaining a low-dimensional dense second behavioral sequence coding vector, and then using this second behavioral sequence coding vector as the input of the subsequent recurrent neural network model, thereby effectively reducing the data sparsity degree, and further reducing the computational complexity of the recurrent neural network model to extract behavioral sequence representations based on the behavioral sequence coding vector, and further improving the accuracy of predicting the user's credit default probability.
[0045] Furthermore, this application uses the action sequence of the historical behavior of borrowers to create a complete behavioral sequence, considering the order of different behaviors from the overall perspective of the customer rather than a single dimension. The main purpose is to construct a credit default model without losing the information of the original behavioral sequence, fully capturing the interaction between behaviors and the sequential information of behaviors, thereby achieving better performance in predicting the credit default probability, and further achieving the technical purpose of improving the accuracy of credit risk identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0047] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0048] Figure 1 It is a schematic diagram of the implementation process of the first embodiment of the credit risk identification method of the present application;
[0049] Figure 2 It is a schematic diagram of the behavior coding mapping of an embodiment of the present application;
[0050] Figure 3 It is a schematic diagram of performing one-hot coding in an embodiment of the present application;
[0051] Figure 4 It is a schematic diagram of the operation of performing Embedding sparse data processing in an embodiment of the present application;
[0052] Figure 5 It is a schematic diagram of the structure of the LSTM model in the embodiment of the present application;
[0053] Figure 6 It is a schematic diagram of the logical processing of the hidden layer vector in the embodiment of the present application;
[0054] Figure 7 It is a schematic diagram of the structure of the credit risk identification device for the device hardware operating environment involved in the embodiment of the present application;
[0055] Figure 8 It is a schematic diagram of the functional modules of the credit risk identification device of the present application.
[0056] The realization, functional features and advantages of the purpose of the present application will be further described with reference to the accompanying drawings in combination with the embodiments. Detailed implementation manners
[0057] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0058] The main solution of the embodiment of the present invention is as follows: Obtain the historical behavior data of the user at different time nodes, sort and splice the historical behavior data of each time node in chronological order to generate a user behavior sequence. Among them, the historical behavior data is unstructured data representing historical credit behaviors, and the user behavior sequence includes behavior characteristics and the chronological information of each behavior characteristic; perform natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector, and perform Embedding sparse data processing on the first behavior sequence encoding vector to obtain a low-dimensional dense second behavior sequence encoding vector; input the second behavior sequence encoding vector into a pre-constructed recurrent neural network model, and decode to obtain behavior sequence characterization data; identify the credit risk of the user according to the behavior sequence characterization data.
[0059] That is to say, this application extracts the original behavior sequence, and uses technologies such as natural language processing and sparse data processing based on unstructured data to transform the unstructured data, which is used as the input of the recurrent neural network model, so as to extract the behavior sequence characterization of the key behavior time series, fully capture the sequential relationship information of the user's behavior in the past, make the feature granularity representing the user's historical credit behavior finer and the information expression more accurate, facilitate more accurately predicting the user's credit default probability, and thus achieve a better prediction effect of the credit risk model and improve the accuracy of credit risk identification.
[0060] The current credit risk identification method is based on a single-dimensional behavior time series, and structured features are derived according to a time window and an aggregation function and input into models such as decision trees and ensemble trees for interaction of different features.
[0061] Compared with this prior art, this application example collects the original behavior data of the borrower (i.e., historical behavior data as unstructured features), and uses technologies such as natural language processing and sparse data processing to transform the original behavior data, thus avoiding using conventional aggregation methods (such as average, ratio or standard deviation, etc.) to aggregate the original behavior information, avoiding loss of some of the most original data information, and thus making the feature granularity representing the user's historical credit behavior finer and the information expression more accurate, facilitating more accurately predicting the user's credit default probability. That is to say, this application changes from extracting structured features from the behavior sequence in the traditional way to directly using unstructured data such as action texts and action flows for modeling. Since more original data is used, it is necessary to combine natural language processing technology and a specific model application solution to solve the problems of unstructured data processing, sparse data transformation and deep sequence feature mining, so as to capture more hidden and valuable information, and thus achieve better credit default probability prediction performance.
[0062] In addition, this application uses the Embedding sparse data processing technology to reduce the dimension of the first behavior sequence encoding vector, which is high-dimensional sparse data, to obtain a low-dimensional dense second behavior sequence encoding vector, and then uses this second behavior sequence encoding vector as the input of the subsequent recurrent neural network model, thus effectively reducing the data sparsity degree, further reducing the computational complexity of the recurrent neural network model to extract the behavior sequence representation based on the behavior sequence encoding vector, and further improving the accuracy of predicting the credit default probability of the user.
[0063] Furthermore, this application uses the action sequence of the borrower's historical behavior to create a complete behavior sequence. From the perspective of the overall customer rather than a single dimension, it considers the sequence order between different behaviors. The main purpose is to construct a credit default model without losing the information of the original behavior sequence, fully capture the interaction between behaviors and the sequence information of behaviors, so as to achieve better performance in predicting the credit default probability, and further achieve the technical purpose of improving the accuracy of credit risk identification.
[0064] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0065] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations.
[0066] Structured data: Structured data, also known as row data, is data logically expressed and implemented by a two-dimensional table structure, strictly following the data format and length specifications, and is mainly stored and managed through a relational database.
[0067] Unstructured data: Data with irregular or incomplete data structures, without a predefined data model, and inconvenient to be represented by a database two-dimensional logical table. It includes all formats of office documents, texts, pictures, HTML, various reports, images, and audio / video information, etc. In this patent, it mainly refers to text-based data.
[0068] Recurrent Neural Network (RNN): A recurrent neural network is a type of recursive neural network that takes sequence data as input, recurses in the evolution direction of the sequence, and all nodes (i.e., recurrent units) are connected in a chain.
[0069] Logistic Regression: Logistic regression is a generalized linear model used to solve binary classification problems. It assumes that the dependent variable belongs to the Bernoulli distribution and selects the sigmoid function as the link function. It is a supervised machine learning method.
[0070] One-Hot Encoding: Also known as one-hot encoding, it mainly uses an N-bit status register to encode N states. Each state has its own independent register bit, and only one bit is valid at any time. One-Hot encoding is the representation of categorical variables as binary vectors. This first requires mapping the categorical values to integer values. Then, each integer value is represented as a binary vector, which is all zero values except for the index of the integer, which is marked as 1.
[0071] Embedding: Represents a word with a low-dimensional vector. The property of this embedding vector is that objects corresponding to vectors with close distances have similar meanings.
[0072] word2vector (W2V): Each word is represented as a fixed-length vector, and these vectors can better express the similarity and analogy relationships between different words.
[0073] Credit Risk: Credit risk refers to the risk that the counterparty fails to fulfill its due debts. Credit risk is also known as default risk, which refers to the possibility that borrowers, security issuers, or counterparties, due to various reasons, are unwilling or unable to fulfill the contract terms and constitute a default, resulting in losses to banks, investors, or counterparties.
[0074] Unless otherwise defined, all technical and scientific terms used in this article have the same meaning as commonly understood by those skilled in the technical field to which this invention belongs. The terms used in this article are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0075] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.
[0076] It should be noted that in this embodiment, the current credit risk identification method is based on a single-dimensional behavioral time series, and structured features are derived according to a time window and an aggregation function and input into models such as decision trees and ensemble trees for interaction of different features. However, the current credit risk identification method often faces the following problems when processing the borrower's behavior sequence into structured features:
[0077] First, information loss. Common aggregation methods, such as average, ratio, standard deviation, etc., all perform aggregation processing on the original information, and these operations will inevitably lose a part of the most original data information.
[0078] Second, the data sparsity is high. When processing unstructured behavior sequences into structured information, due to the huge differences in the behavior sequences of different borrowers, the resulting aggregated features will have a high missing proportion. For example, only 1% of borrowers have credit card repayment records in the month before applying for a loan, and other borrowers do not have such records. Then, when constructing the feature of normal credit card repayment in the month before borrowing, the missing value proportion is 99%.
[0079] Third, the interaction between behaviors and the sequential information of behaviors are ignored. For example, both of two customers have 3 borrowings and 3 repayments in the recent 6 months. One borrows 3 times first and then repays, and the other needs to borrow again before each repayment. Considering from this dimension, the credit risk of the latter is significantly higher than that of the former. However, if according to the traditional feature derivation method, the behavioral differences between these two customers cannot be identified.
[0080] Based on this, in order to overcome the above problems and improve the accuracy of credit risk identification, various embodiments of the credit risk identification method of this application are proposed. Please refer to Figure 1 , Figure 1 which is the schematic diagram of the implementation step process of the first embodiment of the credit risk identification method of this application. In this embodiment, the credit risk identification method of this application may include:
[0081] Step S100, obtain the historical behavior data of the user at different time nodes, sort and splice the historical behavior data at each time node in time sequence to generate a user behavior sequence;
[0082] In this embodiment, the historical behavior data is unstructured data representing historical credit behaviors, and the user behavior sequence includes behavior features and the time sequence information of each behavior feature. Among them, the behavior feature is the behavior feature of the user's financial credit interaction operation, and the behavior feature includes but is not limited to types such as account opening, loan behavior, and repayment behavior.
[0083] In one embodiment, the historical behavior data may be the behavior data of the user's financial credit interaction operations over all past time. In another embodiment, the historical behavior data may be the behavior data of the user's financial credit interaction operations in the past 24 months. In yet another embodiment, the historical behavior data may be the behavior data of the user's financial credit interaction operations in the past 36 months. This embodiment does not make specific limitations in this regard. It is easy to understand that each behavioral feature such as account opening, loan behavior, and repayment behavior in the user's historical behavior data often has corresponding time node information. For further assistance in understanding, an example is given. For example, the user queries a credit card on February 1, 2021, opens a credit card account on March 5, 2021, queries a loan on March 15, 2021, makes a normal repayment of the credit card on April 30, 2021 (the more normal repayments, the lower the credit risk, and the more overdue repayments, the higher the credit risk), and makes a normal repayment of the consumer loan on April 30, 2021. Then the behavioral features here are: the time node corresponding to querying the credit card is February 1, 2021, the time node corresponding to opening the credit card account is March 5, 2021, the time node corresponding to querying the loan is March 15, 2021, the time node corresponding to the normal repayment of the credit card is April 30, 2021, and the time node corresponding to the normal repayment of the consumer loan is April 30, 2021. Then these historical behavior data are sorted and spliced in chronological order to generate a user behavior sequence: querying the credit card, opening the credit card account, querying the loan, normal repayment of the credit card and normal repayment of the consumer loan. Of course, the user behavior sequence may also carry time node identifiers, for example, it may be: 20210201 querying the credit card, 20210305 opening the credit card account, 20210315 querying the loan, 20210430 normal repayment of the credit card and normal repayment of the consumer loan.
[0084] Step S200, perform natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector, and perform Embedding sparse data processing on the first behavior sequence encoding vector to obtain a low-dimensional dense second behavior sequence encoding vector;
[0085] Further, in a feasible embodiment, in the above step S200, the step of performing natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector includes:
[0086] Step A10, according to a preset behavior action dictionary, map each behavioral feature in the user behavior sequence to a behavior code to obtain a behavior code sequence;
[0087] In this embodiment, according to various actions of financial behaviors, a behavior action dictionary is set, and each behavior feature in the borrower behavior sequence (i.e., the user behavior sequence) is mapped to the character encoding in the behavior action dictionary. For the sake of understanding, an example is given. For example, the encoding mapped to "query" in the behavior action dictionary is q, and the encoding mapped to "credit card" is "03". At this time, the behavior number mapped to the behavior feature "query credit card" is "q-03". Another example is that the encoding mapped to "credit card" is "81", and the encoding mapped to "open an account" is "k". At this time, the behavior number mapped to the behavior feature "open an account for credit card" is "k-81". Another example is that the encoding mapped to "credit card" is "81", and the encoding mapped to "normal repayment" is "N". At this time, the behavior number mapped to the behavior feature "normal repayment of credit card" is "81-N", as Figure 2 shown Figure 2 is a schematic diagram of behavior encoding mapping according to an embodiment of the present application. In Figure 2 it, the user behavior sequence is: query credit card, open an account for credit card, query loan, normal repayment of credit card and normal repayment of consumer loan... According to the preset behavior action dictionary, each behavior feature in the user behavior sequence is mapped to a behavior encoding, and the obtained behavior encoding sequence is [q-03], [k-81], [q-02], [81-N, 91-N]....
[0088] Step A20: Perform One-hot encoding on the behavior encoding sequence to obtain a first behavior sequence encoding vector.
[0089] In this embodiment, the concatenated behavior encoding sequence needs to be one-hot encoded. One-hot encoding is a binary representation of categorical variables. First, the categorical values need to be mapped to integer values. Then, each integer value is represented as a binary vector, where the behavior feature is 1 and the rest are 0. For the sake of understanding, an example is given. For example, all actions in the behavior action dictionary are 120-dimensional. For the user behavior sequence A: "query-credit card", "open - business loan" and "credit card-M1" are mapped to behavior encodings. That is, the composed behavior sequence encoding is converted into a 3*120-dimensional vector representation. As Figure 3 shown, in Figure 3 it, the first behavior sequence encoding vector obtained by One-hot encoding of the user behavior sequence A is [1, 0, 0,...], [0, 1, 0,...], [0, 0, 1,...].
[0090] In this embodiment, according to a preset behavior action dictionary, each behavior feature in the user behavior sequence is mapped to a behavior code to obtain a behavior code sequence, and the behavior code sequence is One-hot encoded to obtain a first behavior sequence encoding vector, thereby realizing effective natural language processing of the user behavior sequence (unstructured data).
[0091] After this embodiment performs natural language processing on the user behavior sequence to obtain the first behavior sequence encoding vector, it also performs Embedding sparse data processing on the first behavior sequence encoding vector to obtain a low-dimensional dense second behavior sequence encoding vector. Specifically, after the above natural language processing, the user behavior sequence as unstructured data is converted into the first behavior sequence encoding vector as vector data. However, because the dimension of the first behavior sequence encoding vector is too high, resulting in data sparsity, this embodiment adopts the Embedding sparse data processing technology to perform dimensionality reduction processing on the first behavior sequence encoding vector as high-dimensional sparse data to obtain a low-dimensional dense second behavior sequence encoding vector, as Figure 4 shown Figure 4 is a schematic diagram of the operation of performing Embedding sparse data processing in an embodiment of the present application. Then, the second behavior sequence encoding vector is used as the input of the subsequent recurrent neural network model, thereby reducing the computational complexity of the recurrent neural network model for extracting the behavior sequence representation based on the behavior sequence encoding vector, and further improving the accuracy of credit risk identification.
[0092] After step S200, step S300 is executed to input the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence representation data;
[0093] In this embodiment, the recurrent neural network model can be constructed using the LSTM (Long Short-Term Memory) method, and can also be replaced with the GRU (gated recurrent neural network), bidirectional recurrent neural network, or other conventional recurrent neural network methods. This embodiment does not make specific limitations on this.
[0094] In this embodiment, by inputting the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence representation data, the behavior sequence representation of historical credit interaction behaviors is extracted based on deep learning technology.
[0095] In a possible implementation manner, before the step of inputting the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence representation data, the method further includes:
[0096] Step B10: Train a recurrent neural network model based on the Long Short-Term Memory (LSTM) algorithm.
[0097] In this embodiment, the recurrent neural network model trained based on the LSTM algorithm is an LSTM model, which mainly extracts key information from the encoded vector of the second behavior sequence through three stages:
[0098] (1) Forgetting stage: At the t-th moment of the user's behavior sequence, the LSTM model selectively forgets the unimportant information transmitted at the
[0099] (2) Memory stage: The LSTM model selectively remembers the important behavior information input at the t-th moment;
[0100] Then, the LSTM model adds the results of the above two stages and passes them to the (t + 1)-th moment;
[0101] (3) Output stage: The LSTM model outputs the result at the current t-th moment.
[0102] Using the above three stages, the LSTM model remembers all historical important information and outputs from the initial moment to the end moment of the user behavior sequence, thereby decoding the behavior sequence representation data.
[0103] In this embodiment, since the lending behavior is affected by past behaviors, the LSTM model adopted in this application (i.e., the recurrent neural network model trained based on the LSTM algorithm) is a classic recurrent neural network model. Its advantage is that it can alleviate the problem of gradient disappearance and has the ability of long-term memory, and can capture the influence of past behaviors on current behaviors to predict the future default probability of borrowers. The definition of the LSTM neural network is as follows:
[0104]
[0105] Among them, it, ft, and ot represent the input, forget, and output gates at the t-th moment respectively; xt and ht represent the input value and the hidden output vector; ct is the memory cell vector; ⊙ represents the Hadamard product; W* represents the connection weight; b* represents the corresponding bias; represents the activation function, with the subscripts i, f, o representing the sigmoid activation function, and c and h representing the tanh activation function. During the training process of the LSTM model, the backpropagation error is calculated at each iteration, and each weight is updated accordingly. The structural schematic diagram of this LSTM model is as Figure 5 shown.
[0106] In this embodiment, based on the second behavior sequence encoding vector, an LSTM model is trained, and after iterative convergence, the final LSTM model and its prediction results for each borrower are obtained (the prediction results are behavior sequence characterization data for predicting credit default probability).
[0107] Specifically, based on the constructed LSTM recurrent neural network model, the following ideas can be adopted to extract structured characterizations:
[0108] (1) The predicted value of whether each borrower defaults;
[0109] (2) The average value of the hidden layer vectors of the LSTM model for each borrower (as shown in Figure 6 );
[0110] (3) The last hidden output vector of the LSTM model for each borrower.
[0111] After step S300, step S400 is executed to identify the credit risk of the user based on the behavior sequence characterization data.
[0112] Exemplarily, the steps of identifying the credit risk of the user based on the behavior sequence characterization data include:
[0113] Step C10: Input the behavior sequence characterization data into the trained credit risk prediction model, predict the default probability of the user, and identify the credit risk of the user according to the default probability.
[0114] The technical solution of the embodiment of the present application is to obtain the historical behavior data of the user at different time nodes, sort and splice the historical behavior data at each time node in time sequence to generate a user behavior sequence, where the historical behavior data is unstructured data representing historical credit behaviors, and the user behavior sequence includes behavior characteristics and the time sequence information of each behavior characteristic; perform natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector, and perform Embedding sparse data processing on the first behavior sequence encoding vector to obtain a low-dimensional dense second behavior sequence encoding vector; input the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence characterization data; identify the credit risk of the user according to the behavior sequence characterization data.
[0115] That is, the embodiments of the present application extract the original behavior sequence, and transform the unstructured data by using technologies such as natural language processing and sparse data processing based on the unstructured data, and use it as the input of the recurrent neural network model, so as to extract the behavior sequence representation of the key behavior time sequence, fully capture the sequential relationship information of the user's behavior in the past, make the feature granularity representing the user's historical credit behavior finer and the information expression more accurate, facilitate more accurate prediction of the user's credit default probability, and then achieve a better prediction effect of the credit risk model, and improve the accuracy of credit risk identification.
[0116] The current credit risk identification method is based on a single-dimensional behavior time sequence, and is derived into structured features according to a time window and an aggregation function, and input into models such as decision trees and ensemble trees for interaction of different features.
[0117] Compared with this prior art, the embodiments of the present application collect the original behavior data of the borrower (that is, the historical behavior data as unstructured features), and transform the original behavior data by using technologies such as natural language processing and sparse data processing, so as to avoid using conventional aggregation methods (such as average, ratio or standard deviation, etc.) to aggregate the original behavior information, avoid losing a part of the most original data information, and then make the feature granularity representing the user's historical credit behavior finer and the information expression more accurate, facilitating more accurate prediction of the user's credit default probability. That is, this embodiment is transformed from the traditional extraction of structured features from the behavior sequence to directly using unstructured data such as action text and action flow to build a model. Since more original data is used, it is necessary to combine natural language processing technology and specific model application solutions to solve the problems of unstructured data processing, sparse data transformation and deep sequence feature mining, so as to capture more hidden and valuable information, and thus achieve better credit default probability prediction performance.
[0118] In addition, this embodiment performs dimensionality reduction processing on the first behavior sequence coding vector as high-dimensional sparse data through the Embedding sparse data processing technology to obtain a low-dimensional dense second behavior sequence coding vector, and then uses the second behavior sequence coding vector as the input of the subsequent recurrent neural network model, thereby effectively reducing the data sparsity degree, and further reducing the operation complexity of the recurrent neural network model to extract the behavior sequence representation based on the behavior sequence coding vector, and further improving the accuracy of predicting the user's credit default probability.
[0119] Furthermore, in this embodiment, the action sequence of the borrower's historical behavior is used to create a complete behavior sequence. From the perspective of the overall customer rather than a single dimension, the sequence order among different behaviors is considered. The main purpose is to construct a credit default model without losing the information of the original behavior sequence, fully capture the interaction between behaviors and the sequence information of behaviors, so as to achieve better credit default probability prediction performance, and further achieve the technical purpose of improving the accuracy of credit risk identification.
[0120] Further, based on the first embodiment of the credit risk identification method of the present application, a second embodiment of the credit risk identification method of the present application is proposed.
[0121] In the second embodiment of the credit risk identification method of the present application, the above step S400, the step of identifying the credit risk of the user according to the behavior sequence characterization data includes:
[0122] Step C10, obtain the structural characterization data of the user;
[0123] In this embodiment, the structural characterization data is structured data representing the user attribute characteristics, such as basic age, gender information, the status of the credit account at the current time point, etc.
[0124] Step C20, through a preset classification algorithm, fuse the structural characterization data and the behavior sequence characterization data to obtain fused characterization data, and jointly train a credit risk model;
[0125] Step C30, input the fused characterization data into the trained credit risk model, predict the default probability of the user, and identify the credit risk of the user according to the default probability.
[0126] Exemplarily, the preset classification algorithm can be a logistic regression algorithm. Of course, it can also be replaced with other classification algorithms, such as the NBC (Naive Bayesian Classifier) algorithm, the ID3 (Iterative Dichotomiser 3) decision tree algorithm, the C4.5 decision tree algorithm, the C5.0 decision tree algorithm, the SVM (Support Vector Machine) algorithm, the KNN ( , K-Nearest Neighbor) algorithm, the ANN (Artificial Neural Network) algorithm, etc. This embodiment does not make specific limitations on this.
[0127] In this embodiment, by obtaining the structural characterization data of the user, and through a preset classification algorithm, the structural characterization data and the behavioral sequence characterization data are fused to obtain fused characterization data, and the credit risk model is jointly trained; the fused characterization data is input into the trained credit risk model to predict the default probability of the user, and based on the default probability, the credit risk of the user is identified. Thus, after extracting the key behavioral sequence information (i.e., the behavioral sequence characterization data), combined with the non-temporal structured data of the borrower (i.e., the structural characterization data) for joint modeling, in order to achieve a better model effect. By jointly modeling using behavioral sequence information and structured data, the utilization degree of the data is improved, the credit default probability prediction performance of the model is better, and the accuracy of credit risk identification is further improved.
[0128] Further, in a feasible embodiment, before the step C30 of inputting the fused characterization data into the trained credit risk model to predict the default probability of the user, the method further includes:
[0129] Step D10, obtain the behavioral sequence sample data and the structural sample data, and fuse the behavioral sequence sample data and the structural sample data to obtain a training set and a validation set, where the behavioral sequence sample data is the sample data representing historical credit behaviors, and the structural sample data is the sample data representing the user attribute characteristics;
[0130] In this embodiment, the behavioral sequence sample data and the structural sample data are further obtained through the platform of this institution or the platform of a third-party institution.
[0131] Step D20, based on the training set, perform iterative training on the credit risk model, and evaluate the effect of the credit risk model according to the validation set to obtain a determination result;
[0132] Step D30, if the determination result does not meet the preset standard, continue to perform iterative training on the credit risk model;
[0133] Step D40, if the determination result meets the preset standard, end the iterative training to obtain the trained credit risk model.
[0134] In this embodiment, by obtaining the behavioral sequence sample data and the structural sample data, and fusing the behavioral sequence sample data and the structural sample data to obtain a training set and a validation set, where the behavioral sequence sample data is the sample data representing historical credit behaviors, and the structural sample data is the sample data representing the user attribute characteristics, based on the training set, perform iterative training on the credit risk model, and evaluate the effect of the credit risk model according to the validation set, thereby effectively ensuring the quality of the credit risk model.
[0135] Further, in a feasible embodiment, in D10 of the above steps, the steps of fusing the behavior sequence sample data and the structure sample data to obtain the training set and the validation set include:
[0136] Step E10, through a preset classification algorithm, fuse the behavior sequence sample data and the structure sample data to obtain a fused sample set, where the fused sample set includes multiple training samples and the default labels associated with the features of each training sample;
[0137] Specifically, the fused sample set is a training sample with a default label, and the default label is whether the default is known. Specifically, if training sample A defaults, the value of the default label associated with training sample A is 1; if training sample A does not default, the value of the default label is 0.
[0138] Step E20, divide the fused sample set into a training set and a validation set according to a preset ratio.
[0139] In this embodiment, the fused sample set is divided into a training set and a validation set according to a preset ratio (such as a ratio of 7:3) for subsequent model training.
[0140] Exemplarily, the preset classification algorithm is a logistic regression algorithm. Of course, it can also be replaced with other classification algorithms, such as the NBC (Naive Bayesian Classifier) algorithm, the ID3 (Iterative Dichotomiser 3) decision tree algorithm, the C4.5 decision tree algorithm, the C5.0 decision tree algorithm, the SVM (Support Vector Machine) algorithm, the KNN ( , K-Nearest Neighbor) algorithm, the ANN (Artificial Neural Network) algorithm, etc. This embodiment does not make specific limitations on this.
[0141] In the embodiment of the present application, by extracting the original behavior sequence and applying technologies such as natural language processing and sparse data processing to the unstructured data based on the unstructured data, and using it as the input of the recurrent neural network model, the behavior sequence representation of the key behavior time sequence is extracted, fully capturing the sequential relationship information of the user's behavior in the past time, making the feature granularity representing the user's historical credit behavior finer and the information expression more accurate, facilitating more accurate prediction of the user's credit default probability, achieving a better prediction effect of the credit risk model, and further improving the accuracy of credit risk identification.
[0142] In addition, please refer to Figure 7 , Figure 7This is a schematic structural diagram of a credit risk identification device for the device hardware operating environment involved in the solution of the embodiment of the present application.
[0143] As Figure 7 shown, the credit risk identification device may include: a processor 1001, such as a CPU, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between the processor 1001 and the memory 1005. The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0144] Optionally, the credit risk identification device may further include a rectangular user interface, a network interface, a camera, an RF (Radio Frequency) circuit, sensors, an audio circuit, a WiFi module, etc. The rectangular user interface may include a display screen and an input sub-module such as a keyboard. Optionally, the rectangular user interface may further include a standard wired interface and a wireless interface. The network interface may optionally include a standard wired interface and a wireless interface (such as a WIFI interface).
[0145] Those skilled in the art can understand that Figure 7 the structural diagram of the credit risk identification device shown in
[0146] As Figure 7 shown, the memory 1005, as a computer storage medium, may include an operating device, a network communication module, and a credit risk identification program. The operating device is a program for managing and controlling the hardware and software resources of the credit risk identification device, and supports the operation of the credit risk identification program and other software and / or programs. The network communication module is used to realize the communication between the components inside the memory 1005 and the communication between other hardware and software in the credit risk identification device.
[0147] In Figure 7 the credit risk identification device shown, the processor 1001 is used to execute the credit risk identification program stored in the memory 1005 and perform the following steps:
[0148] Obtain the historical behavior data of the user at different time nodes, sort and splice the historical behavior data at each time node in chronological order to generate a user behavior sequence. Among them, the historical behavior data is unstructured data representing historical credit behaviors, and the user behavior sequence includes behavior characteristics and chronological information of each behavior characteristic;
[0149] Perform natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector, and perform Embedding sparse data processing on the first behavior sequence encoding vector to obtain a second behavior sequence encoding vector that is low-dimensional and dense;
[0150] Input the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence characterization data;
[0151] Identify the credit risk of the user based on the behavior sequence characterization data.
[0152] In some feasible embodiments, the processor 1001 is further configured to execute the credit risk identification program stored in the memory 1005 and perform the following steps:
[0153] According to a preset behavior action dictionary, map each behavior feature in the user behavior sequence to a behavior code to obtain a behavior code sequence;
[0154] Perform One-hot encoding on the behavior code sequence to obtain a first behavior sequence encoding vector.
[0155] In some feasible embodiments, the processor 1001 is further configured to execute the credit risk identification program stored in the memory 1005 and perform the following steps:
[0156] Obtain the structural characterization data of the user, where the structural characterization data is structured data representing the user attribute features;
[0157] Through a preset classification algorithm, fuse the structural characterization data and the behavior sequence characterization data to obtain fused characterization data, and jointly train the credit risk model;
[0158] Input the fused characterization data into the trained credit risk model, predict the default probability of the user, and identify the credit risk of the user based on the default probability.
[0159] In some feasible embodiments, the processor 1001 is further configured to execute the credit risk identification program stored in the memory 1005 and further perform the following steps:
[0160] Obtain behavior sequence sample data and structural sample data, and fuse the behavior sequence sample data and the structural sample data to obtain a training set and a validation set, where the behavior sequence sample data is sample data representing historical credit behaviors, and the structural sample data is sample data representing user attribute features;
[0161] Based on the training set, perform iterative training on the credit risk model, and evaluate the effect of the credit risk model according to the validation set to obtain a determination result;
[0162] If the determination result does not meet the preset standard, continue to perform iterative training on the credit risk model;
[0163] If the determination result meets the preset standard, end the iterative training and obtain the trained credit risk model.
[0164] In some feasible embodiments, the processor 1001 is further configured to execute the credit risk identification program stored in the memory 1005, and further perform the following steps:
[0165] By using a preset classification algorithm, fuse the behavioral sequence sample data and the structural sample data to obtain a fused sample set, where the fused sample set includes multiple training samples and the default labels associated with each training sample;
[0166] Divide the fused sample set into a training set and a validation set according to a preset ratio.
[0167] In some feasible embodiments, the processor 1001 is further configured to execute the credit risk identification program stored in the memory 1005, and further perform the following steps:
[0168] Train a recurrent neural network model based on the long short-term memory network LSTM algorithm.
[0169] The specific implementation manner of the credit risk identification device of the present application is basically the same as that of the above-mentioned embodiments of the credit risk identification method, and will not be elaborated herein.
[0170] In addition, please refer to Figure 8 , Figure 8 which is a schematic diagram of the functional modules of the credit risk identification device of the present application. The present application also provides a credit risk identification device, and the credit risk identification device includes:
[0171] A behavioral sequence extraction module 10, configured to obtain the historical behavioral data of a user at different time nodes, sort and splice the historical behavioral data of each time node in time sequence to generate a user behavioral sequence, where the historical behavioral data is unstructured data representing historical credit behaviors, and the user behavioral sequence includes behavioral features and the time sequence information of each behavioral feature;
[0172] An unstructured data processing module 20, configured to perform natural language processing on the user behavioral sequence to obtain a first behavioral sequence encoding vector, and perform Embedding sparse data processing on the first behavioral sequence encoding vector to obtain a low-dimensional dense second behavioral sequence encoding vector;
[0173] A behavioral sequence characterization module 30, configured to input the second behavioral sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavioral sequence characterization data;
[0174] The credit risk identification module 40 is used to identify the credit risk of a user according to the behavioral sequence characterization data.
[0175] Optionally, the unstructured data processing module 20 is further used for:
[0176] According to a preset behavioral action dictionary, map each behavioral feature in the user behavior sequence to a behavioral code to obtain a behavioral code sequence;
[0177] Perform One-hot encoding on the behavioral code sequence to obtain a first behavioral sequence encoding vector.
[0178] Optionally, the credit risk identification module 40 is further used for:
[0179] Obtain the structural characterization data of the user, where the structural characterization data is structured data representing the user attribute characteristics;
[0180] Through a preset classification algorithm, fuse the structural characterization data and the behavioral sequence characterization data to obtain fused characterization data, and jointly train a credit risk model;
[0181] Input the fused characterization data into the trained credit risk model, predict the default probability of the user, and identify the credit risk of the user according to the default probability.
[0182] Optionally, the credit risk identification device further includes a training module (not shown), and the training module is used for:
[0183] Obtain behavioral sequence sample data and structural sample data, and fuse the behavioral sequence sample data and the structural sample data to obtain a training set and a validation set, where the behavioral sequence sample data is sample data representing historical credit behaviors, and the structural sample data is sample data representing user attribute characteristics;
[0184] Based on the training set, perform iterative training on the credit risk model, and evaluate the effect of the credit risk model according to the validation set to obtain a determination result;
[0185] If the determination result does not meet the preset standard, continue to perform iterative training on the credit risk model;
[0186] If the determination result meets the preset standard, end the iterative training to obtain a trained credit risk model.
[0187] Optionally, the training module is further used for:
[0188] Through a preset classification algorithm, fuse the behavioral sequence sample data and the structural sample data to obtain a fused sample set, where the fused sample set includes multiple training samples and the default labels associated with each training sample;
[0189] Divide the fused sample set into a training set and a validation set according to a preset ratio.
[0190] Optionally, the training module is further configured to:
[0191] Train a recurrent neural network model based on the long short-term memory network (LSTM) algorithm.
[0192] The specific implementation manner of the credit risk identification device of the present application is basically the same as that of the above-mentioned embodiments of the credit risk identification method, and will not be elaborated herein.
[0193] In addition, the present application also provides a computer-readable storage medium, on which a program for credit risk identification is stored. When the credit risk identification program is executed by a processor, the steps of the credit risk identification method of the present application as described above are implemented.
[0194] The specific embodiments of the computer storage medium of the present application are basically the same as those of the above-mentioned embodiments of the credit risk identification method, and will not be elaborated herein.
[0195] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the presence of another identical element in the process, method, article or device including that element.
[0196] The serial numbers of the above-mentioned embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0197] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0198] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present application.
Claims
1. A credit risk identification method, characterized in that, The credit risk identification method includes: Obtain the historical behavior data of the user at different time nodes, sort and splice the historical behavior data of each time node in chronological order to generate a user behavior sequence. Among them, the historical behavior data is unstructured data representing historical credit behaviors, and the user behavior sequence includes behavior characteristics and the chronological information of each behavior characteristic; Perform natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector, and perform Embedding sparse data processing on the first behavior sequence encoding vector to obtain a low-dimensional dense second behavior sequence encoding vector; Input the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence representation data; Obtain the structural representation data of the user. Among them, the structural representation data is structured data representing user attribute characteristics; Through a preset classification algorithm, fuse the structural representation data and the behavior sequence representation data to obtain fused representation data, and jointly train a credit risk model; Input the fused representation data into the trained credit risk model, predict the default probability of the user, and identify the credit risk of the user according to the default probability; The step of performing natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector includes: According to a preset behavior action dictionary, map each behavior characteristic in the user behavior sequence to a behavior code to obtain a behavior code sequence; Perform One-hot encoding on the behavior code sequence to obtain a first behavior sequence encoding vector.
2. The credit risk identification method according to claim 1, characterized in that Before the step of inputting the fused representation data into the trained credit risk model, the method further includes: Obtain behavior sequence sample data and structural sample data, and fuse the behavior sequence sample data and the structural sample data to obtain a training set and a validation set. Among them, the behavior sequence sample data is sample data representing historical credit behaviors, and the structural sample data is sample data representing user attribute characteristics; Based on the training set, perform iterative training on the credit risk model, and evaluate the effect of the credit risk model according to the validation set to obtain a determination result; If the determination result does not meet the preset standard, continue to perform iterative training on the credit risk model; If the determination result meets the preset standard, end the iterative training to obtain a trained credit risk model.
3. The credit risk identification method according to claim 2, characterized in that, The step of fusing the behavior sequence sample data and the structural sample data to obtain a training set and a validation set includes: Through a preset classification algorithm, fuse the behavior sequence sample data and the structural sample data to obtain a fused sample set. Among them, the fused sample set includes multiple training samples and the default labels associated with each training sample; Divide the fused sample set into a training set and a validation set according to a preset ratio.
4. The credit risk identification method according to any one of claims 1 to 3, characterized in that The preset classification algorithm is a logistic regression algorithm.
5. The credit risk identification method according to claim 1, wherein Before the step of inputting the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence characterization data, the method further includes: Training the recurrent neural network model based on the long short-term memory network (LSTM) algorithm.
6. A credit risk identification device, characterized in that, The credit risk identification device includes: A behavior sequence extraction module, configured to obtain historical behavior data of a user at different time nodes, sort and splice the historical behavior data of each time node in time sequence to generate a user behavior sequence, where the historical behavior data is unstructured data representing historical credit behaviors, and the user behavior sequence includes behavior features and time sequence information of each behavior feature; An unstructured data processing module, configured to perform natural language processing on the user behavior sequence to obtain a first behavior sequence encoding vector, and perform Embedding sparse data processing on the first behavior sequence encoding vector to obtain a low-dimensional dense second behavior sequence encoding vector; the unstructured data processing module is further configured to map each behavior feature in the user behavior sequence to a behavior code according to a preset behavior action dictionary to obtain a behavior code sequence; perform One-hot encoding on the behavior code sequence to obtain a first behavior sequence encoding vector; A behavior sequence characterization module, configured to input the second behavior sequence encoding vector into a pre-constructed recurrent neural network model to decode and obtain behavior sequence characterization data; A credit risk identification module, configured to identify the credit risk of the user according to the behavior sequence characterization data, and the credit risk identification module is further configured to obtain structured characterization data of the user, where the structured characterization data is structured data representing user attribute features; fuse the structured characterization data and the behavior sequence characterization data through a preset classification algorithm to obtain fused characterization data, and jointly train a credit risk model; input the fused characterization data into the trained credit risk model, predict the default probability of the user, and identify the credit risk of the user according to the default probability.
7. A credit risk identification device, characterized in that, The credit risk identification device includes: a memory, a processor, and a credit risk identification program stored on the memory and executable on the processor, and when the credit risk identification program is executed by the processor, the steps of the credit risk identification method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, characterized in that, A credit risk identification program is stored on the computer-readable storage medium, and when the credit risk identification program is executed by the processor, the steps of the credit risk identification method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
User credit risk prediction method and system, electronic device and storage medium
CN113837858A