Credit risk identification method and device, electronic equipment and storage medium

By extracting multi-dimensional interaction features and frequency domain features from user behavior time-series data and combining them with a credit risk identification model, the problem that the RMF model cannot capture subtle behavioral information is solved, and more accurate credit risk identification is achieved.

CN115907970BActive Publication Date: 2026-04-17WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WEBANK (CHINA)
Filing Date
2023-01-05
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing credit risk identification methods based on the RMF model cannot effectively capture subtle, high-value user behavior information, resulting in poor identification performance.

Method used

By acquiring time-series data of user behavior, multi-dimensional interaction features, frequency domain features, and differential local features are extracted. Combined with a trained credit risk identification model, multi-angle and interactive analysis is performed to obtain the probability of default.

Benefits of technology

It achieves comprehensive and accurate credit risk identification based on user behavior, overcomes the shortcomings of the RMF model, and improves the identification effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115907970B_ABST
    Figure CN115907970B_ABST
Patent Text Reader

Abstract

The application discloses a credit risk identification method and device, electronic equipment and a storage medium. The credit risk identification method comprises the following steps: obtaining a to-be-predicted sample, wherein the to-be-predicted sample comprises at least one user behavior time sequence data; performing at least one feature extraction on the to-be-predicted sample to obtain at least one user behavior feature, wherein the user behavior feature comprises at least one of multidimensional interaction features, frequency domain features and difference local features; and inputting each user behavior feature into a trained credit risk identification model to obtain a default probability of the to-be-predicted sample. The application solves the technical problem of poor credit risk identification effect in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology in financial technology (Fintech), and more particularly to a credit risk identification method, device, electronic device, and storage medium. Background Technology

[0002] With the continuous development of fintech, especially internet fintech, more and more technologies (such as distributed systems and artificial intelligence) are being applied in the financial field, but the financial industry is also placing higher demands on technology.

[0003] In the financial industry, credit risk identification of users is often required. Currently, the most common feature engineering approach based on credit loan behavior sequences is the "RMF" model, which consists of Recency, Frequency, and Monetary. Other structured features are derived from this, such as the mean, standard deviation, and ratio to the previous window. However, these structured features typically only provide a holistic, general, and single-perspective representation of the sample data. Many behaviors that pose a default risk may be subtle and not reflected in general features, such as a user's borrowing behavior before each repayment. Therefore, credit risk identification based on the RMF model may miss these hidden but valuable pieces of information, resulting in poor credit risk identification performance. Summary of the Invention

[0004] The main objective of this application is to provide a credit risk identification method, device, electronic device, and storage medium, aiming to solve the technical problem of poor credit risk identification performance in existing technologies.

[0005] To achieve the above objectives, this application provides a credit risk identification method, comprising the following steps:

[0006] Obtain a sample to be predicted, wherein the sample to be predicted contains at least one time-series data of user behavior;

[0007] At least one feature is extracted from the sample to be predicted to obtain at least one user behavior feature, wherein the user behavior feature includes at least one of multi-dimensional interaction features, frequency domain features, and differential local features;

[0008] By inputting the aforementioned user behavior features into a trained credit risk identification model, the default probability of the sample to be predicted is obtained.

[0009] This application also provides a credit risk identification device, the credit risk identification device comprising:

[0010] The acquisition module is used to acquire a sample to be predicted, wherein the sample to be predicted contains at least one time-series data of user behavior.

[0011] The feature extraction module is used to extract at least one feature from the sample to be predicted to obtain at least one user behavior feature, wherein the user behavior feature includes at least one of multi-dimensional interaction features, frequency domain features, and differential local features;

[0012] The prediction module is used to obtain the default probability of the sample to be predicted by inputting the user behavior features into the trained credit risk identification model.

[0013] This application also provides an electronic device, which is a physical device, comprising: a memory, a processor, and a program of the credit risk identification method stored in the memory and executable on the processor. When the program of the credit risk identification method is executed by the processor, it can implement the steps of the credit risk identification method as described above.

[0014] This application also provides a storage medium, which is a computer-readable storage medium, on which a program for implementing a credit risk identification method is stored. When the program for the credit risk identification method is executed by a processor, it implements the steps of the credit risk identification method as described above.

[0015] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the credit risk identification method described above.

[0016] This application provides a credit risk identification method, apparatus, electronic device, and storage medium. By acquiring a sample to be predicted, which contains at least one user behavior time-series data, the method obtains the user behavior time-series data of the sample to be predicted. Then, it extracts at least one feature from the sample to be predicted to obtain at least one user behavior feature. The user behavior feature includes at least one of multi-dimensional interaction features, frequency domain features, and differential local features. This extraction of at least one user behavior feature from multi-dimensional interaction features, frequency domain features, and differential local features is achieved. Local analysis, multi-angle analysis, and / or interaction analysis between multi-dimensional data are performed on user behavior, which can obtain more comprehensive user behavior features. Then, by inputting each of the user behavior features into a trained credit risk identification model, the default probability of the sample to be predicted is obtained, thus achieving a more accurate prediction of the default probability of the sample to be predicted. Compared to credit risk identification based on features obtained from RMF models, the user behavior features used in this application for credit risk identification include multi-dimensional information from time and frequency perspectives, interaction information between user behaviors of different dimensions, and subtle local information. This allows for comprehensive and accurate user behavior analysis from multiple angles, considering the correlation between behavioral information, and from the local to the overall perspective. This overcomes the technical shortcomings of RMF-based credit risk identification, which misses these hidden but high-value information, resulting in poor credit risk identification performance. This improves the effectiveness of credit risk identification. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an embodiment of the credit risk identification method in this application;

[0020] Figure 2 This is a schematic diagram illustrating a scenario of segmented aggregation approximation processing in the credit risk identification method of this application.

[0021] Figure 3 This is a flowchart illustrating another embodiment of the credit risk identification method in this application;

[0022] Figure 4This is a schematic diagram of the structure of one embodiment of the credit risk identification device in this application;

[0023] Figure 5 This is a schematic diagram of the hardware operating environment involved in the credit risk identification method in this application embodiment.

[0024] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0025] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] Example 1

[0027] This application provides a credit risk identification method. In the first embodiment of the credit risk identification method of this application, reference is made to... Figure 1 This includes the following steps:

[0028] Step S10: Obtain the sample to be predicted, wherein the sample to be predicted contains at least one user behavior time series data.

[0029] The subject executing the method in this embodiment can be a credit risk identification device, a credit risk identification terminal device, or a server. This embodiment takes a credit risk identification device as an example, which can be integrated into a terminal device such as a smartphone or tablet computer with data processing capabilities.

[0030] In this embodiment, it should be noted that credit risk refers to the risk that the counterparty will not fulfill its due debt. Credit risk, also known as default risk, refers to the possibility that the borrower, securities issuer, or counterparty will default on the contract due to various reasons, resulting in losses to the bank, investor, or counterparty. The probability of default can be used to characterize the level of a user's credit risk.

[0031] Specifically, the user behavior time series data generated by the user to be identified for credit risk within a certain period of time is obtained, and the user behavior time series data shown are spliced ​​together to form a sample to be predicted. Here, user behavior time series data refers to data recorded in chronological order in a certain dimension.

[0032] Users undergoing credit risk identification may generate one or more different user behavior data over a period of time. For example, the same user may simultaneously generate data on credit card application approval, consumer loan repayment, consumer loan borrowing, mortgage repayment, etc. User behavior data can be classified according to dimensions. That is, the sample to be predicted can be a matrix composed of time-series user behavior data from multiple dimensions. For example, it can be classified by loan type, or by borrowing or repayment, etc. The specific rules for dividing dimensions can be determined according to the actual situation and business needs. This embodiment does not impose any restrictions on this. A user behavior data can belong to one or more dimensions.

[0033] Optionally, the step of obtaining the sample to be predicted includes:

[0034] Step S11: Extract user behavior time series based on at least one preset dimension, wherein the user behavior time series contains at least one user behavior time sequence data.

[0035] Step S12: Concatenate the time series of each user behavior into a sample to be predicted.

[0036] In this embodiment, specifically, based on at least one preset dimension, structured user behavior time series are extracted from the underlying data by dimension, and the various user behavior time series are concatenated into a sample to be predicted. The user behavior time series consists of observations of user behavior in a certain dimension over a period of time; that is, user behavior time data is arranged chronologically. For example, a borrower's credit card query records over the past 1, 2, 3…N months are used as one dimension of the user behavior time series, while the borrower's consumer loan query records over the past 1, 2, 3…N months are used as another dimension of the user behavior time series.

[0037] Optionally, the step of concatenating the time series of user behaviors into a sample to be predicted includes:

[0038] Step S121: Perform segmented aggregation approximation processing on each of the user behavior time series to obtain the denser time series corresponding to each dimension.

[0039] Step S122: The densed time series are spliced ​​together to form a sample to be predicted.

[0040] In this embodiment, it should be noted that structured sequence data has the problem of being relatively sparse. Excessively sparse data makes it difficult for the model to identify valuable information and affects the model's performance. Therefore, this embodiment uses segmented aggregation approximation processing to densify the user behavior time series data before proceeding with subsequent feature engineering methods.

[0041] Specifically, the user behavior time series for each dimension are processed using Piecewise Aggregate Approximation (PAA) to obtain the corresponding dense time series for each dimension. These dense time series are then concatenated to form the sample to be predicted. Piecewise Aggregate Approximation involves dividing the sequence into segments with a certain window length, and then applying an aggregation function to each sub-sequence to compress the sequence length and achieve information density. The window length and the type of aggregation function can be determined according to actual needs; this embodiment does not impose any restrictions on them. For example, referring to… Figure 2 The window length is set to 2, and the aggregation function used is the average value function. The result of piecewise aggregation approximation is as follows: Figure 2 As shown.

[0042] Step S20: Extract at least one feature from the sample to be predicted to obtain at least one user behavior feature, wherein the user behavior feature includes at least one of multi-dimensional interaction features, frequency domain features, and differential local features;

[0043] In this embodiment, it should be noted that the feature values ​​obtained after feature extraction can be one or more. Therefore, the user behavior feature can be a single feature value, a feature vector composed of multiple feature values, or a feature matrix composed of multiple feature vectors. The user behavior feature includes at least one of multi-dimensional interaction features, frequency domain features, and differential local features. The multi-dimensional interaction features are used to characterize the interrelationships between time-series user behavior data in multiple dimensions, such as the interaction information between credit card application approval and consumer loan repayment. The frequency domain features are used to characterize the frequency angle information of time-series user behavior data. The differential local features are used to characterize fragmented, short-term abnormal user behavior, such as a sudden surge in large-amount borrowing behavior after no borrowing behavior occurred in the first four months of six months of data, or borrowing behavior occurring before each repayment.

[0044] Specifically, features are extracted from the sample to be predicted using at least one preset feature extraction algorithm. Each feature extraction can yield a corresponding user behavior feature. Optionally, multi-dimensional interaction features can be extracted using the ROCKET (RandOm Convolutional KErnel Transform) algorithm, frequency domain features can be extracted using the Discrete Fourier Transform algorithm, and differential local features can be extracted using the shapelet discovery algorithm. The ROCKET algorithm, the Discrete Fourier Transform algorithm, and the shapelet discovery algorithm are similar to existing technologies and will not be elaborated further here.

[0045] It should be noted that if multiple dimensions of user behavior time series sequences are extracted, features can be extracted from the user behavior time series sequences of different dimensions in the sample to be predicted, and frequency domain features and / or differential local features corresponding to each dimension can be obtained. Alternatively, features can be extracted from multiple user behavior time series sequences in the sample to be predicted together to obtain multi-dimensional interaction features. If only a single dimension of user behavior time series sequence is extracted, features can be extracted from the user behavior time series sequence to obtain the frequency domain features and / or differential local features corresponding to this dimension.

[0046] Optionally, the step of extracting at least one feature from the sample to be predicted to obtain at least one user behavior feature includes:

[0047] Step A10: Obtain at least one set of random convolution kernels and the sequence arrangement rules corresponding to each of the random convolution kernels;

[0048] In this embodiment, it should be noted that for time series data of user behavior in different dimensions, there is information on the characteristics of user behavior changes over time within the dimension. However, there may also be certain correlations between user behaviors in different dimensions. Traditional feature construction of time series data involves processing and extracting features independently for each time series, while ignoring the interactive correlation information between dimensions. For example, there may be a certain correlation between credit card application approval and consumer loan repayment.

[0049] Specifically, at least one set of random convolutional kernels and the sequence arrangement rules corresponding to each random convolutional kernel are obtained. Each set of random convolutional kernels is pre-trained. Each set of random convolutional kernels may include one or more sub-random convolutional kernels. Multiple sub-random convolutional kernels constitute multiple random convolutional layers. The sequence arrangement rules are determined during the training of each random convolutional kernel. The training process of the random convolutional kernels is similar to the existing technology and will not be elaborated here.

[0050] Step A20: Concatenate the user behavior time series according to the sequence arrangement rules to obtain the user behavior matrix corresponding to each sequence arrangement rule;

[0051] Step A30: Using the random convolution kernels corresponding to each user behavior matrix, feature extraction is performed on each user behavior matrix to obtain at least one multi-dimensional interaction feature.

[0052] In this embodiment, specifically, for each group of random convolutional kernels, a target random convolutional kernel and the target subsequence arrangement rule corresponding to the target random convolutional kernel can be obtained from each random convolutional kernel. The user behavior time series are then arranged vertically according to the target subsequence arrangement rule to obtain the target user behavior matrix corresponding to the target random convolutional kernel. Feature extraction is then performed on the target user behavior matrix using the target random convolutional kernel to obtain the multi-dimensional interaction features corresponding to the target random convolutional kernel. Following the above method, feature extraction of the user behavior matrix by each random convolutional kernel can be completed in parallel or sequentially to obtain the multi-dimensional interaction features corresponding to each target random convolutional kernel. Since there is no interaction between different groups of random convolutional kernels, the obtained multi-dimensional interaction features are independent of each other and can represent different information. Furthermore, this embodiment does not require learning weights or hidden layers for feature extraction using random convolutional kernels, which can greatly reduce the computational load.

[0053] In one feasible approach, the multi-dimensional interactive features include at least a matrix positive value proportion and a matrix maximum value, wherein the matrix positive value proportion refers to the proportion of positive values ​​in the output matrix obtained after convolution with a random convolution kernel, and the matrix maximum value refers to the maximum value in the output matrix obtained after convolution with a random convolution kernel.

[0054] Optionally, the step of extracting at least one feature from the sample to be predicted to obtain at least one user behavior feature includes:

[0055] Step B10: Perform a discrete Fourier transform on the sample to be predicted to obtain a discrete Fourier transform coefficient array.

[0056] Step B20: Determine the frequency domain characteristics of user behavior based on the discrete Fourier transform coefficient array.

[0057] In this embodiment, specifically, the user behavior time-series data in the sample to be predicted is subjected to Discrete Fourier Transform (DFT) for each dimension to obtain a DFT coefficient array corresponding to each dimension. All or part of the DFT coefficients in the DFT coefficient array corresponding to each dimension are determined as the frequency domain features of user behavior for each dimension. For example, for each dimension, a portion of the DFT coefficients can be extracted from the DFT coefficient array based on frequency as the frequency domain features of user behavior, or a portion of the DFT coefficients can be extracted from the DFT coefficient array as the frequency domain features of user behavior through feature engineering methods such as principal component analysis.

[0058] For example, the discrete Fourier transform coefficients can be obtained by performing a preset discrete Fourier transform algorithm on the sample to be predicted, wherein the discrete Fourier transform algorithm is:

[0059]

[0060] In the formula, x n This refers to the time-series data of user behavior before the transformation, X k It refers to the k-th discrete Fourier transform coefficient obtained after the transformation;

[0061] Let w = e 2πi / N The Fourier transform process can be represented in matrix form as follows:

[0062]

[0063] It can be seen that the Fourier transform actually multiplies the original sequence by a matrix on the left, which is equivalent to performing a spatial coordinate transformation. After the user behavior time series data is transformed by angle, frequency angle information can be extracted, which improves the comprehensiveness of the features used for credit risk identification, and thus improves the accuracy and precision of credit risk identification.

[0064] Step S30: By inputting each of the user behavior features into the trained credit risk identification model, the default probability of the sample to be predicted is obtained.

[0065] In this embodiment, specifically, the user behavior features are concatenated into a user behavior feature vector or a user behavior feature matrix, and the user behavior feature vector or user behavior feature matrix is ​​input into a pre-trained credit risk identification model to output the default probability of the sample to be predicted. The credit risk identification model and the training process of the model are similar to existing credit risk identification models, and will not be described in detail here.

[0066] In this embodiment, by acquiring a sample to be predicted, which contains at least one time-series data of user behavior, the user behavior time-series data of the sample to be predicted is obtained. Then, at least one feature is extracted from the sample to be predicted to obtain at least one user behavior feature. The user behavior feature includes at least one of multi-dimensional interaction features, frequency domain features, and differential local features. This achieves the extraction of at least one user behavior feature among multi-dimensional interaction features, frequency domain features, and differential local features. Local analysis, multi-angle analysis, and / or interaction analysis between multi-dimensional data are performed on user behavior, which can obtain more comprehensive user behavior features. Then, by inputting each of the user behavior features into a trained credit risk identification model, the default probability of the sample to be predicted is obtained, thus achieving a more accurate prediction of the default probability of the sample to be predicted. Compared to credit risk identification based on features obtained from RMF models, the user behavior features used in this application for credit risk identification include multi-dimensional information from time and frequency perspectives, interaction information between user behaviors of different dimensions, and subtle local information. This allows for comprehensive and accurate user behavior analysis from multiple angles, considering the correlation between behavioral information, and from the local to the overall perspective. This overcomes the technical shortcomings of RMF-based credit risk identification, which misses these hidden but high-value information, resulting in poor credit risk identification performance. This improves the effectiveness of credit risk identification.

[0067] Example 2

[0068] Furthermore, referring to Figure 3 Based on the above embodiments of this application, in the second embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, before the step of obtaining the initial credit risk identification result of the sample to be evaluated by inputting the sample to be evaluated into the first risk model, the method further includes:

[0069] Step C10: Obtain the reference sequence;

[0070] In this embodiment, it should be noted that a pre-trained shapelet discovery algorithm extracts differential sequence segments from the sample to be predicted, composed of multiple consecutive user behavior time-series data with the highest similarity to the pre-trained shapelet (i.e., the reference sequence). When there are multiple dimensions of user behavior time-series data, the shapelets corresponding to each dimension may be the same or different. The shapelet corresponding to each dimension is determined separately, and differential sequence segments are determined separately for each dimension. The number of pre-trained shapelets can be one or more, and the differential sequence segments corresponding to each shapelet can be determined in parallel or sequentially. The training process of the shapelet discovery algorithm and the discovery and training process of the shapelets are similar to existing technologies and will not be elaborated upon here.

[0071] In one feasible approach, the discovery and training process of shapelets includes:

[0072] Training samples are obtained, and shapelets of the sequences are extracted using a learning method. For ease of explanation, exemplarily, the number of training samples is I, each sample is a training sequence consisting of N user behavior time-series data, the true value of whether a user has violated the rules is denoted as Y, the sliding window length is L, and the i-th training sequence contains J = N - L + 1 sub-training sequences of length L. The total number of extracted shapelets is K. The minimum Euclidean distance between each training sequence and the shapelet is calculated using a preset minimum Euclidean distance algorithm, wherein the minimum Euclidean distance algorithm is:

[0073]

[0074] In the formula, M i,k The minimum Euclidean distance between the i-th sequence and the k-th shapelet is given, where L is the length of the sliding window, l is the index of the data in the sequence, j is the index of the sub-training sequence in the i-th sequence, and T is the index of the sub-training sequence. i,j+l-1 S refers to the time-series data of the user behavior of the (j+l-1)th time sequence in the i-th sequence. k,l It refers to the l-th reference data in the k-th reference sequence.

[0075] The distance matrix M can be obtained by calculating the minimum Euclidean distance for each sample. I×K The shapelet of the sequence can be extracted using a learning method. For ease of explanation, assume Y∈{0,1} I Furthermore, the length L of the shapelet is fixed, and a linear model is used for prediction. The linear model is as follows:

[0076]

[0077] In the formula, the predicted value Coefficient vector W∈R K Furthermore, a bias term W0∈R was added.

[0078] The model's loss function uses cross-entropy loss, defined as:

[0079]

[0080] Adding a L2 regularization term, the objective function becomes:

[0081]

[0082] Then, by optimizing the objective function using stochastic gradient descent, the final shapelet can be obtained. Compared to the method of selecting candidates by truncating subsequences from the original data (which is consistent with the method of screening features, such as using the information gain criterion), which faces the problem of computation due to the large number of candidates, the shapelets obtained by the optimization idea are "hidden" in the sequence and may not find a completely matching subsequence in the data. This solution is also faster.

[0083] Specifically, obtain the pre-trained shapelet, i.e., the reference sequence.

[0084] Step C20: Determine the length of the reference sequence as the length of the sliding window;

[0085] In this embodiment, specifically, the length of the reference sequence is determined, and the length of the reference sequence is determined as the length of the sliding window.

[0086] Step C30: Traverse the time-series data of each user behavior in the sample to be predicted through the sliding window to determine the target subsequence with the highest similarity to the reference sequence;

[0087] In this embodiment, specifically, the user behavior time-series data in the sample to be predicted are traversed through the sliding window. Each sliding window determines a sub-sequence to be compared. Each sub-sequence to be compared contains the same amount of user behavior time-series data as the reference sequence. The similarity between each sub-sequence to be compared and the reference sequence is determined. The target sub-sequence with the highest similarity to the reference sequence is determined from each sub-sequence to be compared. The similarity can be Euclidean distance, Manhattan distance, Mahalanobis distance, etc.

[0088] Optionally, the step of determining the target subsequence with the highest similarity to the reference sequence by traversing each user behavior time-series data in the sample to be predicted through the sliding window includes:

[0089] Step C31: Using the sliding window, multiple pairs of subsequences to be compared are sequentially extracted from each of the user behavior time series data in the sample to be predicted, wherein each subsequence to be compared contains at least one user behavior time series data.

[0090] In this embodiment, specifically, the user behavior time series data in the sample to be predicted is traversed through the sliding window. Each sliding window can extract a subsequence to be compared, and each subsequence to be compared contains the same amount of user behavior time series data as the reference sequence.

[0091] Step C32: Calculate the Euclidean distance between each of the subsequences to be compared and the reference sequence, and determine the target subsequence with the smallest Euclidean distance from each of the subsequences to be compared.

[0092] In this embodiment, specifically, the Euclidean distance between each of the subsequences to be compared and the reference sequence is calculated, and the minimum Euclidean distance is determined. Then, the target subsequence corresponding to the minimum Euclidean distance is determined from each of the subsequences to be compared. In one feasible implementation, the algorithm for the minimum Euclidean distance is as follows:

[0093]

[0094] In the formula, M i,k It refers to the minimum Euclidean distance between the i-th sequence and the k-th reference sequence, L is the length of the sliding window, l is the index of the data in the sequence, j is the index of the subsequence to be compared in the i-th sequence, and T i,j+l-1 S refers to the time-series data of the user behavior of the (j+l-1)th time sequence in the i-th sequence. k,l It refers to the l-th reference data in the k-th reference sequence, where the i-th sequence is the sequence of time-series user behavior data corresponding to each dimension in the sample to be predicted.

[0095] Step C40: Determine the differential local features based on the target subsequence.

[0096] In this embodiment, specifically, the similarity between the target subsequence and the reference sequence can be determined as a differential local feature. When there are multiple reference sequences, the similarity between each reference sequence and its corresponding target subsequence can be concatenated into a similarity vector, and the similarity vector can be determined as a differential local feature.

[0097] In this embodiment, shapelets can effectively capture local trend features in time series. In some financial scenarios, credit risk identification through local features is more accurate than credit risk identification through global features. It can identify abnormal customer behavior in a timely and accurate manner, take countermeasures in advance, and has good interpretability. It is easy to verify with expert experience, thus improving the effectiveness of credit risk identification.

[0098] Example 3

[0099] Furthermore, embodiments of this application also provide a credit risk identification device, referring to... Figure 4 The credit risk identification device is applied to a credit risk identification party and includes:

[0100] The first acquisition module 10 is used to acquire a sample to be predicted, wherein the sample to be predicted contains at least one user behavior time series data.

[0101] The feature extraction module 20 is used to extract at least one feature from the sample to be predicted to obtain at least one user behavior feature, wherein the user behavior feature includes at least one of multi-dimensional interaction features, frequency domain features, and differential local features;

[0102] The prediction module 30 is used to obtain the default probability of the sample to be predicted by inputting each of the user behavior features into the trained credit risk identification model.

[0103] Optionally, the first acquisition module 10 is further configured to:

[0104] User behavior time series is extracted based on at least one preset dimension, wherein the user behavior time series contains at least one user behavior time sequence data;

[0105] The time series of user behaviors described above are concatenated to form a sample to be predicted.

[0106] Optionally, the first acquisition module 10 is further configured to:

[0107] Each user behavior time series is segmented and aggregated to obtain a denser time series for each dimension.

[0108] The densed time series are spliced ​​together to form the sample to be predicted.

[0109] Optionally, the feature extraction module 20 is further configured to:

[0110] Obtain at least one set of random convolution kernels and the sequence arrangement rules corresponding to each of the random convolution kernels;

[0111] The user behavior time series are concatenated according to the sequence arrangement rules to obtain the user behavior matrix corresponding to each sequence arrangement rule.

[0112] By using the random convolution kernels corresponding to each of the user behavior matrices, feature extraction is performed on each of the user behavior matrices to obtain at least one multi-dimensional interaction feature.

[0113] Optionally, the feature extraction module 20 is further configured to:

[0114] Perform a discrete Fourier transform on the sample to be predicted to obtain a discrete Fourier transform coefficient array;

[0115] Based on the discrete Fourier transform coefficient array, the frequency domain characteristics of user behavior are determined.

[0116] Optionally, the feature extraction module 20 is further configured to:

[0117] Obtain the reference sequence;

[0118] The length of the reference sequence is determined as the length of the sliding window;

[0119] By traversing the time-series data of each user behavior in the sample to be predicted through the sliding window, the target subsequence with the highest similarity to the reference sequence is determined.

[0120] Differential local features are determined based on the target subsequence.

[0121] Optionally, the feature extraction module 20 is further configured to:

[0122] Multiple pairs of subsequences to be compared are sequentially extracted from each of the user behavior time series data in the sample to be predicted through the sliding window, wherein each subsequence to be compared contains at least one user behavior time series data.

[0123] Calculate the Euclidean distance between each of the proposed subsequences and the reference sequence, and determine the target subsequence with the smallest Euclidean distance from each of the proposed subsequences.

[0124] The credit risk identification device provided by this invention employs the credit risk identification method in the above embodiments, solving the technical problem of poor credit risk identification effect in the prior art. Compared with the prior art, the beneficial effects of the credit risk identification device provided by the embodiments of this invention are the same as the beneficial effects of the credit risk identification method provided in the above embodiments, and other technical features in this credit risk identification device are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.

[0125] Example 4

[0126] Furthermore, embodiments of the present invention provide an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the credit risk identification method or the conversion qualification cutoff parameter determination method in the above embodiments.

[0127] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as Bluetooth headsets, mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0128] like Figure 5 As shown, an electronic device may include a processing unit (such as a central processing unit, graphics processing unit, etc.) that can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from a storage device into random access memory (RAM). The RAM also stores various programs and arrays required for the operation of the electronic device. The processing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0129] Typically, the following systems can be connected to the I / O interface: input devices including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. Communication devices allow electronic devices to communicate wirelessly or wiredly with other devices to exchange arrays. Although electronic devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.

[0130] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, it performs the functions defined above in the methods of embodiments of this disclosure.

[0131] The electronic device provided by this invention employs the credit risk identification method or the conversion qualification cutoff parameter determination method in the above embodiments, solving the technical problem of poor credit risk identification effect in the prior art. Compared with the prior art, the beneficial effects of the electronic device provided by the embodiments of this invention are the same as the beneficial effects of the credit risk identification method or the conversion qualification cutoff parameter determination method provided in the above embodiments, and other technical features in this electronic device are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.

[0132] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.

[0133] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0134] Example 6

[0135] Furthermore, this embodiment provides a computer-readable storage medium having computer-readable program instructions stored thereon, which are used to execute the credit risk identification method or the conversion qualification cutoff parameter determination method in the above embodiments.

[0136] The computer-readable storage medium provided in this embodiment of the invention may be, for example, a USB flash drive, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0137] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0138] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to: acquire a sample to be predicted, the sample to be predicted containing at least one time-series data of user behavior; perform at least one feature extraction on the sample to be predicted to obtain at least one user behavior feature, wherein the user behavior feature includes at least one of multi-dimensional interaction features, frequency domain features, and differential local features; and obtain the default probability of the sample to be predicted by inputting each of the user behavior features into a trained credit risk identification model.

[0139] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0141] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0142] The computer-readable storage medium provided by this invention stores computer-readable program instructions for executing the above-described credit risk identification method or conversion qualification cutoff parameter determination method, thus solving the technical problem of poor credit risk identification performance in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the embodiments of this invention are the same as the beneficial effects of the credit risk identification method or conversion qualification cutoff parameter determination method provided in the above-described embodiments, and will not be repeated here.

[0143] Example 7

[0144] Furthermore, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the credit risk identification method or the conversion qualification cutoff parameter determination method described above.

[0145] The computer program product provided in this application solves the technical problem of poor credit risk identification effect in existing technologies. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this invention are the same as the beneficial effects of the credit risk identification method or the method for determining conversion qualification cutoff parameters provided in the above embodiments, and will not be repeated here.

[0146] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.

Claims

1. A credit risk identification method characterized by, The credit risk identification method includes the following steps: Obtain a sample to be predicted, wherein the sample to be predicted contains at least one time-series data of user behavior; Feature extraction is performed on the sample to be predicted to obtain user behavior features, wherein the user behavior features include multi-dimensional interaction features, frequency domain features, and differential local features; By inputting the aforementioned user behavior features into a trained credit risk identification model, the default probability of the sample to be predicted is obtained. The step of extracting features from the sample to be predicted to obtain user behavior features includes: Obtain the reference sequence; The length of the reference sequence is determined as the length of the sliding window; By traversing the time-series data of each user behavior in the sample to be predicted through the sliding window, the target subsequence with the highest similarity to the reference sequence is determined. Based on the target subsequence, determine the differential local features; Obtain at least one set of random convolution kernels and the sequence arrangement rules corresponding to each of the random convolution kernels; The user behavior time series are concatenated according to the sequence arrangement rules to obtain the user behavior matrix corresponding to each sequence arrangement rule. By using the random convolution kernels corresponding to each of the user behavior matrices, feature extraction is performed on each of the user behavior matrices to obtain at least one multi-dimensional interaction feature; Perform a discrete Fourier transform on the sample to be predicted to obtain a discrete Fourier transform coefficient array; Based on the discrete Fourier transform coefficient array, the frequency domain characteristics of user behavior are determined.

2. The credit risk identification method as described in claim 1, characterized in that, The steps for obtaining the sample to be predicted include: User behavior time series is extracted based on at least one preset dimension, wherein the user behavior time series contains at least one user behavior time sequence data; The time series of user behaviors described above are concatenated to form a sample to be predicted.

3. The credit risk identification method as described in claim 2, characterized in that, The step of concatenating the various user behavior time series into a sample to be predicted includes: Each user behavior time series is segmented and aggregated to obtain a denser time series for each dimension. The densed time series are spliced ​​together to form the sample to be predicted.

4. The credit risk identification method as described in claim 1, characterized in that, The step of determining the target subsequence with the highest similarity to the reference sequence by traversing the time-series user behavior data in the sample to be predicted through the sliding window includes: Multiple pairs of subsequences to be compared are sequentially extracted from each of the user behavior time series data in the sample to be predicted through the sliding window, wherein each subsequence to be compared contains at least one user behavior time series data. Calculate the Euclidean distance between each of the proposed subsequences and the reference sequence, and determine the target subsequence with the smallest Euclidean distance from each of the proposed subsequences.

5. A credit risk identification device, the credit risk identification device comprising: The acquisition module is used to acquire a sample to be predicted, wherein the sample to be predicted contains at least one time-series data of user behavior. The feature extraction module is used to extract features from the sample to be predicted to obtain user behavior features, wherein the user behavior features include multi-dimensional interaction features, frequency domain features, and differential local features. The prediction module is used to obtain the default probability of the sample to be predicted by inputting the user behavior features into the trained credit risk identification model. The feature extraction module is further configured to: acquire a reference sequence; determine the length of the reference sequence as the length of a sliding window; traverse the time-series data of each user behavior in the sample to be predicted through the sliding window to determine the target subsequence with the highest similarity to the reference sequence; determine differential local features based on the target subsequence; acquire at least one set of random convolution kernels and the sequence arrangement rules corresponding to each random convolution kernel; concatenate each user behavior time series according to each sequence arrangement rule to obtain a user behavior matrix corresponding to each sequence arrangement rule; extract features from each user behavior matrix using the random convolution kernel corresponding to each user behavior matrix to obtain at least one multi-dimensional interaction feature; perform a discrete Fourier transform on the sample to be predicted to obtain a discrete Fourier transform coefficient array; and determine the frequency domain features of user behavior based on the discrete Fourier transform coefficient array.

6. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps of the credit risk identification method according to any one of claims 1 to 4.

7. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and the computer-readable storage medium stores a program for implementing the credit risk identification method. The program for implementing the credit risk identification method is executed by a processor to implement the steps of the credit risk identification method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Risk assessment method and device based on convolutional neural network

    CN110490424A

  • Credit data processing method and device, electronic equipment and computer program medium

    CN114519630A