Default prediction method and system, and storage medium
By introducing KAN-based GRU and LSTM models and blank period mechanisms into the loan default prediction model, the shortcomings of existing models to predict and adapt to changes in financial markets in the early stage are solved, and more accurate and practical default predictions are achieved.
Patent Information
- Application Number
- CN202510347975.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
The existing loan default forecast model performed poorly in early predictions, failed to identify high-risk borrowers in a timely manner, and relied on data from the same year, unable to adapt to the rapid changes in the financial market, and ignored the rich information in time series data.
Using KAN-based GRU and LSTM models, the time dependence and dynamic data features in time series data are extracted through the feature extraction layer. The KAN layer decomposes hidden state vectors to obtain nonlinear relationships, the fully connected layer performs nonlinear transformation, and the output layer generates default probability. At the same time, a blank period mechanism was introduced to construct a data set to include window period data and predict observation period data, and a model was trained to predict the default probability after the blank period.
It improves the accuracy and practicality of loan default predictions, can identify default risks in advance, reduce potential losses, adapt to the dynamic financial environment, and enhances the generalization ability and adaptability of the model.
Smart Images

Figure CN120219077A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a default prediction method, system, and storage medium. Background Art
[0002] In the financial field, loan default prediction is crucial for the risk management of financial institutions. Accurately predicting loan defaults helps financial institutions identify high-risk borrowers in advance and take preventive measures such as adjusting credit strategies and strengthening post-loan management in a timely manner, thereby effectively reducing the non-performing loan ratio and improving the asset quality and financial health level. With the development of data mining and machine learning technologies, the application of time series anomaly detection in loan default prediction has become increasingly widespread. However, there are still many problems to be solved in the current models for loan default prediction.
[0003] Existing mainstream models, such as models based on long short-term memory networks (LSTMs), can identify high-risk borrowers to a certain extent, but their performance in early prediction is not satisfactory. Early prediction is of great significance for financial institutions to formulate risk response strategies and reduce potential losses. If the model cannot give an accurate prediction at an earlier stage before loan default occurs, financial institutions will have difficulty in planning in advance, resulting in passive responses when risks occur and being unable to adjust loan strategies, communicate with borrowers, or take other risk mitigation measures in a timely manner.
[0004] Many existing models rely on out-of-sample (OOS) data from the same year for training and testing, which makes it difficult for the models to integrate new data quickly. The financial market environment changes rapidly, and new data such as new market dynamics, changes in economic situations, and changes in borrowers' behavior patterns are constantly emerging. However, due to relying on data from a specific year, existing models cannot adapt to these changes in a timely manner, resulting in the prediction results being unable to accurately reflect the current real risk situation. In addition, this way of using data also limits the full utilization of extensive historical data, prolongs the model development cycle, and cannot fully explore the potential information in historical data, affecting the prediction performance of the model.
[0005] When dealing with the loan default prediction problem, some studies simplify it to a static binary classification problem, ignoring the rich time patterns and trend information contained in time series data. In fact, loan data has obvious time series characteristics, and borrowers' repayment behaviors are correlated at different time points. Past repayment situations have an important impact on future default risks. Ignoring these time series characteristics will cause the model to fail to comprehensively capture the key information in the data, thereby reducing the accuracy and reliability of the prediction.
[0006] In practical applications, the ability to predict loan default risk in advance is crucial, but existing studies rarely attempt to introduce "blank intervals" in model design to simulate early prediction scenarios. Without such simulations, it is difficult to effectively train and evaluate the model in real early prediction scenarios, and it is impossible to accurately evaluate the model's performance in predicting default risk in advance, which greatly reduces the effectiveness of the model in practical applications.
[0007] For example, Chinese patent application CN201910814229.1 discloses a credit delinquency prediction method and system for fusion machine learning, which collects and preprocesses a number of credit factor data, calculates and sorts the importance of the credit factor data in the preprocessing results, and deletes redundancy to obtain selected credit factor data. A training sample is constructed based on the credit factor data, and a credit delinquency prediction model is established and trained using LSTM based on the training sample, and the optimal parameters are determined. After obtaining the optimal model, credit delinquency prediction is performed. This prior art also has the above-mentioned problems.
[0008] In summary, the various deficiencies of the current loan default prediction model limit the effective management of loan default risks by financial institutions. Developing innovative models that can overcome these deficiencies, improve the accuracy and practicality of loan default prediction, and meet the risk management needs of financial institutions in a dynamic financial environment has become an important issue that needs to be solved in this field. Summary of the invention
[0009] The purpose of the present invention is to provide a default prediction method and system, and a storage medium, which partially solve or alleviate the above-mentioned deficiencies in the prior art and can predict loan default risks in advance.
[0010] In order to solve the above-mentioned technical problems, the present invention specifically adopts the following technical solutions: A first aspect of the present invention is to provide a default prediction method, comprising: Constructing a default prediction model, the default prediction model comprising: Several feature extraction layers are used to extract time dependencies and dynamic data features in time series data and output hidden state vectors; The KAN layer is used to decompose the hidden state vector output by the feature extraction layer into a univariate function to obtain the nonlinear relationship in the hidden vector; The fully connected layer is used to perform nonlinear transformation on the univariate function output by the KAN layer; The output layer is used to convert the output of the fully connected layer into a probability between 0 and 1 as the probability of predicting default; Construct a data set, where the data set includes a training set and a test set; the samples in the training set and the test set include window period data and prediction observation period data; the window period data includes time series data of the data period and time series data of the blank period, and the timeline of the blank period is located between the data period and the prediction observation period; and the timeline of the samples in the training set is before the timeline of the samples in the test set. Use the data set to train the default prediction model. Input the time series data to be predicted into the trained default prediction model to predict the default probability during the prediction observation period after the blank period.
[0011] As an improvement, the feature extraction layer is a GRU layer or an LSTM layer; and a batch normalization layer is arranged between GRU layers or between LSTM layers.
[0012] As an improvement, the feature extraction layer has two layers.
[0013] As an improvement, when the feature extraction layer is a GRU layer, the feature extraction layer includes a first GRU layer and a second GRU layer; the first GRU layer includes 128 processing units, and the second GRU layer includes 64 processing units. After the data enters the first GRU layer, the hidden state vectors h of all time steps of the data are output. gur1 ; The batch normalization layer performs batch normalization on the hidden state to obtain the normalized hidden state vector h. gur1norm ; The normalized hidden state vector h. gur1norm After entering the second GRU layer, the hidden state vector h of the last time step is returned. gru2 ; When the feature extraction layer is an LSTM layer, the feature extraction layer includes a first LSTM layer and a second LSTM layer; the first LSTM layer includes 128 processing units, and the second LSTM layer includes 64 processing units. After the data enters the first LSTM layer, the hidden state vectors h of all time steps of the data are output. lstm1 ; The batch normalization layer performs batch normalization on the hidden state to obtain the normalized hidden state vector h. lstm1norm ; The normalized hidden state vector h. lsmt1norm After entering the second LSTM layer, the hidden state vector h of the last time step is returned. lstm2 .
[0014] As an improvement, the window period is 18 months, where the blank period is 3 to 8 months; the prediction observation period is 3 months.
[0015] As an improvement, the output dimension of the KAN layer is 1 and it includes 10 functions.
[0016] As an improvement, a mask layer is provided before the feature extraction layer, and the mask layer labels the padding values in the input time series data.
[0017] As an improvement, before training the default prediction model using the data set, the data in the data set is preprocessed, specifically including: Structurally organize the data set according to the loan serial number, and sort the time series data according to the remaining legal due months; Perform random undersampling on the majority class samples in the data set to reduce the difference between the majority class samples and the minority class samples.
[0018] The present invention also provides a default prediction system, including: A default prediction model construction module for constructing a default prediction model, and the default prediction model includes: A number of feature extraction layers for extracting time-dependent relationships and dynamic data features in the time series data and outputting hidden state vectors; A KAN layer for decomposing the hidden state vectors output by the feature extraction layer into univariate functions to obtain the non-linear relationships in the hidden vectors; A fully connected layer for performing non-linear transformation on the univariate functions output by the KAN layer; An output layer for converting the output of the fully connected layer into a probability between 0 and 1 as the probability of predicting default; A data set construction module for constructing a data set, and the data set includes a training set and a test set; the samples in the training set and the test set include window period data and prediction observation period data; the window period data includes time series data of the data period and time series data of the blank period, and the time line of the blank period is located between the data period and the prediction observation period; and the time line of the samples in the training set is before the time line of the samples in the test set; A training module for training the default prediction model using the data set; A prediction module for inputting the time series data to be predicted into the trained default prediction model to predict the default probability during the prediction observation period after the blank period.
[0019] The present invention also provides a storage medium, and a computer program is stored in the storage medium; when the computer program is executed, the above-mentioned default prediction method can be realized.
[0020] Beneficial effects: The present invention introduces GRU and LSTM models based on KAN. KAN brings new structures or processing methods to the models, endowing them with unique advantages in processing time series data. Time series data (such as the changes in loan-related data over time) usually contains complex non-linear relationships, which may be difficult for traditional models to accurately capture. The GRU and LSTM models based on KAN can more effectively handle these complex relationships, improving the ability to understand and model data.
[0021] Introducing a "blank period" makes it possible to establish a model for early prediction of loan defaults. The blank interval plays a special role in the time series, changing the structure of the data and the way the model processes the data, providing new ideas and methods for early prediction. For financial institutions, such a model becomes a key tool for proactive risk management. Traditional risk assessments may rely more on post hoc data analysis, while early prediction models can detect potential default risks in advance, enabling financial institutions to take proactive measures rather than reacting passively. Early identification of default risks allows banks to take preventive measures, such as adjusting credit policies (such as tightening loan approval criteria, adjusting loan amounts, etc.), which helps screen out more creditworthy loan applicants and reduce the likelihood of defaults. At the same time, timely intervention can also mitigate financial losses and avoid loan losses and related costs caused by defaults.
[0022] The proposed model demonstrates enhanced performance when processing OOT (Out Of Time) data under different conditions. OOT data refers to new data that is temporally different from the training data, and the ability to process such data is an important indicator for measuring the generalization ability and adaptability of the model. The performance of the model when processing OOT data is verified through experimental results, indicating that the performance improvement of the model is not a theoretical speculation but a conclusion obtained through actual data testing and analysis, increasing the credibility of the model. This feature improves its applicability to near-real-time anomaly detection by integrating the latest data in a dynamic financial environment. In a dynamic financial environment, data is constantly changing, and timely acquisition and processing of the latest data are crucial for anomaly detection. The model can utilize this latest data to more promptly detect anomalies such as loan defaults, meeting the needs of financial institutions for near-real-time risk monitoring. Brief Description of the Drawings
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts do not necessarily draw according to the actual scale. Obviously, the following-described drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 It is the flowchart of Embodiment 1; Figure 2 It is the schematic diagram of the GRU-KAN model architecture. The Chinese-English comparison in the figure is as follows: Figure 3 It is the schematic diagram of the LSTM-KAN model architecture. The Chinese-English comparison in the figure is as follows: Figure 4 It is the schematic diagram of the dataset; Figure 5 It is the performance comparison chart between the present invention and the existing model when the blank period is 3 to 8 months. Among them, (a) is the accuracy, (b) is the recall, (c) is the F1 score, and (d) is the area under the curve (AUC); Figure 6 It is the structural diagram of Embodiment 2. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0026] In this text, suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of explaining the present invention, and they have no specific meaning in themselves. Therefore, "module", "component", or "unit" can be used interchangeably. In this text, terms such as "upper", "lower", "inner", "outer", "front", "rear", "one end", "the other end", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it cannot be understood as a limitation of the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In this text, unless otherwise clearly specified and limited, terms such as "installed", "provided with", "connected", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. In this text, "and / or" includes any and all combinations of one or more of the listed related items. In this text, "a plurality" means two or more, that is, it includes two, three, four, five, etc.
[0027] Example 1: As Figure 1 shown, the present invention provides a default prediction method, and its specific steps include: S1 Construct a default prediction model, and the default prediction model successively includes: 1. A data preprocessing layer, which is used to preprocess the input time series data, such as loan-related data, to prepare for subsequent model calculations.
[0028] Specifically, the above time series data refers to the time series data in the time series anomaly detection task. The data preprocessing layer preprocesses the input time series data, such as loan-related data, which is the first step for the model to process data. The preprocessing operations may include data cleaning (removing noise, incorrect data, etc.), data normalization (making data with different features have similar scales to facilitate model learning), data encoding (such as encoding categorical variables), etc. To prepare for subsequent model calculations and ensure that the data input into the model is clean, standardized, and suitable for model processing. For example, there may be missing values, outliers, etc. in loan data, and these problems can be processed through preprocessing to improve the training effect and prediction accuracy of the model.
[0029] 2. A mask layer, which is used to label the padding values in the input time series data.
[0030] When processing loan-related time series data, since the sequence lengths of different samples may be different, in order to be able to batch these data and input them into the model for training, the shorter time series data is usually padded to make their length consistent with the longest sequence. The role of the mask layer is to identify which data is real and valid, and which is padded data added to make up the length. Ensure that the model ignores these irrelevant padded data during training and calculation, and only focuses on learning and processing real and valid data. This can prevent the model from being disturbed by the padded values during training, thereby more accurately capturing the characteristics and patterns in the time series data, and improving the training effect and prediction accuracy of the model.
[0031] 3. Several feature extraction layers, used to extract time dependency and dynamic data features in time series data, and output hidden state vectors. In this embodiment, the feature extraction layer is two layers. The feature extraction layer can be a GRU layer or an LSTM layer; and a batch normalization layer is set between the GRU layers or the LSTM layers.
[0032] The feature extraction layer is used to extract the time dependency and dynamic data features in the time series data, and output a hidden state vector. Specifically, the hidden state vector characterizes the time dependency and dynamic data features contained in the input time series data. It is a number of vectors generated inside the model, representing the state information at a certain point in time. Through the gating mechanism of GRU or LSTM, these models can automatically learn and capture complex patterns in sequence data, thereby encoding time dependency and dynamic features into low-dimensional hidden state vectors. These vectors not only contain local information at each time step, but also reflect the overall time series characteristics through the transfer of state, providing a comprehensive and efficient information representation for subsequent tasks such as default prediction. In this embodiment, the feature extraction layer is two layers, and a GRU (gated recurrent unit) layer or a LSTM (long short-term memory network) layer can be selected. GRU and LSTM are both variants of recurrent neural networks (RNNs), which can effectively process time series data and capture the dependencies of data in the time dimension.
[0033] For example, in loan data, the repayment behaviors of borrowers are correlated at different time points. GRU or LSTM layers can learn these time dependencies to extract dynamic data features related to default. Among them, time dependency refers to the correlation between the current state and past states in time series data. In loan analysis, this is usually manifested in the relationship between a borrower's repayment behavior and their past repayment history. By analyzing these relationships, key patterns affecting default risk can be identified. Dynamic data features refer to those features that change over time in a time series, as opposed to static features (such as a borrower's age, income level, occupation, etc.). These features reflect the behavior patterns of borrowers in the time series and can provide important information about the borrowers' credit risk. In loan data, the following dynamic features can be extracted: Repayment timeliness: Whether the borrower repays on time each month (yes or no).
[0034] Change in repayment amount: The change in the actual monthly repayment amount relative to the amount due.
[0035] Number of consecutive on-time repayments: The number of months the borrower has repaid on time consecutively so far.
[0036] Repayment delay frequency: How many months in the past there have been repayment delays.
[0037] Change in credit score: The trend of the borrower's credit score over time.
[0038] Change in borrowing amount: If revolving borrowing is allowed, the change in the amount borrowed by the borrower each time.
[0039] Figure 2 Shows the model architecture with GRU as the feature extraction layer.
[0040] In the case where the feature extraction layer is a GRU layer, the feature extraction layer includes a first GRU layer and a second GRU layer; the first GRU layer includes 128 processing units, and the second GRU layer includes 64 processing units.
[0041] The feature extraction layer consists of a first GRU layer and a second GRU layer. This multi-layer structure helps the model to dig deeper into data features. The first GRU layer is set with 128 processing units, and the second GRU layer is set with 64 processing units. The number of processing units determines the learning ability and complexity of the model. A larger number of processing units (such as 128 in the first GRU layer) can capture richer feature information, while reducing the number of processing units subsequently (64 in the second GRU layer) helps to screen and refine the features, avoid overfitting of the model, and reduce the computational complexity. Of course, it can be foreseen that the number of GRU layers can be increased or decreased according to the actual situation.
[0042] After the data enters the first GRU layer, the hidden state vectors h of all time steps of the data are output. gur1 ; GRU (Gated Recurrent Unit), as a type of recurrent neural network structure, can effectively capture the temporal dependencies in time series data. In the scenario of loan default prediction, this means that it can learn the associations between loan data at different time points. For example, the repayment behaviors of borrowers at different times in the past have an impact on the current default risk. Outputting the hidden state vectors of all time steps is to completely retain the feature change information of the data in the time series and provide a rich data basis for the processing of subsequent layers. Among them, a time step can be understood as each individual time point in the sequence. A time step is the basic unit of sequence data and reflects the change and evolution process of the data. In stock prediction, each time step may correspond to a data point of one day or one hour. In default prediction, the data is monthly, so the time step is one month.
[0043] The batch normalization layer performs batch normalization on the hidden state to obtain the normalized hidden state vector h. gur1norm ; Batch normalization is a commonly used technique in deep learning. Its main function is to normalize each batch of time series data so that the time series data has a similar distribution. During the model training process, if the distribution of time series data varies greatly, it will lead to difficulties in model training, slow convergence speed, and even problems such as gradient disappearance or gradient explosion. Batch normalization helps to accelerate model training, improve the stability and training efficiency of the model by standardizing the time series data, adjusting the mean of the time series data to 0 and the variance to 1, enabling the model to learn data features faster and more stably.
[0044] The normalized hidden state vector h gur1norm After entering the second GRU layer, it returns the hidden state vector h of the last time step. gru2。The second GRU layer further extracts key features based on the time series data processed by the first GRU layer. Only the hidden state vector of the last time step is returned because, after the comprehensive feature capture and batch normalization processing of the first layer, the hidden state vector of the last time step can comprehensively reflect the key information of the entire time series data. It condenses the feature change trends of the loan data over multiple time steps and is of great value for judging whether a loan defaults. Subsequently, the model can perform more accurate predictive analysis based on this key hidden state vector.
[0045] Figure 2 Shows the model architecture with LSTM as the feature extraction layer.
[0046] When the feature extraction layer is an LSTM layer, the feature extraction layer includes a first LSTM layer and a second LSTM layer; the first LSTM layer includes 128 processing units, and the second LSTM layer includes 64 processing units. Similarly to GRU, the setting of multiple levels with a decreasing number of processing units has its purpose. LSTM (Long Short-Term Memory network) itself has unique advantages in processing time series data and can effectively capture long-term dependencies. By setting multiple LSTM layers with a gradually decreasing number of processing units, on the one hand, the first LSTM layer (128 processing units) can capture as many features and information in the time series data as possible, fully exploring the potential patterns in the time series; on the other hand, the second LSTM layer (64 processing units) further screens and refines the features on the basis of the first layer, removing some unimportant or redundant information, reducing the complexity of the model, avoiding overfitting, and also reducing the computational amount.
[0047] After the data enters the first LSTM layer, the hidden state vectors h of all time steps of the output data lstm1 ; Through its unique gating mechanism (forget gate, input gate, and output gate), the LSTM layer can process each time step in the time series, combine the input information of the current time step with the hidden state of the previous time step, and thus output the hidden state vector reflecting the data characteristics of each time step. In the context of loan default prediction, this means that it can learn the associations between the borrower's repayment behaviors, credit status, etc. at different time points, providing rich feature representations for subsequent analysis. Outputting the hidden state vectors of all time steps retains the complete information of the data in the time dimension and provides a comprehensive data basis for the further processing of subsequent layers.
[0048] The batch normalization layer performs batch normalization on the hidden state to obtain the normalized hidden state vector h lstm1norm ; Similar to GRU, it will not be elaborated here.
[0049] The normalized hidden state vector h lsmt1norm After entering the second LSTM layer, return the hidden state vector h of the last time step lstm2 . After being processed by the first LSTM layer and the batch normalization layer, the data already contains rich and stable feature information. The second LSTM layer further processes this information. Through its gating mechanism, it focuses on extracting the information that is most crucial for the final prediction result. And returning the hidden state vector of the last time step is because this vector synthesizes the information of all previous time steps and condenses the core features of the entire time series data. For example, it can reflect the overall change trend and key features of the loan data over a period of time, which has important reference value for subsequent judgment of whether the loan defaults. The subsequent model layers can perform further analysis and prediction based on this key hidden state vector to obtain the probability of loan default.
[0050] 4. KAN layer, used to decompose the hidden state vector output by the feature extraction layer into univariate functions to obtain the non-linear relationships in the hidden vector
[0051] The core task of the KAN (Kolmogorov - Arnold Networks) layer is to decompose the hidden state vector output by the feature extraction layer (such as the GRU layer or LSTM layer) into univariate functions. The feature extraction layer has already extracted the time-dependent relationships and dynamic data features from the original data and output them in the form of a hidden state vector, but there may be complex non-linear relationships between these features. The KAN layer, through this decomposition operation, deeply explores the non-linear relationships in the hidden vector, providing support for the subsequent model to more accurately understand and process the data.
[0052] In the field of loan default prediction, the relationships between data show obvious non-linear characteristics. The connections between various factors such as borrowers' credit ratings, income levels, and debt situations and whether they finally default are not simple linear relationships. For example, borrowers with high credit ratings do not necessarily not default, and there is no direct linear correspondence between income levels and defaults. There may be some complex interactions and potential factors affecting the occurrence of defaults. These non-linear relationships increase the difficulty of loan default prediction, making it difficult for simple linear models to accurately predict.
[0053] Due to the non-linear characteristics of loan default prediction data, the existence of the KAN layer is particularly important. It can help the model better capture these non-linear relationships, thereby improving the accuracy of loan default prediction. By mining the complex non-linear features in the data, the model can more comprehensively and deeply understand the default risk of borrowers, providing a more reliable decision-making basis for financial institutions, such as more accurately evaluating loan applications, adjusting loan amounts and interest rates, formulating risk management strategies, etc., effectively reducing the losses caused by loan defaults.
[0054] 5. Dense layer, used to perform non-linear transformation on the univariate function output by the KAN layer.
[0055] The Dense layer performs non-linear transformation on the univariate function output by the KAN layer. Through linear transformation and non-linear transformation of the activation function, the features output by the KAN layer are further combined and transformed, so as to extract higher-level feature representations.
[0056] 6. Regularization layer, specifically the Dropout layer. Dropout is a commonly used regularization technique that prevents the model from overfitting and improves the generalization ability of the model by randomly inactivating some neurons during the training process.
[0057] During the model training process, if the model is too complex, overfitting may occur, that is, the model performs well on the training set but poorly on the test set or new data. The Dropout layer randomly discards the outputs of some neurons, so that the model cannot overly rely on certain specific neurons, thereby improving the generalization ability of the model.
[0058] 7. Output layer, used to convert the output of the Dense layer into a probability between 0 and 1 as the probability of predicting default. Specifically, through the Sigmoid activation function, the output of the model is converted into a probability value between 0 and 1 for binary classification tasks, and it is used to judge whether there is default in loan default prediction.
[0059] The characteristic of the Sigmoid function is that it can map the input to the interval [0, 1], and the output value can be understood as the probability that the sample belongs to the positive class (default). Converting the output of the model into an intuitive probability value facilitates financial institutions or relevant personnel to evaluate the default risk of loans based on this probability, so as to make corresponding decisions, such as whether to approve loans, adjust loan interest rates, etc.
[0060] S2. Construct a data set, which includes a training set and a test set; the samples in the training set and the test set include window period data and predicted observation period data; the window period data includes data period data and a blank period, and the time line of the blank period is located between the data period and the predicted observation period; and the time line of the samples in the training set is before the time line of the samples in the test set.
[0061] Another key problem of the existing model is its dependence on using data from the same year for training and testing. This limitation restricts the applicability of the model in real-time detection and prediction, making it necessary to wait until the data is complete before model training can be carried out. In practical applications, since the financial environment and the financial status of borrowers are dynamically changing, banks need to continuously analyze data as new information emerges. If the model can only rely on historical data for training and testing, it cannot flexibly incorporate the latest data into the prediction, and the prediction ability of the model will not be able to timely reflect the current risk situation. Therefore, it will lead to a delay in the model's response to newly emerging risks, thus affecting its accuracy and reliability. To solve these problems, the present invention uses an out-of-time (OOT) test set for model training, thereby improving the accuracy and practical applicability of loan default prediction.
[0062] The OOT test set refers to a data set in which the test set data is in a different time period from the training data used for model training in the time dimension, and is usually data in a new time period outside the time range of the training data. For example, if the training data of the model is time series data from 2010 to 2020, then the OOT data set may be data from 2021 and later. That is, there is a time difference between the training set and the test set, and they cannot overlap.
[0063] The OOT data is strictly taken from a new time period outside the time range of the training data (such as data after 2021 is used to test a model trained based on 2010 - 2020), and can accurately test the adaptability of the model to long-term trends, seasonal fluctuations, and structural changes in time series data. For example, during 2020, the repayment ability of borrowers may show non-cyclical changes, and the OOT test can reveal whether the model has the ability to identify the impact of such unexpected events, while the OOS test is difficult to detect this cross-cycle risk due to time overlap.
[0064] As Figure 4 shown, in this embodiment, the time line of the entire test set is after the training set (the training set also includes a window period and a predicted observation period, not shown in the figure). That is, an OOT test set is used to test the model.
[0065] In the present invention, the window period refers to the complete time range covered when extracting features for each sample, including two parts: the data period and the blank period. For example, the window period of a certain sample is from January 2020 to June 2023, where the data period is from January 2020 to December 2022, and the blank period is from January 2023 to June 2023.
[0066] The data period refers to the historical data time period that can be actually observed within the window period, including features such as the borrower's repayment behavior, credit score, and macroeconomic indicators. As the core input for model training, it is used to extract the historical behavior patterns of borrowers. For example, the repayment records, income changes, etc. of a certain loan sample are all from the data period.
[0067] The blank period is the data-free interval after the data period and before the prediction observation period. For example, the historical data (data period) of a sample ends in December 2022, but the time point to be predicted is July 2023, then the period from January 2023 to June 2023 is the blank period.
[0068] The prediction observation period refers to the target time period for which the model needs to predict whether the borrower will default within this time period. The length of the observation period needs to match the business requirements.
[0069] More specifically, in this embodiment, the window period is 18 months, where the blank period is 3 to 8 months; the prediction observation period is 3 months. Of course, the specific lengths of the window period, the blank period, and the prediction observation period can be adjusted according to actual needs and the prediction ability of the model, and are not limited in the present invention.
[0070] Figure 5It shows the performance comparison between GRU-KAN and LSTM-KAN provided in this embodiment and existing models with the window period set to 3 to 8 months. It can be seen that the GRU-KAN model outperforms the LSTM-KAN and other baseline models in terms of accuracy, recall, and F1-score at each interval. Although the GRU-KAN accuracy is not the highest, it maintains the advantage of the F1-score due to its high recall. In terms of AUC, the LSTM-KAN performs best only when the blank interval is set to 3 months, with an average value of 0.9278. The GRU-KAN consistently achieves the best results in the intervals from 4 to 8 months, demonstrating its robustness and excellent performance under different conditions. It is worth noting that the accuracy, recall, and F1-score of the proposed model at a 5-month blank interval are the same as those of other baseline models at a 3-month interval. This indicates that the proposed model can make equally accurate default predictions two months in advance. This additional reaction time provides a significant advantage for financial institutions, enabling them to more effectively implement early risk management strategies and response measures. This ability is crucial for reducing potential losses and optimizing the decision-making process.
[0071] Setting the blank period indeed largely enables the use of existing data to predict default risks in a longer prediction observation period. Introducing the blank period mechanism during model training expands the time span of model prediction, enabling the model not to be limited to predicting the situation in the immediate time period based on the most recent data, but to be able to predict future time points or time periods (i.e., the prediction observation period after the blank period) based on relatively historical data information.
[0072] In financial lending business, financial institutions can use this method to predict loan default risks much earlier, and then take corresponding measures in advance, such as adjusting credit strategies, strengthening post-loan management, and giving early risk warnings, etc., so as to more effectively reduce potential default losses and enhance the forward-looking and effectiveness of risk management. At the same time, this also enables the model to have stronger adaptability and generalization ability in the time dimension, and can better cope with the lag of data acquisition and the complex and changeable market environment in reality.
[0073] In addition, it is worth noting that there is no order of execution between step S1 and step S2 in this embodiment, and the two steps can also be executed in parallel.
[0074] S3 uses the data set to train the default prediction model.
[0075] After constructing the dataset, the default prediction model is trained using the dataset, and a series of indicators including accuracy, precision, recall, F1 score and AUC are used to comprehensively and accurately evaluate the performance of the model in the loan default prediction task.
[0076] S4 inputs the data to be predicted into the trained default prediction model to predict the default probability during the observation period after the prediction blank period.
[0077] The setting of the blank period allows the model to predict the default risk at a more distant point in the future based on historical data (data period). For example, if the current data ends in March 2025 and the blank period is 5 months, the model can predict the default probability from August to November 2025 (assuming the forecast observation period is 3 months). This buys financial institutions a longer risk response time and avoids passive response due to data lag.
[0078] Through the prediction results, institutions can identify high-risk loans in advance and intervene before default occurs (such as strengthening post-loan tracking, adjusting repayment plans, etc.), thereby reducing the actual default rate.
[0079] Embodiment 2: Figure 6 As shown, the present invention also provides a default prediction system, comprising: The default prediction model building module is used to build a default prediction model, and the default prediction model includes: Several feature extraction layers are used to extract time dependencies and dynamic data features in the data and output hidden state vectors; The KAN layer is used to decompose the hidden state vector output by the feature extraction layer into a univariate function to obtain the nonlinear relationship in the hidden vector; The fully connected layer is used to perform nonlinear transformation on the univariate function output by the KAN layer; The output layer is used to convert the output of the fully connected layer into a probability between 0 and 1 as the probability of predicting default; A data set construction module is used to construct a data set, wherein the data set includes a training set and a test set; the samples in the training set and the test set include window period data and predicted observation period data; the window period data includes data period data and blank period data, and the blank period timeline is located between the data period and the predicted observation period; and the timeline of the samples in the training set is located before the timeline of the samples in the test set; A training module, used to train the default prediction model using a data set; The prediction module is used to input the data to be predicted into the trained default prediction model to predict the default probability during the observation period after the prediction blank period.
[0080] Embodiment 3: The present invention further provides a storage medium, in which a computer program is stored; when the computer program is executed, the above-mentioned default prediction method can be implemented.
[0081] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0082] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a computer terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0083] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.
Claims
1. A default prediction method, characterized in that include: Constructing a default prediction model, the default prediction model comprising: Several feature extraction layers are used to extract time dependencies and dynamic data features in time series data and output hidden state vectors; A KAN layer, used to decompose the hidden state vector output by the feature extraction layer into a single variable function to obtain a nonlinear relationship in the hidden state vector; A fully connected layer, used for performing nonlinear transformation on the univariate function output by the KAN layer; An output layer, used to convert the output of the fully connected layer into a probability between 0 and 1 as a predicted probability of default; Constructing a data set, the data set comprising a training set and a test set; the samples in the training set and the test set comprise window period data and forecast observation period data; the window period data comprises time series data of the data period and time series data of the blank period, the time line of the blank period is between the data period and the forecast observation period; and the time line of the samples in the training set is before the time line of the samples in the test set; Using the data set to train the default prediction model; The time series data to be predicted is input into the trained default prediction model to predict the default probability within the specified prediction observation period after the preset blank period.
2. A default prediction method according to claim 1, characterized in that: The feature extraction layer is a GRU layer or an LSTM layer; and a batch normalization layer is arranged between the GRU layers or between the LSTM layers.
3. A default prediction method according to claim 2, characterized in that: The feature extraction layer consists of two layers.
4. A default prediction method according to claim 3, characterized in that: In the case where the feature extraction layer is a GRU layer, the feature extraction layer includes a first GRU layer and a second GRU layer; the first GRU layer includes 128 processing units, and the second GRU layer includes 64 processing units; After the data to be predicted enters the first GRU layer, the hidden state vector h of all time steps of the data to be predicted is output. gur1 ; The batch normalization layer performs the hidden state vector h gur1 Perform batch normalization to obtain the normalized hidden state vector h gur1norm ; Normalized hidden state vector h gur1norm After entering the second GRU layer, the hidden state vector h of the last time step is returned. gru2 ; In the case where the feature extraction layer is an LSTM layer, the feature extraction layer includes a first LSTM layer and a second LSTM layer; the first LSTM layer includes 128 processing units, and the second LSTM layer includes 64 processing units; After the data to be predicted enters the first LSTM layer, the hidden state vector h of all time steps of the data to be predicted is output. lstm1 ; The batch normalization layer performs the hidden state vector h lstm1 Perform batch normalization to obtain the normalized hidden state vector h lstm1norm ; Normalized hidden state vector h lsmt1norm After entering the second LSTM layer, the hidden state vector h of the last time step is returned. lstm2 .
5. A default prediction method according to claim 2, characterized in that: The window period is 18 months, of which the blank period is 3 to 8 months; the predicted observation period is 3 months.
6. A default prediction method according to claim 1, characterized in that: The output dimension of the KAN layer is 1 and includes 10 functions.
7. A default prediction method according to claim 1, characterized in that: A mask layer is provided before the feature extraction layer, and the mask layer marks the filling values in the input time series data.
8. A default prediction method according to claim 1, characterized in that Before using the data set to train the default prediction model, the time series data in the data set is preprocessed, specifically including: The dataset is structured and organized by loan serial number, and the time series data is sorted by the number of months remaining to legal maturity; Randomly undersample the majority class samples in the dataset to reduce the difference between majority class samples and minority class samples.
9. A default prediction system, characterized in that include: The default prediction model building module is used to build a default prediction model, and the default prediction model includes: Several feature extraction layers are used to extract time dependencies and dynamic data features in time series data and output hidden state vectors; The KAN layer is used to decompose the hidden state vector output by the feature extraction layer into a univariate function to obtain the nonlinear relationship in the hidden state vector; The fully connected layer is used to perform nonlinear transformation on the univariate function output by the KAN layer; The output layer is used to convert the output of the fully connected layer into a probability between 0 and 1 as the probability of predicting default; A data set construction module is used to construct a data set, wherein the data set includes a training set and a test set; the samples in the training set and the test set include window period data and forecast observation period data; the window period data includes time series data of the data period and time series data of the blank period, and the time line of the blank period is between the data period and the forecast observation period; and the time line of the samples in the training set is before the time line of the samples in the test set; A training module, used to train the default prediction model using a data set; The prediction module is used to input the time series data to be predicted into the trained default prediction model to predict the default probability within the specified prediction observation period after the preset blank period.
10. A storage medium, characterized in that: The storage medium stores a computer program; when the computer program is executed, the default prediction method described in any one of claims 1 to 8 can be implemented.
Citation Information
Patent Citations
Credit prediction overdue method and system fused with machine learning
CN110675243A