Credit risk prediction method based on personalized federal continuous learning
By adopting a personalized federal continuous learning method in credit risk assessment, the problem of insufficient data privacy and model adaptability in traditional methods is solved, efficient and personalized credit risk prediction is achieved, and the generalization ability and practicality of the model are improved.
Patent Information
- Application Number
- CN202510639057.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
AI Technical Summary
Traditional credit risk assessment methods have data leakage, privacy issues and insufficient model adaptability when dealing with large-scale and dynamically changing data, making it difficult to effectively integrate heterogeneous data and adapt to new data distribution and business needs.
Using a method based on personalized federated continuous learning, local model parameters are uploaded to the server through each client, and layer by layer is analyzed and optimized. The contribution of global model parameters is weighted and averaged, the global model parameters are updated, and incremental updates are performed on the local model to achieve personalized adaptability and continuous learning.
It realizes that while protecting data privacy, it integrates heterogeneous data, improves the generalization ability and prediction accuracy of the model, can adapt to new data distribution and business needs in real time, reduces the risk of model obsoleteness, and improves the practicality and robustness of the model.
Smart Images

Figure CN120163645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of credit risk prediction, and in particular to a credit risk prediction method based on personalized federated continuous learning. Background Art
[0002] With the rapid development of fintech and the acceleration of digital transformation, credit risk assessment has become a core link for financial institutions to manage risks and optimize decisions. Traditional credit risk assessment methods mainly rely on historical data and static models, and these methods have obvious limitations in dealing with large-scale and dynamically changing data. For example, traditional methods usually require centralized data processing and model training, which may not only lead to data leakage and privacy issues, but also be difficult to effectively integrate heterogeneous data from different financial institutions or data sources. At the same time, due to the dynamic changes of data distribution and risk patterns over time, traditional models are difficult to adapt to the new stage requirements in a timely manner and cannot effectively capture the latest credit risk changes. In addition, centralized models lack personalized adaptability, cannot make full use of the characteristics of local data, and are difficult to meet the specific needs of different financial institutions or regions.
[0003] In recent years, as a distributed machine learning framework, federated learning allows different clients to independently train models on local data and share model parameters in a secure manner, so as to achieve the optimization of the global model while protecting data privacy. However, when traditional federated learning methods are used to handle credit risk prediction tasks, there are still problems such as insufficient model adaptability and lack of continuous learning ability. For example, the global model is difficult to adapt to the local data distribution of different clients, resulting in poor model performance on some clients; at the same time, traditional federated learning methods are difficult to effectively accumulate knowledge and adapt to the new stage requirements when facing dynamically changing data. At the same time, continuous learning technology provides a new idea for solving the model update problem in a dynamic data environment. Continuous learning accumulates the knowledge of each stage through a deep neural network to build a more stable model, which can not only help achieve better results in future stages, but also perform well in previous stages. However, how to achieve global optimization in the federated learning framework while retaining the personalized characteristics of local models is the key challenge for realizing efficient credit risk prediction. Summary of the Invention
[0004] The purpose of the present invention is to overcome the problems of the prior art, and provides a credit risk prediction method based on personalized federated continuous learning.
[0005] The purpose of the present invention is achieved through the following technical solutions: A credit risk prediction method based on personalized federated continuous learning, the method includes the following steps: Each client uploads all layer parameters of its local model to the server; the local model is a credit risk prediction model trained based on local lending sample data; The server receives all layer parameters of the local models uploaded by each client, and performs layer-by-layer analysis on all layer parameters of each local model through a matching algorithm to obtain the contribution degree of each layer parameter of each local model in the corresponding layer global parameters of the global model. Furthermore, based on the contribution degree, it optimizes the global parameters of each layer of the global model and sends the optimized global parameters to each client; The client receives the optimized global parameters and applies the optimized global parameters to the corresponding layer of the local model; Incrementally update the local model based on the training dataset of the current stage; Use the local model after incremental update to perform a credit risk prediction task and output a prediction result.
[0006] In one example, applying the optimized global parameters to the corresponding layer of the local model includes: Evaluate the importance of the local model parameters, freeze the parameters with the top P% importance rankings in the local model, and use the optimized global parameters to replace the 1 - P% parameters with lower rankings in the local credit risk prediction model, where P is the parameter importance ratio threshold.
[0007] In one example, after sending the optimized global parameters to each client, it further includes: Perform weighted average processing on the final global parameters of the global model according to the class distribution of the local lending sample data to update the global model parameters.
[0008] In one example, the local model uses elastic weight consolidation loss, or cross-entropy loss and elastic weight consolidation loss to optimize the model parameters.
[0009] In one example, incrementally updating the local model based on the training dataset of the current stage includes: During the update period, preprocess the newly added lending sample data to obtain the training dataset of the current round, and use the training dataset of the current round to perform incremental training on the local model obtained from the previous round of training. During the incremental training process, use the parameter knowledge extracted from the local model obtained from the previous training round to perform parameter optimization to obtain the local model corresponding to the current training round; Using the parameter knowledge extracted from the local model obtained from the previous training round to perform parameter optimization includes: Compare the parameters of the local model in the current round of training task with the parameter knowledge extracted from the local model obtained in the previous round of training task, and incorporate the obtained parameter differences into the loss function for optimizing the model parameters in the current round of model training task, so that the parameters with larger weights in the local model obtained in the previous round of training task maintain larger weights in the loss function of the current round of training task.
[0010] In one example, after applying the optimized global parameters to the corresponding layer of the local model or after applying the optimized global parameters to the corresponding layer of the local model, it further includes: Determine whether the set update period of the local model is reached. If so, determine whether the update times of the local model are an integer multiple of the constant S. If so, return to the step where each client uploads the parameters of any layer of its own local model to the server until the optimized global parameters are applied to the corresponding layer of the local model; if not, perform an incremental update on the local model and execute the credit risk prediction task.
[0011] It should be further noted that the technical features corresponding to the above examples can be combined or replaced with each other to form a new technical solution.
[0012] Compared with the prior art, the beneficial effects of the present invention are: 1. In one example, the present invention constructs a federated continuous learning framework, which not only realizes the continuous update of the model in the time dimension, but also effectively integrates heterogeneous data from different financial institutions or data sources through the multi-client knowledge sharing mechanism in the space dimension, realizing effective group intelligence collaborative credit risk prediction of resources. Compared with the traditional centralized learning method, the present invention protects data privacy while making full use of the advantages of distributed data, that is, training the local model with local data under the federated learning architecture can make full use of local characteristics, so that the local model has personalized adaptability to meet the specific needs of different financial institutions or regions, significantly improving the generalization ability and prediction accuracy of the model. In addition, through the continuous learning mechanism, the model can adapt to new data distributions and business requirements in real time, effectively capture the latest credit risk changes, reduce the risk of model obsolescence, and improve the practicality and robustness of the model.
[0013] Furthermore, the server can identify hidden elements with similar characteristics by layer-by-layer analyzing the parameters of any layer of each client through a matching algorithm, so as to more precisely capture the characteristics of different layers of each local model, and thus more accurately evaluate the contribution of the parameters of each local model to the global model parameters, effectively solving the problem of performance degradation caused by directly averaging weights in traditional federated learning, and ensuring that the global model parameters can better reflect the common characteristics of client data.
[0014] 2. In one example, the importance of model parameters is evaluated. According to the evaluation results, the parameters ranked in the top P% in terms of importance are frozen, thereby protecting these key parameters from interference in subsequent training stages, alleviating the catastrophic forgetting problem, and ensuring that the model does not degrade its performance on old tasks when learning new tasks. The parameters in the last 1 - P% are directly replaced with the optimized global parameters, further optimizing the local model parameters, improving the stability and adaptability of the model, and enabling the model to maintain good performance under different data distributions and stage requirements.
[0015] Meanwhile, freezing the important parameters, that is, restricting the variation range of the important parameters of the local model, realizes an efficient model update in terms of resources, avoids frequent global updates and a large amount of data transmission, significantly reduces the computational and communication costs, and is especially suitable for resource - constrained scenarios. Compared with the traditional global model update method, the present invention reduces unnecessary consumption of computing resources and improves the efficiency of model update through local updates. At the same time, by evaluating the importance of parameters and freezing the important parameters, the model update process is further optimized, ensuring the stability and effectiveness of the model.
[0016] 3. In one example, regularization constraints are introduced during the training process: the elastic weight consolidation loss and the cross - entropy loss function. By minimizing the cross - entropy loss, the model can continuously adjust its internal parameters to make the prediction results closer to the true labels, thereby optimizing the model's ability to identify normal samples and also better identify high - risk samples. The elastic weight consolidation loss is used to avoid catastrophic forgetting, enabling the model to retain the memory of knowledge in the old stage when learning a new stage.
[0017] 4. In one example, the final weights of the global model are weighted - averaged according to the class distribution of local lending sample data. The parameters after weighted - averaging can more accurately reflect the class distribution of local data, thereby constructing a model that is more adaptable to the global data distribution and improving the generalization ability of the global model on global data.
[0018] 5. In one example, during the incremental update process, the parameters of the local model in the current round of the training stage are compared with the model parameter knowledge extracted in the previous round of the training stage, and the parameter differences obtained from the comparison are incorporated into the loss function of the current round of the model training stage. When the model learns a new stage, it is constrained by the regularization term, making the parameters with larger weights in the previous round of the training stage maintain higher weights in the current round of the training stage, thereby further improving the accuracy of credit risk prediction.
[0019] 6. In one example, by determining whether the number of updates of the local model is an integer multiple of a constant S, the personalized federated step or the local model update step is executed, thereby controlling the timing of global model parameter aggregation and the local model update frequency, which not only helps to optimize the personalized training of the local model, but also improves communication efficiency and speeds up the model convergence rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings. The accompanying drawings used herein are provided to further understand the present application and form a part of the present application. The same reference numerals are used in these drawings to represent the same or similar parts. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application.
[0021] Figure 1 The flowchart of the method provided for an example of the present invention; Figure 2 The schematic diagram of the deep neural network provided for an example of the present invention; Figure 3 The schematic diagram of personalized protection provided for an example of the present invention; Figure 4 The schematic diagram of continuous training of the local model provided for an example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. At the same time, the technical features involved in different examples of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0023] In one example, a credit risk prediction method based on personalized federated continuous learning is as Figure 1 shown, and the method includes the following steps: S1: Each client uploads all layer parameters of its respective local model to the server.
[0024] Among them, the local model is a credit risk prediction model trained by each client (such as a bank client) based on its own local data, which can capture the unique patterns and regularities of local data, so as to provide diverse information for the global model. The credit risk prediction model can be a deep neural network, a convolutional neural network, a recurrent neural network, a Transformer model, etc. The local data is local lending sample data, which can be text data, image data, etc. Further, the local lending sample data includes loan-related information, basic information of borrowers, credit assessment-related data. The loan-related information includes loan amount, loan term, loan interest rate, loan status, loan purpose. The basic information of borrowers includes annual income, debt-to-income ratio, home ownership and years of work. The credit assessment-related data includes loan grade, credit score and the earliest credit record time. The repayment and historical records include loan disbursement date, last repayment date and total amount paid. The above local data serves as the training dataset of the local model, and the dataset includes abnormal data, that is, high-risk lending samples.
[0025] Optionally, after each client uploads all layer parameters of its own local model to the server, in cooperation with the subsequent personalized federated collaborative learning steps S2-S3, the parameter update starts layer by layer from the first layer parameters of the local model. The model parameters include weights, biases, etc. In this example, it specifically refers to the weights of the model. Further, the layer is the hierarchical structure of the local model, such as the hidden layer in a deep neural network.
[0026] Preferably, before step S1, it further includes: S0: Initialize the local model and the global model.
[0027] Specifically, the model initialization includes defining the model structure, configuring learning parameters, etc. Among them, the global model is a credit risk prediction model obtained by aggregating the local model parameters of all clients, which can reflect the comprehensive characteristics of all clients' data, and thus learn more comprehensive and general feature representations. The local model and the global model can have the same network structure. In this example, the local model and the global model are preferably deep neural networks for performing credit risk prediction, that is, the local model and the global model are credit risk prediction models. The deep neural network includes an input layer, hidden layers and an output layer. The input layer receives the original data (local data) and passes it into the network. The hidden layers process the weighted sum of multiple neurons and non-linear activation functions to extract high-level features of the local data layer by layer. The output layer then converts the features extracted by the hidden layer into the final prediction result according to the stage requirements.
[0028] Further, constructing a credit risk prediction model with a deep neural network structure includes: Constructing a deep neural network with one input layer, several hidden layers and one output layer, such asFigure 2 As shown. After a series of linear processing and activation processing on the input loan sample through the weight matrix and bias vector of the deep neural network, it is then processed by logistic regression, and finally the final credit risk prediction result is output through the output layer. On the local client, the deep neural network first maps the input loan sample data to a low-dimensional feature space, and through the weighted summation and non-linear activation function processing of multiple layers of neurons in the hidden layer, extracts the feature patterns that can effectively represent credit risk layer by layer. These feature patterns reflect the differences between different borrowers and the potential laws related to credit risk. For example, a borrower with multiple historical overdue records will have a higher risk proportion in subsequent credit risk prediction. Finally, the output layer of the model converts the features extracted by the hidden layer into specific credit risk prediction results according to the stage requirements. The present invention mines the internal correlation relationship between the borrower's behavior characteristics and credit risk through the model loan sample, and then realizes credit risk prediction.
[0029] By introducing a deep neural network as the local risk prediction model, its powerful feature extraction ability can be fully utilized to map complex credit risk data to a low-dimensional feature space, and through multi-layer non-linear transformation, extract the feature patterns that can effectively represent credit risk. This automated feature extraction method not only avoids the subjectivity of relying on manual experience to select features in traditional methods, but also can capture the subtle differences between different borrowers and the potential laws related to credit risk, thus significantly improving the model's ability to identify credit risk, especially the recognition accuracy for high-risk samples.
[0030] Optionally, the local model and global model of the present invention are not only applicable to credit risk prediction, but also applicable to various application scenarios such as financial abnormal transaction detection, financial fraud detection, intelligent investment decision-making, etc. Taking financial abnormal transaction detection as an example, at this time, the local model and global model are abnormal transaction detection models, and the network structure is any one of a deep neural network, a convolutional neural network, a recurrent neural network, and a Transformer model. At this time, the present invention provides an abnormal transaction detection method based on personalized federated continuous learning, including the following steps: Each client uploads all layer parameters of its local model to the server respectively; the local model is an abnormal transaction detection model trained based on local credit card sample data; The server receives all layer parameters of the local models uploaded by each client, and through a matching algorithm, analyzes all layer parameters of each local model layer by layer to obtain the contribution degree of each layer parameter of each local model in the corresponding layer global parameters of the global model, and then optimizes the corresponding layer global parameters of the global model based on the contribution degree, and sends the optimized global parameters to each client; The client receives the optimized global parameters and applies the optimized global parameters to the corresponding layer of the local model; Incrementally update the local model based on the training dataset at the current stage; Use the local model after incremental update to perform the abnormal transaction detection task and output the prediction result.
[0031] Specifically, the data types of the credit card sample data input into the local model are text, images, etc. The credit card sample data includes the basic information of the cardholder and the credit card transaction information. The basic information of the cardholder includes gender, education level, marital status, and age. The credit card transaction information includes transaction time, transaction location, transaction amount, and transaction frequency. At this time, after a series of linear processing and activation processing of the input credit card sample through the weight matrix and bias vector of the deep neural network, it is then processed by logistic regression, and finally the final abnormal transaction detection result is output through the output layer. On the local client, the deep neural network first maps the input credit card sample data to a low-dimensional feature space, and through the weighted summation and non-linear activation function processing of multiple layers of neurons in the hidden layer, extracts the feature patterns that can effectively represent abnormal transactions layer by layer. Finally, the model output layer converts the features extracted by the hidden layer into specific abnormal transaction detection results according to the stage requirements.
[0032] It should be noted that the technical concepts of the model performing credit risk prediction tasks, financial abnormal transaction detection tasks, financial fraud detection tasks, and intelligent investment decision-making tasks in the method of the present invention are the same, and the only differences lie in the sample data input into the model and the prediction purposes.
[0033] S2: The server receives all layer parameters of the local models uploaded by each client, and analyzes all layer parameters of the local models layer by layer through a matching algorithm, and then identifies the hidden elements with similar features, obtains the contribution degree (matching result) of the layer parameters of each local model in the corresponding layer global parameters of the global model, and then optimizes the global parameters of each layer of the global model layer by layer based on the contribution degree, rather than simply performing element-wise averaging, and broadcasts the optimized global parameters to each client.
[0034] Among them, the matching algorithm can be the cosine similarity algorithm or other similarity measurement methods. At this time, the weighted aggregation of the server layer parameters (weights) specifically includes the following steps:
[0035] Suppose we have K clients, and each client has a local model , where represents the client number; represents the number of layers of the model.
[0036] The server first collects the weights of the th layer of the local model from each client 。Then, these weights are analyzed through a matching algorithm to identify hidden elements with similar characteristics. The matching result of the matching algorithm is a set of similarity weights, indicating the contribution degree of each client weight in the global model.
[0037] Based on the matching result, the server calculates the global weight of the th layer of the global model , and the formula is as follows: ; where represents the contribution weight of client in the weight of the th layer, where and ; This way of weighted aggregation of hierarchical parameters can ensure that the weights of the global model can better reflect the common characteristics of all client data, and at the same time avoid the performance degradation problem that may be caused by directly averaging the weights.
[0038] Finally, the server distributes the calculated global weight of the th layer back to all clients.
[0039] S3: The client receives the optimized global parameters and applies the optimized global parameters to the corresponding layer of the local model.
[0040] S4: Incrementally update the local model based on the training dataset of the current stage.
[0041] Specifically, the local model is trained using a continuous learning strategy, with each round of model training as a training stage. For the local model in the first stage, when the client trains the local model in the first stage, it preprocesses the currently available credit sample data for training to obtain a training credit sample set; the data preprocessing methods include but are not limited to handling missing data, deleting redundant features, constructing new features, balancing data categories, etc., aiming to prepare for training a high-quality model. Then, based on the training credit sample set, the local model is trained using the backpropagation method to optimize and adjust the network parameters, so as to continuously improve the model's ability to identify credit risks and prediction accuracy, and obtain the local model corresponding to the first stage.
[0042] Based on the existing local model, the local model is further trained and optimized by introducing new stage training data to adapt to possible changes in abnormal patterns in the new data. This method allows the local model to continuously learn and update its knowledge base, thereby improving the ability to identify newly emerging abnormal lending behaviors, while maintaining the memory of old credit risk predictions, and effectively integrating new knowledge without retraining the entire model.
[0043] S5: Perform a credit risk prediction task using the locally updated model after incremental update and output the prediction results.
[0044] For a client (such as a bank), when facing a loan application, the preprocessed credit samples are used as input, and the latest locally updated model through a dynamic update method is used to predict the loan sample to be borrowed, predicting whether the loan will default.
[0045] The method of the present invention realizes the personalized optimization of the model and the continuous improvement of the credit risk prediction ability through the combination of federated learning and continuous learning. First, by initializing the deep neural network model, it provides basic support for credit risk prediction; subsequently, the client participates in federated learning by uploading the model weights layer by layer, and uses the optimized global parameters to update the parameters of the corresponding layer of the local model, ensuring that the model adapts to the local data characteristics while being globally optimized. On this basis, a continuous learning strategy is adopted, and the local model is incrementally updated based on the data in the new stage, dynamically adapting to data changes, timely capturing new abnormal patterns, and at the same time retaining the memory of the old data. It can be seen that the present invention innovatively constructs a federated continuous learning framework, which not only realizes the continuous update of the model in the time dimension, but also effectively integrates heterogeneous data from different financial institutions or data sources through the multi-client knowledge sharing mechanism in the space dimension, realizing the effective group intelligence collaborative credit risk prediction of resources.
[0046] Compared with the traditional centralized learning method, the method of the present invention has strong data privacy protection. Specifically, the present invention adopts a federated learning mechanism of "data does not leave the local" to ensure the security and privacy of each client's data. By sharing model parameters instead of raw data, the efficient transfer and fusion of knowledge are realized, significantly enhancing data privacy protection. Compared with the traditional data sharing method, the present invention realizes the efficient transfer and fusion of knowledge by sharing model parameters without sharing raw data, significantly enhancing data privacy protection, meeting the requirements of increasingly strict privacy regulations. This method not only protects data privacy but also improves the generalization ability of the model.
[0047] While protecting data privacy, the advantages of distributed data are fully utilized, that is: training the local model using local data under the federated learning architecture can make full use of local characteristics, and then make the local model have personalized adaptability to meet the specific needs of different financial institutions or regions, significantly improving the generalization ability and prediction accuracy of the model.
[0048] Furthermore, through the continuous learning mechanism, the dynamic adaptation ability of the model is enhanced. Specifically, through the continuous learning mechanism of the present invention, the model parameters can be updated in real time to better adapt to new data distributions and business requirements. This method significantly improves the dynamic adaptation ability of the model in complex financial environments. The model can adapt to new data distributions and business requirements in real time, effectively capture the latest credit risk changes, reduce the risk of model obsolescence, and improve the practicality and robustness of the model. Compared with traditional static models, the dynamic adaptation ability of the present invention significantly improves the performance of the model in practical applications, reduces the risk of model obsolescence, and improves the practicality and robustness of the model.
[0049] In addition, the federated continuous learning framework of the method of the present invention has good scalability and can flexibly handle data and business requirements of different scales. For example, it is not only applicable to credit risk prediction, but also applicable to various application scenarios such as financial fraud detection and intelligent investment decision-making, with high practicality and economic value. Compared with existing solutions for single application scenarios, the present invention can be flexibly applied to a variety of financial business scenarios through a general federated continuous learning framework, with higher practicality and economic value. In addition, the framework design of the present invention is simple and flexible, easy to expand to more clients and application scenarios, and has broad applicability and scalability.
[0050] In one example, the local model uses elastic weight consolidation loss, or cross-entropy loss and elastic weight consolidation (EWC) loss to optimize the model parameters. In this example, the loss function of the local model includes cross-entropy loss and EWC loss, where the cross-entropy loss function is used to measure the difference between the model prediction output and the true label, and thus evaluate the prediction performance of the model. The cross-entropy loss function of the c-th client The calculation expression is: ; where 2 is the number of classes; is the true label (usually represented in one-hot encoding, with the target class corresponding to 1 and the rest being 0); is the probability distribution predicted by the model.
[0051] The elastic weight consolidation loss measures the importance of parameters by introducing the Fisher information matrix and adds a regularization term in the optimization process to prevent large changes to important parameters, thereby preventing the model from forgetting the knowledge of old tasks when learning new tasks. EWC ensures the stability of these parameters in subsequent training by assigning larger weights to important parameters, thus achieving continuous accumulation and consolidation of knowledge in multi-task learning scenarios. Specifically, the elastic weight consolidation loss The calculation expression is: ; wherein, represents the weight coefficient of the regularization term; is the diagonal element of the Fisher information matrix; is the current parameter; is the best parameter obtained by training on the old task. In this example, the elastic weight consolidation method is used to extract the parameter knowledge of the model, and there is no need to save the real credit samples in the previous stage. The elastic weight consolidation method takes the importance of the parameters in the model as knowledge, and the information transmitted between models is the information of the importance of model parameters. These knowledges will be transmitted to the next stage, thereby restricting the change range of important parameters to guide the next stage to refer to these knowledges when training the model.
[0052] During the training process, the total loss function of each local client can be expressed as the weighted sum of the cross-entropy loss and the EWC loss, and the calculation expression is: ; wherein, represents the total loss function.
[0053] In one example, after the client applies the optimized global parameters to the corresponding layer of the local model, it further includes:
[0054] Freeze the parameters of the corresponding layer (keep the weights of the matched federated layer unchanged), and update the parameters of the corresponding subsequent layer of the local model according to the global parameters of the subsequent layer sent back. Preferably, the parameters of the subsequent layer of the model are updated layer by layer (progressively) until the parameters of all layers of the model are updated, so as to maintain the adaptability of the local model to local data.
[0055] In one example, after the local client receives the updated global weights sent back by the server, it will compare these weights with the original weights before its own upload, and according to the characteristics and requirements of the local data, selectively adjust and optimize the global weights, retain the weight information highly relevant to the local data, and at the same time blur or regularize other less important weights to ensure that the model can still adapt to the characteristics of the local data while being globally optimized, so as to achieve personalized protection.
[0056] Specifically, as Figure 3 shown, after applying the optimized global parameters to the corresponding layer of the local model, it includes:
[0057] Evaluate the importance of the local model parameters, freeze the parameters with the top P% importance ranking in the local model, and use the optimized global parameters to replace the 1 - P% parameters with the lower ranking in the local credit risk prediction model, where P is the parameter importance ratio threshold.
[0058] More specifically, after receiving the global weights, the client performs personalized processing, including the following sub-steps:
[0059] (1) Calculate the importance weights of each parameter : ; Among them, is the loss function, which is used to measure the difference between the predicted value and the true value of the model; represents the parameter label; is the total number of training steps, is the training step label; is the th parameter of the th layer of the local model of the th client; represents that when the parameter changes slightly, the change rate of the loss function . During the training process, by calculating the partial derivative and using optimization algorithms such as gradient descent, the model parameters can be updated to minimize the loss function .
[0060] (2) Organize the importance weights of all parameters into a matrix , which is used to guide the subsequent parameter update: ; Among them, is the total number of parameters.
[0061] (3) Freeze the top P% of the parameters according to the importance weights, and impose regularization constraints on the remaining 1 - P% of the parameters. The weight of the th layer after the client update can be expressed as: ; ; Among them, represents personalized processing; represents the th parameter of the th layer of the global model; represents the learning rate; represents the regularization loss function, and its calculation expression is: ; Among them, represents the weight coefficient of the regularization term; After the client updates the weight of the th layer, continue to train the subsequent layers on the local data until the last layer .
[0062] After receiving the updated global weights, the example client performs personalized protection and regularization constraint processing to ensure the adaptability of the model to local data while protecting the key parameters (the first P% parameters) of the local model and avoiding the negative impact of global updates on the performance of the local model. At the same time, by restricting the variation range of the important parameters of the local model, an efficient model update of resources is achieved, avoiding frequent global updates and a large amount of data transmission, and significantly reducing the computing and communication costs, especially suitable for resource-constrained scenarios.
[0063] In one example, after sending the optimized global parameters to each client, it further includes:
[0064] Performing weighted average processing on the final weights of the global model according to the class distribution of local lending sample data to update the global model parameters.
[0065] Specifically, steps S1 - S3 are personalized federated collaboration steps. Uploading any layer of parameters of the local model, the server calculates the corresponding layer of global parameters and sends the global parameters back to the local model, which is defined as one round of communication. In each round of communication, weighted average processing is performed on the final global parameters (final weights) of the global model according to the class distribution of local lending sample data to adapt to the global data distribution. The calculation expression is: ; Among them, represents the training process; represents the client 's local dataset;
[0066] After each round of communication, the client performs weighted average on the global model according to the class distribution of local data to construct a model that better adapts to the global data distribution. Assuming that the weight ratio of each client is , then the final global parameters (final weights) of the global model through weighted average processing are: ; Among them, can be adjusted according to the client data volume or other factors.
[0067] Optionally, the above update of global model parameters includes the following sub-steps: (1) The client uploads the class distribution information to the server; (2) The server calculates the weight ratio of each client according to the class distribution information provided by all clients. For example, if the proportion of default clients in the local data of a certain client is relatively high, the server may assign higher weights to the model parameters of this client.
[0068] (3) The server uses the calculated weight ratios to perform weighted averaging on the model parameters of all clients to generate new global model parameters.
[0069] Alternatively, the above global model parameter update includes the following sub-steps: (1) Each client independently calculates the weight ratio according to the class distribution of the local data; for example, the client can assign a weight ratio to the local model parameters according to the proportion of default customers in the local data.
[0070] (2) The client uploads the weighted model parameters to the server; (3) The server aggregates the weighted parameters uploaded by all clients to generate new global model parameters.
[0071] Through the above final global parameter update method, this example solves the problem of performance degradation caused by directly averaging weights in traditional federated learning, and further improves the adaptability of the model to local data and the global optimization ability.
[0072] Optionally, combining the above examples, at this time, personalized federated learning includes the following steps:
[0073] (1) Initialize the global model and local models. In each round of federated learning, the server sends requests to all participating clients k to collect the weight information of the current layer of each client. These weights are the model parameters trained by the client on the local data, reflecting the characteristics and patterns of the local data.
[0074] (2) After the server receives the weights of all clients, it analyzes these weights through a matching algorithm. The purpose of the matching algorithm is to identify the similarities between the weights of different clients, so as to determine the contribution degree of each client's weight to the global model. Based on the matching results, the server calculates the global weights and distributes them back to all clients. The global weights are a comprehensive representation of the weights of all clients, aiming to balance the contributions of different clients while optimizing the performance of the global model.
[0075] (3) After receiving the global weights, client k will perform personalized adjustment on the global weights according to the characteristics of the local data. The purpose of this process is to protect the key parameters of the local model and avoid the negative impact of global updates on the performance of the local model. The client selectively retains or adjusts the global weights by analyzing the importance of the local data to ensure the adaptability of the model to the local data. After completing the personalized adjustment, the client gradually optimizes the subsequent layer parameters of the model in a step-by-step manner until all layer parameters of the model are updated.
[0076] In one example, such as Figure 4As shown, the local model is incrementally updated based on the training dataset of the current stage, that is, the local model is continuously trained, including:
[0077] When the set model update period arrives, preprocess the newly added credit sample data available for training during the update period to obtain the training dataset for the current model training stage; if it is the first training stage, preprocess the currently existing credit samples to construct the local model corresponding to the first training round.
[0078] If it is not the first training stage, use the training dataset of the current round to optimize the parameters of the local model obtained from the previous round of training. During the incremental training process, use the parameter knowledge extracted from the local model obtained from the previous training stage to optimize the parameters, obtain the local model corresponding to the current training stage, and extract the parameter knowledge of the model.
[0079] In this example, by incrementally updating the local model, not only the adaptability of the model to new data is improved, but also the time cost and computational resource consumption of model retraining are reduced. At the same time, the present invention also takes into account the data imbalance problem, and processes the data imbalance problem through data augmentation techniques or other methods, significantly improving the recognition ability of the model on minority class samples and further enhancing the robustness and reliability of the model.
[0080] Furthermore, using the parameter knowledge extracted from the local model obtained from the previous training stage to optimize the parameters specifically includes: comparing the parameters of the model in the current training stage with the parameter knowledge extracted from the local model obtained from the previous stage, incorporating the difference between the two into the loss function for optimizing the model parameters in the current stage, and making the parameters with larger weights in the credit risk prediction model obtained from the previous stage maintain larger weights in the loss function of the current stage. When minimizing the loss function of the current stage, since the difference between the new and old stage parameters is incorporated into the consideration of the loss function, the change range of the parameters is restricted. The more important parameters in the previous stage have larger weights in the current loss function and have a greater impact on the loss function of the current stage compared to general parameters. By minimizing the loss function, strict restrictions on the change range of different parameters are achieved, thereby realizing the retention of the knowledge learned in the previous stage.
[0081] Taking two stages 、 as an example, the total training dataset of the two stages contains the training datasets corresponding to the two stages respectively and . After constructing the local model of stage , use the EWC method to transfer the model knowledge to stage . The weight information of the model parameters As knowledge, it is in the stage extracted by EWC and transmitted to the stage , the restricted stage parameter weight information The specific principle of constructing the loss function is introduced as follows: ① According to Bayes' formula, using the prior probability , marginal probability , and conditional probability , the posterior probability can be calculated through the following formula : .
[0082] ② Since the dataset is completely divided into the stage and the stage , the posterior probability calculation formula is equivalent to: ; Among them, represents the conditional probability of the stage ; represents the posterior probability of the stage ; represents the marginal probability of the stage .
[0083] ③ Since each subset affects the posterior probability, based on the Laplace approximation method, the loss function of EWC can be expressed as: ; Among them, represents the total loss; represents the loss of the stage ; is the Laplace approximation information of the posterior probability calculated based on the dataset ; is the relative importance of the stage; represents the value of the th parameter in the current model parameters; represents the value of the parameter learned on the stage .
[0084] ④ When constructing the model in the stage , the loss of the parameter in the stage is a part of the total loss, and the change range of the parameter relative to the parameter in the stage is another part of the total loss. The finally selected parameter is the parameter that makes achieve the minimum value. In the ideal case, the stage has better performance; at the same time stage Among the more important parameters in the current stage, the change range is smaller compared to the current parameters.
[0085] With the help of a continuous learning framework, transfer the knowledge contained in the existing model to assist in the next-stage modeling. In the credit risk prediction scenario, to balance the two factors of model performance and prediction accuracy, the elastic weight consolidation, a regularization-based method, is adopted to achieve continuous learning. This method transfers the knowledge learned in the previous stage to the next stage by suppressing the change range of important parameters, providing a solution for improving the performance of the local model in a complex financial environment.
[0086] In one example, after applying the optimized global parameters to the corresponding layer of the local model, freezing the parameters of the corresponding layer, and completing the parameter update of all layers of the local model, it further includes:
[0087] Judge whether the set local model update period is reached. If so, judge whether the update times of the local model are an integer multiple of the constant S. If so, execute steps S1 - S3, that is, perform personalized federated learning; if not, perform an incremental update on the local model and execute the credit risk prediction task.
[0088] Among them, S represents the condition for triggering the aggregation of the global model after a certain number of updates of the local model. For example, if S is 5, it means that after every 5 updates of the local model, the client will upload the parameters of the local model to the server for global model aggregation.
[0089] Furthermore, the model period here can be a fixed time period (such as half a day or one day). Since the number of loan sample data obtained within each fixed time period may vary greatly, for example, the number of newly added loan sample data in a certain period is less, while the number of newly added loan sample data in a certain period is more. Then, in order to update the local model in a timely manner to obtain a more accurate prediction, the current available loan sample data for incremental training reaching a certain data volume (threshold data volume) can be used as an update period.
[0090] This example determines whether to perform personalized federated learning or only incremental update by setting the local model update period and judging whether the update times of the local model are an integer multiple of S, ensuring the timeliness and efficiency of model updates, and using the dynamically updated local model to output credit risk prediction results to provide decision-making support for financial institutions.
[0091] Combining the above examples, the preferred example of the method of the present invention is obtained, including the following steps: S100: Initialize the local model and the global model; S200: The client uploads the weights of all layers of the local model; S300: The server aggregates the weights of each layer of the local model uploaded by the client based on the matching algorithm, and then optimizes the global weights of each layer of the global model; S400: The client performs personalized federation on the global weights sent back by the server; S500: Incrementally update the credit risk prediction model based on the training data set of the current stage; S600: Execute the credit risk prediction stage and output the prediction result; S700: Determine whether the set local model update period has been reached. If so, determine whether the number of local model updates is an integer multiple of S. If so, execute steps S200 - S400 for personalized federation. Otherwise, return to step S500 to start a new stage. After all training is completed, execute step S600.
[0092] During the process of training the local model of the present invention, the local model and the global model are initialized, and then each round of training the model is taken as a training stage. Among them, for the first training stage, the existing loan sample data is directly preprocessed to obtain the corresponding training data set. Among them, data preprocessing is used to clean, transform and prepare the original data to make the data suitable for model training, including handling missing values, outliers and duplicate values, selecting the most relevant features, scaling and encoding the features, and performing data transformation and dimensionality reduction, so as to improve the training efficiency and generalization ability of the model. For subsequent other training stages, the corresponding training data set is obtained by preprocessing the newly added credit sample data within the update period (the time interval between the previous training round and the current training round). And after each round of training is completed, the parameter knowledge of the model is extracted and passed to the next stage. In the process of training the model in the next stage, parameter optimization is carried out by combining the parameter knowledge extracted from the credit risk prediction model obtained in the previous training stage, so as to guide the update of the model, thereby realizing the continuous learning of the model. When the model is applied, the federated idea is incorporated, that is, at the beginning of each round of stage, it is judged whether the set local model update period is reached. If the update period is reached, the system further judges whether the local model has been updated an integer multiple of S times. If it is an integer multiple of S times of update, return to step S200 to re-perform personalized federated learning to further optimize the local model. In the process of layer-by-layer uploading the model weights on the client side for personalized federation, through an innovative personalized federated learning mechanism, the balance between global optimization and local personalization of the model is achieved. Specifically, the client first uploads the current layer weight of its own model to the server; after the server receives the weights of all clients, it analyzes the similarity of these weights through a matching algorithm and calculates the weights of the global model. Subsequently, the server distributes the global weights back to each client. After receiving the global weights, the client makes a personalized adjustment to the global weights in combination with the characteristics of the local data. Through this strategy of layer-by-layer matching aggregation and personalized federation, the system not only improves the adaptability of the model to local data, but also significantly improves the communication efficiency and reduces the number of communication rounds.
[0093] Based on the above method, the present invention not only significantly improves the accuracy, adaptability and flexibility of credit risk prediction, but also effectively solves the problems of data privacy protection, model performance degradation, low communication efficiency and data imbalance existing in the prior art, and provides a more efficient, reliable and practical solution for the field of credit risk prediction.
[0094] The above specific implementation manners are detailed descriptions of the present invention. It cannot be determined that the specific implementation manners of the present invention are only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions and substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A credit risk prediction method based on personalized federated continuous learning, characterized in that: The method comprises the following steps: Each client uploads all layer parameters of its own local model to the server; the local model is a credit risk prediction model trained based on local loan sample data; The server receives all layer parameters of the local model uploaded by each client, and analyzes all layer parameters of each local model layer by layer through a matching algorithm to obtain the contribution of each layer parameter of each local model to the global parameters of the corresponding layer of the global model, and then optimizes the global parameters of each layer of the global model based on the contribution, and sends the optimized global parameters to each client; The client receives the optimized global parameters and applies the optimized global parameters to the corresponding layers of the local model; Incrementally update the local model based on the training dataset of the current stage; Use the local model that has been incrementally updated to perform credit risk prediction tasks and output the prediction results.
2. The credit risk prediction method based on personalized federated continuous learning according to claim 1 is characterized in that: The step of applying the optimized global parameters to the corresponding layer of the local model includes: Evaluate the importance of local model parameters, freeze the top P% parameters in the local model, and use the optimized global parameters to replace the bottom 1-P% parameters in the local credit risk prediction model, where P is the parameter importance ratio threshold.
3. The credit risk prediction method based on personalized federated continuous learning according to claim 1 is characterized in that: After sending the optimized global parameters to each client or applying the optimized global parameters to the corresponding layer of the local model, the method further includes: The final global parameters of the global model are weighted averaged according to the category distribution of the local loan sample data to update the global model parameters.
4. The credit risk prediction method based on personalized federated continuous learning according to claim 1 is characterized in that: The local model uses elastic weight consolidation loss, or cross entropy loss and elastic weight consolidation loss to optimize model parameters.
5. The credit risk prediction method based on personalized federated continuous learning according to claim 1 is characterized in that: The incremental updating of the local model based on the training data set in the current stage includes: During the update cycle, the newly added loan sample data is preprocessed to obtain the training data set of the current round, and the local model obtained by the previous round of training is incrementally trained using the training data set of the current round. During the incremental training process, the parameter knowledge extracted from the local model obtained by the previous training round is used to perform parameter optimization to obtain the local model corresponding to the current training round; Utilize the parameter knowledge extracted from the local model obtained from the previous training round to perform parameter optimization, including: The parameters of the local model in the current round of training task are compared with the parameter knowledge extracted from the local model obtained from the previous round of training task, and the parameter differences obtained from the comparison are incorporated into the loss function of the model parameter optimization of the current round of model training task, so that the parameters with larger weights in the local model obtained in the previous round of training task have larger weights in the loss function of the current round of training task.
6. The credit risk prediction method based on personalized federated continuous learning according to claim 1, characterized in that: After applying the optimized global parameters to the corresponding layers of the local model, the method further includes: Determine whether the set local model update cycle has been reached. If so, determine whether the number of updates of the local model is an integer multiple of the constant S. If so, return to the step where each client uploads all parameters of its local model to the server until the optimized global parameters are applied to the corresponding layer of the local model; if not, perform incremental updates to the local model and execute the credit risk prediction task.
Citation Information
Patent Citations
Federal learning method and system based on client weight evaluation
CN114564746A
Credit risk prediction method based on continuous learning
CN116452320A
Abnormal transaction detection method based on continuous learning and knowledge base playback strategy
CN118940200A
Abnormal transaction detection method based on continuous learning and double-auto-encoder architecture
CN118967138A
Dynamic weight aggregation federal learning method and system for credit risk assessment
CN119831069A
Cited By
Closed-loop continuous learning method and system for traffic participant behavior prediction model
CN120950999A