Federal learning stability improvement method and system based on block chain and knowledge distillation

Through the federated learning method of blockchain and knowledge distillation, the problem of low matching between labelless public data and client data volume in the financial field is solved, and accurate evaluation and optimization of data heterogeneity and target heterogeneity are achieved, which improves the stability and reliability of federated learning.

CN120258091APending Publication Date: 2025-07-04ZHENGZHOU UNIV

Patent Information

Application Number
CN202510317899.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the financial field, the labelless public data for financial customers in the federated learning scenario does not match the client data volume, which leads to different feature representations and decision boundaries that may be learned during model training, and there are data heterogeneity problems.

Method used

Through the federated learning method based on blockchain and knowledge distillation, dual-model training is carried out to monitor heterogeneity changes in real time, obtain fault risk assessment values, and based on this, determine whether to send a single point of failure risk instruction, perform smart contract verification or complete federated learning optimization, and combine knowledge distillation technology to migrate the data characteristics of the personalized model to the global model.

Benefits of technology

The matching degree between the label-free public data and the client data volume is improved, the stability and reliability of the federated learning architecture are improved, data heterogeneity and target heterogeneity are solved, and the universality of the model is ensured to take into account the needs of personalized models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258091A_ABST
    Figure CN120258091A_ABST
Patent Text Reader

Abstract

The invention discloses a federated learning stability improvement method and system based on a block chain and knowledge distillation, and relates to the technical field of federated learning. The federal learning stability improvement method based on the block chain and the knowledge distillation comprises the following steps of performing double-model training; obtaining a fault risk assessment value; and verifying the smart contract. According to the invention, dual-model training is carried out through the obtained local private financial data of each participant client, the dual-model training result is uploaded to the constructed block chain data sharing platform, the fault risk assessment value is obtained and whether a single-point fault risk instruction is sent is determined, if yes, smart contract verification is carried out, and if not, the smart contract verification is carried out. And otherwise, completing federated learning optimization, so that the effect of improving the matching degree of the financial customer-oriented unlabeled public data and the client data volume in the federated learning scene is achieved, and the problem of low matching degree of the financial customer-oriented unlabeled public data and the client data volume in the federated learning scene in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of federated learning, and particularly to a method and system for improving the stability of federated learning based on blockchain and knowledge distillation. Background Art

[0002] With the rapid development of technologies such as the Internet of Things, big data, and artificial intelligence, federated learning, as an emerging machine learning paradigm, has gradually received wide attention. Blockchain technology is a decentralized and immutable distributed ledger technology. Introducing blockchain technology into federated learning can effectively solve the mutual trust problem among participating users and improve the security and reliability of the system. Knowledge distillation is a technology for model compression and acceleration. It transfers the knowledge of a large model (teacher model) to a small model (student model), thereby achieving the lightweight and high efficiency of the model. Introducing knowledge distillation technology into federated learning can reduce the complexity and computational amount of the model and improve the operation efficiency of the system while ensuring the model performance. In the future, with the continuous development and improvement of technologies, federated learning systems based on blockchain and knowledge distillation are expected to be widely applied and promoted in more fields.

[0003] In the prior art, by storing the model parameters of federated learning on the blockchain, the security and immutability of the model parameters can be ensured. At the same time, during the federated learning process, knowledge distillation technology is used to compress the complex knowledge of the large model into the small model, thereby reducing the size of model updates and communication costs, which helps to improve the efficiency and stability of federated learning.

[0004] For example, the client selection optimization method and device for an unstable federated learning scenario disclosed in the invention patent announcement with the publication number CN114841368B includes: S1, analyzing the unstable factors of unstable federated learning; S2, modeling the influence of the client set, client local data, and client local training status on the model training performance; S3, modeling the client selection problem, that is, modeling the influence of the unstable client set, unstable client local data, and unstable client local training status on client selection; S4, proposing a client selection method based on the upper confidence interval and greedy selection to select the optimal client combination.

[0005] For example, the patent application with the publication number: CN115511102A discloses a federated learning system and method based on fairness and reputation mechanism, including: S1. Publish training task information, accept the bid information submitted by participants based on the training task information, and create a corresponding database; S2. Calculate the fairness utility value and reputation utility value of the participants according to the record information in the database; S3. Obtain the set of selected participants according to the fairness utility value and reputation utility value; S4. Send the global model to the selected participants; S5. Aggregate the uploaded models of the new round to obtain the global model of the new round; S6. Verify the performance of the uploaded models of the new round, update the record information in the participant database corresponding to the uploaded models at the same time, and feedback the corresponding rewards to the participants according to the bid information; Repeat S2 to S6 until the global model reaches the expected effect or the federated learning reaches the predetermined number of rounds.

[0006] However, in the process of implementing the technical solution of the present invention in the embodiments of the present application, it is found that the above technology has at least the following technical problems:

[0007] In the prior art, in the financial field, high-quality labeled data is often scarce, and the labeling process is time-consuming and laborious. Secondly, the amount of client data often varies greatly. Some clients may have a large amount of unlabeled data, while some clients may only have a small amount of labeled data, which may cause the model to learn different feature representations and decision boundaries during the training process, and then lead to a high degree of heterogeneity in the data of each client (such as different financial institutions, banks or fintech companies). There is a problem that the matching degree between the unlabeled public data for financial customers and the amount of client data in the federated learning scenario is not high. Summary of the Invention

[0008] The embodiments of the present application provide a method and system for improving the stability of federated learning based on blockchain and knowledge distillation, which solve the problem that the matching degree between the unlabeled public data for financial customers and the amount of client data in the federated learning scenario is not high in the prior art, and realize the improvement of the matching degree between the unlabeled public data for financial customers and the amount of client data in the federated learning scenario.

[0009] The embodiment of the present application provides a method for improving the stability of federated learning based on blockchain and knowledge distillation, including the following steps: Step 1, perform dual-model training according to the obtained local private financial data of each participating party's client, and upload the results of the dual-model training to the constructed blockchain data sharing platform; Step 2, construct a federated learning architecture, and simultaneously monitor the changes in heterogeneity in the constructed federated learning architecture in real time to obtain a failure risk assessment value, where the failure risk assessment value is used to evaluate the degree of heterogeneity difference in the federated learning architecture; Step 3, judge whether to send a single-point failure risk instruction based on the obtained failure risk assessment value. If so, perform smart contract verification; otherwise, complete the federated learning optimization.

[0010] Further, the specific construction process of the federated learning architecture is as follows: When performing local updates, use knowledge distillation to transfer the data features of the personalized model to the global model; after the local updates are completed, perform dual-model aggregation based on the blockchain data sharing platform and then perform global model verification; after performing global model verification, obtain a federated learning architecture through a set of participating nodes, where the set of participating nodes is used to aggregate the global model obtained from the dual-model training and the processes of dual-model aggregation and global model verification.

[0011] Further, the specific steps for obtaining the global model verification include: encrypt the address of the local global model through a hash algorithm and upload it to the blockchain data sharing platform; obtain model aggregation nodes according to the reputation values of all participating nodes, and perform node aggregation based on the obtained model aggregation nodes to obtain the local global model; upload the locally aggregated global model to the blockchain data sharing platform and request verification; if the verification request is passed, prompt each participating node to request the latest local global model; if the verification request fails, prompt the data owner node, model training node, and verification node to request secondary dual-model training; the node aggregation means aggregating the local global models corresponding to the model aggregation nodes, aggregated data owner nodes, model training nodes, and verification nodes; the secondary dual-model training means combining the last locally aggregated global model that passed cross-validation with the personalized model and re-performing dual-model training.

[0012] Further, the specific process for obtaining the fault risk assessment value is as follows: Obtain the fault risk parameters of the model parameters during the dual-model training process. The fault risk parameters include the average update frequency, average update amplitude, and actual transaction frequency. The actual transaction frequency represents the transaction frequency of local private financial data on the blockchain network. Obtain the local objective loss function, global model weight factor, and local global model weight factor of the participant client. At the same time, combine the obtained fault risk parameters and the reference fault risk data in the database to obtain the fault risk assessment value. The reference fault risk data includes the reference update frequency, reference update amplitude, reference transaction efficiency, update frequency weight factor, update amplitude weight factor, transaction efficiency weight factor, and local objective loss function weight factor.

[0013] Further, the specific limiting expression of the fault risk assessment value is:

[0014]

[0015] In the formula, g is the number of the participant client, g = 1, 2,..., G, G is the total number of participant clients, e is the natural constant, FP represents the fault risk assessment value of heterogeneity in the federated learning architecture, q represents the global model weight factor, α1 represents the update frequency weight factor, A represents the average update frequency of the model parameters during the dual-model training process, A0 represents the reference update frequency, α2 represents the update amplitude weight factor, B represents the average update amplitude of the model parameters during the dual-model training process, B0 represents the reference update amplitude, α3 represents the transaction efficiency weight factor, C represents the actual transaction efficiency of local private financial data on the blockchain network, C0 represents the reference transaction efficiency, α4 represents the local objective loss function weight factor, F g (w g ) represents the local objective loss function of the g-th participant client, w g represents the local global model weight factor of the g-th participant client.

[0016] Further, the specific process for determining whether to send a single-point fault risk instruction based on the obtained fault risk assessment value is as follows: Determine whether the obtained fault risk assessment value is not less than the preset fault risk assessment value in the database. If the obtained fault risk assessment value is not less than the preset fault risk assessment value in the database, the blockchain data sharing platform sends a single-point fault risk instruction. If the obtained fault risk assessment value is less than the preset fault risk assessment value in the database, the blockchain data sharing platform sends a federated learning optimization completion instruction.

[0017] The embodiment of the present application provides a system for improving the stability of federated learning based on blockchain and knowledge distillation, including: a dual-model training module, a failure risk assessment value acquisition module, and a smart contract verification module; wherein, the dual-model training module is used to perform dual-model training based on the obtained local private financial data of each participating party client, and upload the results of the dual-model training to the constructed blockchain data sharing platform; the failure risk assessment value acquisition module is used to construct a federated learning architecture, and at the same time monitor the changes in heterogeneity in the constructed federated learning architecture in real time to obtain a failure risk assessment value, and the failure risk assessment value is used to evaluate the degree of difference in heterogeneity in the federated learning architecture; the smart contract verification module is used to determine whether to send a single-point failure risk instruction based on the obtained failure risk assessment value, and if so, perform smart contract verification, otherwise complete the federated learning optimization.

[0018] One or more technical solutions provided in the embodiment of the present application have at least the following technical effects or advantages:

[0019] 1. By performing dual-model training based on the obtained local private financial data of each participating party client, uploading the results of the dual-model training to the constructed blockchain data sharing platform, and at the same time obtaining the failure risk assessment value and determining whether to send a single-point failure risk instruction, and if so, performing smart contract verification, otherwise completing the federated learning optimization, it realizes more accurate avoidance of heterogeneity, and further realizes the improvement of the matching degree between the unlabeled public data and the client data volume for financial customers in the federated learning scenario, effectively solving the problem of low matching degree between the unlabeled public data and the client data volume for financial customers in the existing federated learning scenario.

[0020] 2. By using knowledge distillation to transfer the data features of the personalized model to the global model during local update, after the local update is completed, performing global model verification after dual-model aggregation based on the blockchain data sharing platform, and after global model verification, obtaining the federated learning architecture through the set of participating nodes, it realizes more accurate construction of the federated learning architecture, and further realizes more accurate optimization of the stability of the federated learning architecture.

[0021] 3. By obtaining the failure risk parameters of the model parameters during the dual-model training process, at the same time obtaining the local objective loss function, global model weight factor, and local global model weight factor of the participating party client, and combining the obtained failure risk parameters and the reference failure risk data in the database to obtain the failure risk assessment value, it realizes the improvement of the accuracy of obtaining the failure risk assessment value, and further realizes more accurate evaluation of the degree of difference in heterogeneity of the federated learning architecture. Description of the Drawings

[0022] Figure 1Flowchart of the method for improving the stability of federated learning based on blockchain and knowledge distillation provided by the embodiments of the present application;

[0023] Figure 2 Schematic diagram of the knowledge distillation process provided by the embodiments of the present application;

[0024] Figure 3 Flowchart of local update provided by the embodiments of the present application;

[0025] Figure 4 Schematic diagram of the structure of the system for improving the stability of federated learning based on blockchain and knowledge distillation provided by the embodiments of the present application;

[0026] Figure 5 Schematic diagram of the federated learning system based on blockchain and knowledge distillation provided by the embodiments of the present application;

[0027] Figure 6 Schematic diagram of the global model training process provided by the embodiments of the present application;

[0028] Figure 7 Comparison chart of the accuracy rates of the global models of each algorithm on all private test sets provided by the embodiments of the present application;

[0029] Figure 8 Comparison chart of the accuracy rates of the global model on each participating party's client under different weights provided by the embodiments of the present application. Detailed implementation manners

[0030] The embodiments of the present application provide a method and system for improving the stability of federated learning based on blockchain and knowledge distillation, which solve the problem that the matching degree between the unlabeled public data for financial customers and the client data volume in the federated learning scenario in the prior art is not high. By performing dual-model training on the local private financial data of each participating party's client and uploading the results of the dual-model training to the constructed blockchain data sharing platform, then constructing a federated learning architecture based on the blockchain data sharing platform, and simultaneously monitoring the changes in heterogeneity in the constructed federated learning architecture in real time to obtain a fault risk assessment value, and finally judging whether the blockchain data sharing platform sends a single-point fault risk instruction based on the obtained fault risk assessment value. If so, perform smart contract verification, otherwise complete the federated learning optimization, achieving an improvement in the matching degree between the unlabeled public data for financial customers and the client data volume in the federated learning scenario.

[0031] The technical solution in the embodiments of the present application for solving the problem that the matching degree between the unlabeled public data for financial customers and the client data volume in the above-mentioned federated learning scenario is not high has the following general idea:

[0032] Double-model training is performed on the local private financial data of each participating party's client, and the results of the double-model training are uploaded to the constructed blockchain data sharing platform. At the same time, a fault risk assessment value is obtained and it is judged whether to send a single-point fault risk instruction. If so, smart contract verification is performed; otherwise, federated learning optimization is completed, achieving the effect of improving the matching degree between unlabeled public data and client data volume for financial customers in the federated learning scenario.

[0033] To better understand the above technical solution, the above technical solution will be described in detail below in combination with the accompanying drawings of the specification and specific implementation manners.

[0034] As Figure 1 shown, it is a flowchart of the method for improving the stability of federated learning based on blockchain and knowledge distillation provided by an embodiment of the present application. The method for improving the stability of federated learning based on blockchain and knowledge distillation provided by an embodiment of the present application includes the following steps: Step 1, double-model training is performed on the local private financial data of each participating party's client, and the results of the double-model training are uploaded to the constructed blockchain data sharing platform; Step 2, a federated learning architecture is constructed based on the blockchain data sharing platform, and at the same time, the change situation of heterogeneity in the constructed federated learning architecture is monitored in real time to obtain a fault risk assessment value, where the fault risk assessment value is used to evaluate the difference degree of heterogeneity in the federated learning architecture, and the federated learning architecture is an architecture constructed based on the blockchain data sharing platform and combined with knowledge distillation; Step 3, based on the obtained fault risk assessment value, it is judged whether the blockchain data sharing platform sends a single-point fault risk instruction. If so, smart contract verification is performed; otherwise, federated learning optimization is completed.

[0035] Among them, the participating party's client includes a Global Model and a Personalized Model; the blockchain data sharing platform contains a blockchain network; the blockchain network is composed of a preset number of participating nodes; the participating nodes are the nodes corresponding to the participating party's client; the participating nodes include data owner nodes, model training nodes, model aggregation nodes, and verification nodes; heterogeneity includes data heterogeneity and target heterogeneity; smart contract verification is used to verify whether there is illegal intrusion in the participating party's client.

[0036] In this embodiment, the global model is used to learn the data features of the personalized model; the personalized model is used to fit the distribution characteristics of the local private financial data; the blockchain data sharing platform is built through blockchain technology; the participating nodes are the nodes corresponding to the participating party clients; data heterogeneity represents the data distribution differences between different participating party clients in the federated learning scenario; target heterogeneity represents the different requirements, service objectives, and functional requirements of different participating party clients for the global model and the personalized model in the federated learning scenario; the smart contract has predefined rules for data upload, model update, dual-model verification, and dual-model aggregation.

[0037] Specifically, the local private financial data includes but is not limited to customer information, transaction records, and credit history. One participating node corresponds to one participating party client. At the end of each round of dual-model training, only one participating node will be selected as the model aggregation node, and each participating node has a private dataset and the reputation value of the participating node itself.

[0038] It should be added that, as Figure 2 shown, it is a schematic diagram of the knowledge distillation process provided by the embodiment of the present application. Knowledge Distillation (KD) is a model compression technology that transfers the knowledge in a complex and high-performance teacher model to a smaller student model, thereby improving the performance of the student model. This method was initially proposed by Hinton et al. and has achieved success in many deep learning tasks. The loss function of the student model can be simplified as follows:

[0039] L student =L CE +D KL (P teacher IIP student );

[0040]

[0041] In the formula, L CE and D KL are the cross-entropy loss and the KL divergence loss respectively, P teacher and P student are the predicted values of the teacher model and the student model respectively, T is the hyperparameter average temperature, i represents the index of the sample number division corresponding to the participating party client, z is the logical value of the teacher model, z i represents the logical value of the teacher model under the divided index, and D KL (P teacher IIP student ) represents the KL divergence corresponding to the predicted values of the teacher model and the student model (the teacher model fits the student model).

[0042] The federated learning optimization method of the solution of this application can effectively balance the generality of the global model and the personalized needs of the participating party clients, can effectively solve data heterogeneity and target heterogeneity, and further improves the matching degree between the unlabeled public data and the client data volume for financial clients in the federated learning scenario.

[0043] Furthermore, the specific construction process of the federated learning architecture is as follows: when performing local updates, knowledge distillation is used to transfer the data features of the personalized model to the global model, and knowledge distillation has the advantages of knowledge transfer and model enhancement; after the local updates are completed, double-model aggregation is performed based on the blockchain data sharing platform and then global model verification is carried out; after the global model verification, the participating node set is obtained to form the federated learning architecture, and the participating node set is used to aggregate the global model obtained from the double-model training and the processes of double-model aggregation and global model verification.

[0044] As Figure 3 shown, it is the local update flow chart provided by the embodiment of this application. Local updates are used to make an update request for the global model and mark it to obtain the local global model; the update request means requesting the latest on-chain global model to update the global model when the personalized model remains unchanged; the on-chain global model means the global model stored on the blockchain network after being aggregated and updated a preset number of times (set by a preset person); the local global model means the global model used for double-model aggregation after local updates.

[0045] The specific steps for obtaining global model verification include: encrypting the address of the local global model through a hashing algorithm and uploading it to the blockchain data sharing platform; obtaining model aggregation nodes according to the reputation values of all participating nodes, and performing node aggregation based on the obtained model aggregation nodes to obtain the local global model; uploading the locally globally modeled node-aggregated model to the blockchain data sharing platform and requesting verification; node aggregation means aggregating the local global models corresponding to the model aggregation nodes, the aggregated data owner nodes, the model training nodes, and the verification nodes; if the verification request is passed, each participating node is prompted to request the latest local global model; if the verification request fails, the reputation value of the corresponding model aggregation node will be deducted points, and at the same time, the data owner nodes, the model training nodes, and the verification nodes are prompted to request secondary double-model training; secondary double-model training means combining the local global model that passed cross-verification last time with the personalized model and re-performing double-model training.

[0046] In this embodiment, the specific process of the verification request is as follows: submitting a verification request through a smart contract and attaching the hash value of the local global model and the reputation value of the model aggregation node; the verification node performs cross-verification on the hash value of the local global model and the reputation value of the model aggregation node respectively, and uploads the cross-verification results to the blockchain network.

[0047] It should be added that to effectively address the problem of model performance degradation caused by data heterogeneity and the objective heterogeneity between the participating party's client and the server replaced by the blockchain, this application adopts a federated average distillation dynamic optimization algorithm composed of three strategies: coexistence of dual models, local update, and dual model aggregation.

[0048] Among them, the specific process of dual model aggregation is as follows: The loss functions of the personalized model and the global model are as follows:

[0049]

[0050] In the formula, represents the cross-entropy loss of the personalized model, represents the cross-entropy loss of the global model, D KL represents the KL divergence loss, P1 and P2 respectively represent the prediction values of the personalized model and the global model corresponding to the blockchain network, D KL (P1IIP2) represents the KL divergence loss corresponding to the prediction values of the personalized model and the global model corresponding to the blockchain network, represents the personalized model, whose goal is to better fit the distribution of local private financial data, represents the global model, whose goal is to learn the knowledge of the personalized model and improve the generalization performance of the global model by achieving prediction consistency (distillation). After dual model aggregation, a more general global model, that is, the local global model, is obtained.

[0051] Since the local update of each participating party's client is not directly trained on the copy of the global model, and the key tasks of federated learning are different at different times, therefore, the weights of the parameters of the local global model from local private data and the personalized model are adjusted according to the communication rounds, and the loss function of the local global model is rewritten as follows:

[0052]

[0053] In the formula, represents the local global model, α and β represent the parameters controlling the proportion of knowledge from local private financial data or the personalized model. When α = 1 and β = 0, the federated average distillation dynamic optimization algorithm degenerates into FedAvg, P'1' and P'2' respectively represent the prediction values of the local personalized model and the local global model corresponding to the blockchain network, D KL (P'1'IIP'2') represents the KL divergence loss corresponding to the prediction values of the local personalized model and the local global model corresponding to the blockchain network.

[0054] The pseudocode of the federated average distillation dynamic optimization algorithm in this example is as follows:

[0055] Input: (T = 200, t, G = 10, g, E = 5, b)

[0056] Server (blockchain) executes:

[0057] 1. Initialization: global model global0

[0058] 2. for t ← 1 to T do

[0059] 3. for participating party client g ← 1 to G do

[0060] 4. Participating party client update:

[0061] 5. end for

[0062] 6. Select the dual - model aggregation of participating party clients:

[0063] 7. Dual - model verification

[0064] 8. end for

[0065] ClientUpdate:

[0066] 1. Initialization: local personalized model and global model:

[0067] 2. Global model update:

[0068] 3. for b ← 1 to E do

[0069] 4. Update using local private data

[0070] 5. end for

[0071] It should be understood that although this example takes into account the differences in the distribution of local private financial data corresponding to different participating party clients, it will reduce the contribution of the participating party client with a relatively small amount of local private financial data. Therefore, in this example, the aggregation weights of each participating party client are the same, that is That is, the contributions of each participating party client are the same and equally important. At the same time, it is also a protection of the privacy of local private financial data, thus achieving the stability and robustness of the global model in the case of a large distribution of local private financial data.

[0072] Further, the specific process for obtaining the fault risk assessment value is as follows: Obtain the fault risk parameters of the model parameters during the dual-model training process. The fault risk parameters include the average update frequency, average update amplitude, and actual transaction frequency. The actual transaction frequency represents the transaction frequency of local private financial data on the blockchain network. Obtain the local objective loss function, global model weight factor, and local global model weight factor of the participant client. At the same time, combine the obtained fault risk parameters and the reference fault risk data in the database to obtain the fault risk assessment value. The reference fault risk data includes the reference update frequency, reference update amplitude, reference transaction efficiency, update frequency weight factor, update amplitude weight factor, transaction efficiency weight factor, and local objective loss function weight factor.

[0073] Among them, the specific constraint expression of the fault risk assessment value is:

[0074]

[0075] In the formula, g is the number of the participant client, g = 1, 2,..., G, G is the total number of participant clients, e is the natural constant, FP represents the fault risk assessment value of heterogeneity in the federated learning architecture, q represents the global model weight factor, α1 represents the update frequency weight factor, A represents the average update frequency of the model parameters during the dual-model training process, A0 represents the reference update frequency, α2 represents the update amplitude weight factor, B represents the average update amplitude of the model parameters during the dual-model training process, B0 represents the reference update amplitude, α3 represents the transaction efficiency weight factor, C represents the actual transaction efficiency of local private financial data on the blockchain network, C0 represents the reference transaction efficiency, α4 represents the local objective loss function weight factor, F g (w g ) represents the local objective loss function of the g-th participant client, w g represents the local global model weight factor of the g-th participant client.

[0076] It should be added that the specific steps for obtaining the local objective loss function are as follows: The loss function of local private financial data is: Given the local private financial data (X g ) generated from different distributions P g , Y g ) of G participant clients, the federated learning (FedAvg) of each participant client starts with obtaining the global model weight factor from the global model. Then, each participant client performs local updates and optimizes the local objective through the gradient descent method to obtain the local objective loss function: In the formula, δ represents the learning rate, represents the gradient operator.

[0077] After local update, the participating party client passes the local global model weight factor to the parameter server, and the parameter server aggregates it through weighted average to obtain the global model weight factor, that is In the formula, N represents the number of samples of all participating party clients, N g represents the number of samples corresponding to the g-th participating party client.

[0078] The smaller the local objective loss function (but not equal to 0), the higher the degree of distributed global optimization of the corresponding local private financial data, and the smaller the global model weight factor, the more convergent the global model in the dual-model training process.

[0079] In this embodiment, the reference update frequency, reference update amplitude, and reference transaction efficiency are respectively represented by the results of summing and averaging the historical update frequency, historical update amplitude, and historical transaction efficiency of the model parameters of the personalized model and the global model in the historical training process in the database; the actual transaction efficiency represents the ratio of the amount of local private financial data actually traded during the dual-model training period on the blockchain network to the duration of the dual-model training period, and the average update frequency and average update amplitude respectively represent the average values of the model parameter update frequency and update amplitude of the personalized model and the global model during the training process. Among them, the update frequency represents the number of changes of the model parameters during the dual-model training process, usually obtained through a counter, the update amplitude represents the amount of change that occurs each time the model parameters are updated, usually obtained by real-time monitoring through a sensor, and the amount of local private financial data actually traded is obtained through a counter.

[0080] The update frequency weight factor, update amplitude weight factor, transaction efficiency weight factor, and loss function weight factor are respectively the influence degrees of the preset average update frequency, average update amplitude, actual transaction efficiency, and loss function in the database on the process of obtaining the fault risk assessment value. Specifically, the database stores the preset weight factors corresponding to the average update frequency, average update amplitude, actual transaction efficiency, and loss function. There is a preset mapping relationship between these weight factors and the average update frequency, average update amplitude, actual transaction efficiency, and loss function. This mapping relationship can be one-to-one or many-to-one. In practical applications, the real-time average update frequency, average update amplitude, actual transaction efficiency, and loss function can be input into this mapping relationship to quickly obtain the corresponding weight factors, providing an important quantitative index for evaluating the stability and reliability of the federated learning architecture, and then calculating the fault risk assessment value more accurately.

[0081] In this example, the value ranges of the update frequency weight factor, update amplitude weight factor, transaction efficiency weight factor, and loss function weight factor are all limited between 0 and 1, and the sum of the four is 1.

[0082] The aforementioned database was established before the design of the method for improving the stability of federated learning based on blockchain and knowledge distillation, and is used to store various types of set data. The database includes, but is not limited to, reference update frequency, reference update amplitude, reference transaction efficiency, preset fault risk assessment value, and the dual-model training period. Various values therein are directly set by technical personnel. Among them, the setting basis of the preset fault risk assessment value can be determined according to the actual application scenario of the federated learning architecture. For example, the preset fault risk assessment value is represented by the result of summing and averaging the historical fault risk assessment values of heterogeneity in the federated learning architecture in the database. In addition, various values in the database can be set and fine-tuned by technical personnel according to actual debugging.

[0083] It should be understood that the fault risk assessment value increases as the global model weight factor, local objective loss function, update frequency deviation (i.e., |A - A0|), update amplitude deviation (i.e., |B - B0|), and transaction efficiency deviation (i.e., |C - C0|) increase. Among them, the local objective loss function also affects the value of the update amplitude deviation. When the local objective loss function increases, it means that the deviation between the prediction result of the global model and the real situation increases. In order to reduce the loss as soon as possible, this usually means that the global model needs a larger adjustment to improve its performance, resulting in an increase in the update amplitude deviation.

[0084] The update frequency deviation also indirectly affects the value of the transaction efficiency deviation. When the update frequency deviation of the model parameters increases, this may affect the real-time performance of the global model. For example, if the update frequency decreases, the global model may not be able to adapt to market changes in a timely manner, resulting in a lag in transaction decisions. If the update frequency increases, it may increase communication and computing costs and reduce the actual transaction efficiency.

[0085] By considering the above indirect influence mechanism, the composition and change reasons of the fault risk assessment value can be more comprehensively understood, which helps to balance the performance, update frequency, and transaction cost of the global model, and then improves the matching degree of unlabeled public data and client data volume for financial customers in the federated learning scenario, effectively solving the problem of low matching degree of unlabeled public data and client data volume for financial customers in the existing federated learning scenario.

[0086] Further, the specific process of determining whether to send a single point of failure risk instruction based on the obtained failure risk assessment value is as follows: Determine whether the obtained failure risk assessment value is not less than the preset failure risk assessment value in the database. If the obtained failure risk assessment value is not less than the preset failure risk assessment value in the database, it indicates that the degree of heterogeneity in the federated learning architecture does not meet the expected requirements of federated learning optimization, and the blockchain data sharing platform sends a single point of failure risk instruction. If the obtained failure risk assessment value is less than the preset failure risk assessment value in the database, it indicates that the degree of heterogeneity in the federated learning architecture meets the expected requirements of federated learning optimization, and the blockchain data sharing platform sends a federated learning optimization completion instruction.

[0087] In this embodiment, the single point of failure risk instruction is used to prompt the preset personnel to check the model update process. In the federated learning architecture, a certain participating node has no redundancy or backup mechanism, resulting in the failure of the participating node, that is, a single point of failure (Single Point of Failure, abbreviated as SPOF). In this example, by comparing the numerical change relationship between the failure risk assessment value and the preset failure risk assessment value, the refined management and control of the failure risk in the federated learning architecture are realized. This helps to timely discover and handle potential risks, avoid the occurrence of single points of failure, reduce manual intervention and decision-making time, improve the response speed and accuracy, and thus improve the stability and reliability of the federated learning architecture.

[0088] As Figure 4 shown, it is a schematic structural diagram of a federated learning stability improvement system based on blockchain and knowledge distillation provided by an embodiment of the present application. The federated learning stability improvement system based on blockchain and knowledge distillation provided by an embodiment of the present application includes: a dual-model training module, a failure risk assessment value acquisition module, and a smart contract verification module. Among them, the dual-model training module is used to perform dual-model training according to the obtained local private financial data of each participating party's client and upload the results of the dual-model training to the constructed blockchain data sharing platform. The failure risk assessment value acquisition module is used to construct a federated learning architecture based on the blockchain data sharing platform and simultaneously monitor the change of heterogeneity in the constructed federated learning architecture in real time to obtain a failure risk assessment value. The failure risk assessment value is used to evaluate the degree of heterogeneity in the federated learning architecture. The federated learning architecture is an architecture constructed according to the blockchain data sharing platform and combined with knowledge distillation. The smart contract verification module is used to determine whether the blockchain data sharing platform sends a single point of failure risk instruction based on the obtained failure risk assessment value. If so, it performs smart contract verification; otherwise, it completes the federated learning optimization.

[0089] In this embodiment, as Figure 5As shown in the figure, it is a schematic diagram of a federated learning system based on blockchain and knowledge distillation provided by an embodiment of the present application. Blockchain technology can well solve the hidden danger of single-point failure of the central server, and knowledge distillation can effectively enhance the performance of the global model. Existing research on federated learning has made great progress in solving data heterogeneity, but there are still deficiencies in solving target heterogeneity and cannot solve or alleviate the conflict between the personalized needs of clients and the generality of the global model. The specific contributions of the solution of the present application are as follows:

[0090] (1) A dual-model strategy is proposed, enabling participating client parties, namely financial institutions and banks, to simultaneously have a personalized model and a global model. The global model participates in optimization during the aggregation process, while the personalized model does not need to be replaced with the current global model in each iteration, thus improving the flexibility of dual-model training and helping to better adapt to specific task requirements.

[0091] (2) By adopting knowledge distillation technology, the personalized model can transfer knowledge to the global model, where the weights of the local global model loss function change dynamically, enabling the global model to learn the data characteristics of all clients and effectively improving the generalization performance and generality of the global model.

[0092] (3) Blockchain technology is used to overcome the hidden danger of single-point failure of the central server, ensuring the security and transparency of the global model transfer process. Finally, experiments verify the effectiveness of the proposed federated learning dynamic optimization algorithm based on blockchain and knowledge distillation in improving the generalization accuracy and robustness of the global model, while maintaining the high accuracy of the personalized model.

[0093] In this application, the dataset is divided into 4 levels according to the degree of Non-IID (IID, Non-IID (Dirichlet coefficients are 0.5, 0.1, 0.05) respectively), and different levels of datasets are further divided into 10 parts according to the number of participating parties. The data volume of each participating party is random, and there are extreme cases.

[0094] As Figure 6 shown in the figure, it is a schematic diagram of the global model training process provided by an embodiment of the present application. FedAvg and FedKD are used as two baselines, and the global models of the three algorithms are compared under four data distributions such as IID and Non-IID (0.5, 0.1, 0.05). Due to the different ultimate goals of the participating client parties and the central server replaced by the blockchain in federated learning, the participating client parties aim to train a personalized model that conforms to their own data distribution for themselves, while the server aims to train a single, generalizable model for all clients and new participants.

[0095] This application tested the performance of the global model in the algorithm of this application and the global models in FedAvg and FedKD on a private test set. As Figure 7 shown, it is a comparison chart of the accuracy of the global models of each algorithm provided by the embodiments of this application on all private test sets. According to Figure 7 the data analysis in, the algorithm proposed in this application is superior to other algorithms in the performance of the global model. Especially when there are extreme distributions or large differences in the data volume among the participating party clients, the average accuracy of the algorithm in this application is the highest. In addition, the smoothness and stability of the result curve also show superiority, indicating that the global model of the algorithm in this application has stronger stability and robustness when facing clients with large data heterogeneity differences. These results further verify the effectiveness of the algorithm in this application in dealing with heterogeneity.

[0096] As Figure 8 shown, it is a comparison chart of the accuracy of the global model on each participating party client under different weights provided by the embodiments of this application. The loss function of the local global model includes the cross-entropy loss of the local data and the KL divergence loss of learning from the personalized model. Therefore, the weights from the latter two loss functions have a great impact on the local model. Therefore, comparative experiments are carried out with different (α, β) values. The values and experimental results are as Figure 8 shown. Taking α = 0.5 and β = 0.5 as the benchmark, when β is decreased and α is increased, it is observed that the accuracy of the global model increases, the variance decreases, and the fluctuation of the accuracy is also suppressed; when α is decreased and β is increased, the accuracy of the global model decreases, the variance becomes larger, and the fluctuation of the accuracy becomes larger. The algorithm in this application selects dynamically changing weight values. As the number of iterations increases, α is gradually increased and β is gradually decreased, improving the accuracy of the global model while enhancing the generalization ability and robustness of the model.

[0097] In summary, the embodiments of this application perform dual-model training on the local private financial data of each participating party client obtained, upload the results of the dual-model training to the constructed blockchain data sharing platform, and at the same time obtain the fault risk assessment value and determine whether to send a single-point fault risk instruction. If so, intelligent contract verification is carried out; otherwise, federated learning optimization is completed, thereby achieving more precise avoidance of heterogeneity, and further improving the matching degree between the unlabeled public data and the client data volume for financial customers in the federated learning scenario, effectively solving the problem of low matching degree between the unlabeled public data and the client data volume for financial customers in the existing federated learning scenario.

[0098] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0099] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0100] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks by the instructions executed on the computer or other programmable device.

[0102] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0103] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. A method for improving the stability of federated learning based on blockchain and knowledge distillation, characterized in that It includes the following steps: Step 1: Perform dual-model training based on the obtained local private financial data of each participating party's client, and upload the results of the dual-model training to the constructed blockchain data sharing platform; Step 2: Construct a federated learning architecture, and simultaneously monitor the changes in heterogeneity in the constructed federated learning architecture in real time to obtain a fault risk assessment value, which is used to evaluate the degree of heterogeneity difference in the federated learning architecture; Step 3: Based on the obtained fault risk assessment value, determine whether to send a single-point fault risk instruction. If so, perform smart contract verification; otherwise, complete the federated learning optimization.

2. The method for improving the stability of federated learning based on blockchain and knowledge distillation according to claim 1, wherein The participating party's client includes a global model and a personalized model; The blockchain data sharing platform contains a blockchain network; The blockchain network consists of a preset number of participating nodes; The participating nodes include data owner nodes, model training nodes, model aggregation nodes, and verification nodes.

3. The method for improving the stability of federated learning based on blockchain and knowledge distillation according to claim 1, characterized in that, The heterogeneity includes data heterogeneity and target heterogeneity; The data heterogeneity represents the data distribution difference between different participating party clients in the federated learning scenario; The smart contract verification is used to verify whether there is an illegal intrusion in the participating party's client.

4. The method for improving the stability of federated learning based on blockchain and knowledge distillation according to claim 1, wherein The specific construction process of the federated learning architecture is as follows: When performing local updates, use knowledge distillation to transfer the data features of the personalized model to the global model; After the local update is completed, perform dual-model aggregation based on the blockchain data sharing platform and then perform global model verification; After performing the global model verification, obtain a federated learning architecture through a set of participating nodes. The set of participating nodes is used to aggregate the global model obtained from the dual-model training and the processes of dual-model aggregation and global model verification.

5. The method for improving the stability of federated learning based on blockchain and knowledge distillation according to claim 4, wherein The local update is used to send an update request to the global model and mark it to obtain a local global model; The update request means that when the personalized model remains unchanged, request the latest on-chain global model to update the global model; The on-chain global model refers to the global model stored on the blockchain network after being aggregated and updated a preset number of times; The local global model refers to the global model used for dual-model aggregation after local update.

6. The method for improving the stability of federated learning based on blockchain and knowledge distillation according to claim 4, wherein The specific obtaining steps of the global model verification include: Encrypt the address of the local global model through a hash algorithm and upload it to the blockchain data sharing platform; Obtain model aggregation nodes according to the reputation values of all participating nodes, and perform node aggregation based on the obtained model aggregation nodes to obtain a local global model; Upload the locally aggregated global model after node aggregation to the blockchain data sharing platform and request verification; The node aggregation refers to aggregating the local global models corresponding to the model aggregation nodes, aggregated data owner nodes, model training nodes, and verification nodes; If the verification request is passed, prompt each participating node to request the latest local global model; If the verification request fails, prompt the data owner node, model training node, and verification node to request secondary dual-model training; The secondary dual-model training means combining the local global model that passed the cross-validation last time with the personalized model and re-performing dual-model training.

7. The method for improving the stability of federated learning based on blockchain and knowledge distillation according to claim 1, wherein The specific obtaining process of the fault risk assessment value is as follows: Obtain the fault risk parameters of the model parameters during the dual-model training process. The fault risk parameters include the average update frequency, the average update amplitude, and the actual transaction frequency. The actual transaction frequency represents the transaction frequency of local private financial data on the blockchain network; Obtain the local objective loss function, the global model weight factor, and the local global model weight factor of the participant client. At the same time, combine the obtained fault risk parameters and the reference fault risk data in the database to obtain the fault risk assessment value; The reference fault risk data includes the reference update frequency, the reference update amplitude, the reference transaction efficiency, the update frequency weight factor, the update amplitude weight factor, the transaction efficiency weight factor, and the local objective loss function weight factor.

8. The method for improving the stability of federated learning based on blockchain and knowledge distillation according to claim 7, wherein The specific limit expression of the fault risk assessment value is: Wherein, g is the number of participating party clients, g = 1, 2,..., G, G is the total number of participating party clients, e is the natural constant, FP represents the failure risk assessment value of heterogeneity in the federated learning architecture, q represents the global model weight factor, α1 represents the update frequency weight factor, A represents the average update frequency of model parameters during the dual-model training process, A0 represents the reference update frequency, α2 represents the update amplitude weight factor, B represents the average update amplitude of model parameters during the dual-model training process, B0 represents the reference update amplitude, α3 represents the transaction efficiency weight factor, C represents the actual transaction efficiency of local private financial data on the blockchain network, C0 represents the reference transaction efficiency, α4 represents the local objective loss function weight factor, F g (w g ) represents the local objective loss function of the g-th participating party client, w g represents the local global model weight factor of the g-th participating party client.

9. The method for improving the stability of federated learning based on blockchain and knowledge distillation according to claim 1, wherein The specific process for judging whether to send a single-point fault risk instruction based on the obtained fault risk assessment value is: Judge whether the obtained fault risk assessment value is not less than the preset fault risk assessment value in the database: If the obtained fault risk assessment value is not less than the preset fault risk assessment value in the database, the blockchain data sharing platform sends a single-point fault risk instruction; If the obtained fault risk assessment value is less than the preset fault risk assessment value in the database, the blockchain data sharing platform sends a federated learning optimization completion instruction.

10. A federated learning stability improvement system based on blockchain and knowledge distillation, characterized in that, Including: A dual-model training module, a fault risk assessment value acquisition module, and a smart contract verification module; Among them, the dual-model training module is used to perform dual-model training according to the obtained local private financial data of each participant client and upload the results of the dual-model training to the constructed blockchain data sharing platform; The fault risk assessment value acquisition module is used to construct a federated learning architecture and simultaneously monitor the changes in heterogeneity in the constructed federated learning architecture to obtain the fault risk assessment value. The fault risk assessment value is used to evaluate the degree of difference in heterogeneity in the federated learning architecture; The smart contract verification module is used to judge whether to send a single-point fault risk instruction based on the obtained fault risk assessment value. If so, perform smart contract verification; otherwise, complete the federated learning optimization.

Citation Information

Patent Citations

  • Client selection optimization method and device for unstable federated learning scenarios

    CN114841368B

  • Federal learning system and method based on fairness and reputation mechanism

    CN115511102A

Cited By

  • Financial data risk analysis method and system based on large model

    CN120509982A

  • Safe and credible federal learning method for large model illusion perception

    CN121998129A

  • A secure and reliable federated learning method for large model hallucination perception

    CN121998129B