Federal learning-based knowledge base updating method and system

Through the knowledge base update method based on federated learning, dynamically screening the participating device subset, performing encryption model training and layered knowledge distillation, combining smart contracts and blockchain, the problems of data privacy leakage and inefficiency in knowledge base updates are solved, and efficient, secure and trusted knowledge base updates are achieved.

CN120196640AInactive Publication Date: 2025-06-24GUANGZHOU QIANJIN INTELLECTUAL PROPERTY SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510319942.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology has problems such as data privacy leakage risks, data inconsistency, uneven data quality and inefficient updates during the knowledge base update process. It is especially difficult to achieve efficient, secure and trustworthy knowledge base updates in multi-source heterogeneous data environments.

Method used

The knowledge base update method based on federated learning is adopted, and the subset of participating devices is dynamically screened, model training is performed based on local knowledge base differential data, encrypted local knowledge update parameters, and global knowledge fusion and privacy enhancement processing are carried out through a layered knowledge distillation architecture, and global knowledge update instructions are finally generated, and the legality and immutability of the update instructions are ensured through smart contracts and blockchain.

Benefits of technology

It realizes efficient, secure and trustworthy knowledge base updates, ensuring the improvement of data privacy protection, knowledge consistency and update efficiency, and at the same time promoting knowledge sharing and collaborative updates between multiple clients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196640A_ABST
    Figure CN120196640A_ABST
Patent Text Reader

Abstract

The invention relates to a knowledge base updating method and system based on federated learning, and belongs to the technical field of federated learning and knowledge base management.The method comprises the steps that client device subsets are dynamically screened, model training is executed based on local knowledge base difference data, and encrypted local knowledge updating parameters are generated; performing global knowledge fusion and privacy enhancement processing by using the hierarchical knowledge distillation architecture to generate a global knowledge update instruction; the legality and non-tampering property of the update instruction are ensured through block chain evidence storage; each client node obtains an update instruction from the alliance chain, executes local knowledge base version synchronization and feeds back an update effect, in addition, an exception handling mechanism is designed to guarantee system stability and safety, the knowledge base is updated efficiently and safely on the premise that data privacy is protected, and the user experience is improved. And the credibility and the non-tampering property of the updating process are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of federated learning and knowledge bases, and specifically to a method and system for updating a knowledge base based on federated learning. Background Art

[0002] With the rapid development of information technology and the explosive growth of data volume, knowledge bases have been widely used in many fields such as enterprises, scientific research, and medical care. The update of knowledge bases is crucial for maintaining their effectiveness and practicality. Traditional knowledge base update methods usually rely on centralized data collection and processing, which is not only inefficient but also has the risk of data privacy leakage. Especially when dealing with multi-source heterogeneous data, problems such as data inconsistency, uneven data quality, and excessive data volume make the update process complex and time-consuming.

[0003] As an emerging machine learning paradigm, federated learning can use the data of multiple clients for joint model training without sharing the original data, effectively protecting data privacy. However, there are still some deficiencies in the existing federated learning technologies when applied to knowledge base updates. For example, the heterogeneity and dynamics of client devices make it a challenge to select appropriate participating devices. Problems such as multi-source knowledge fusion, privacy protection, and update effect evaluation involved in the knowledge base update process have not been fully solved.

[0004] In addition, during the process of knowledge base update, how to ensure the legality and immutability of update instructions, and how to efficiently perform knowledge synchronization and version management are also important issues that the existing technologies need to face. Traditional centralized update mechanisms are difficult to meet the requirements of multi-client collaborative updates in a distributed environment, and are prone to data inconsistency and update conflicts. Summary of the Invention

[0005] The main purpose of this application is to provide a method and system for updating a knowledge base based on federated learning, aiming to solve the deficiencies in the existing technologies and achieve efficient, secure, and trustworthy knowledge base updates.

[0006] To achieve the above invention objective, this application proposes a knowledge base update method based on federated learning, which is characterized by including the following steps: S1. Dynamically screen a subset of devices participating in the current round of federated learning from online client devices according to the dynamic contribution degree evaluation rule; S2. On the screened client devices, perform model training based on the local knowledge base difference data to generate encrypted local knowledge update parameters; S3. Input the local parameters of each client into a hierarchical knowledge distillation architecture, and sequentially perform global knowledge fusion and privacy enhancement processing to generate global knowledge update instructions; S4. Combine the global instructions with the client signature information to generate a verification request, and write it into the consortium chain after performing legal validity verification through a smart contract; S5. Each client node obtains the verified update instructions from the consortium chain, performs local knowledge base version synchronization, and feeds back the update effect.

[0007] Further, the dynamic contribution degree evaluation rule in S1 includes: periodically collecting the device status data of each client, including network bandwidth, computing load, and storage space; calculating the knowledge contribution degree benchmark value of the client according to the historical contribution record; dynamically adjusting the selection ratio based on the current number of online devices and network bandwidth conditions; sending a federated learning task package containing training configuration parameters to the selected clients.

[0008] Further, the calculation dimensions of the knowledge contribution degree benchmark value include: data quality score: calculated based on the improvement range of the knowledge base accuracy before and after the training task; parameter upload success rate: counting the proportion of successfully uploaded parameters in the client's historical tasks; version update timeliness: quantified according to the average time interval for the client to complete the knowledge base version update.

[0009] Further, S2 includes: generating a difference data set of the local knowledge base through a version comparison engine, where the difference data set contains new knowledge entities and associated relationships; loading the federated learning basic model in a trusted execution environment (TEE), and injecting dynamic noise to initialize the model parameters; using a homomorphic encryption algorithm to encrypt the training gradients to generate locally updated parameters with digital signatures; performing gradient clipping and quantization compression on the encrypted parameters, and the compression ratio is dynamically adjusted according to the client's network bandwidth.

[0010] Further, the execution steps of the hierarchical knowledge distillation architecture in S3 include: in the global knowledge distillation layer, using a multi-head attention mechanism to perform feature alignment on multi-source local parameters; in the local knowledge enhancement layer, performing differential privacy processing and dimension compression on the aligned feature vectors; feeding back the global knowledge features to each local model through a cross-layer feedback channel for parameter calibration.

[0011] Further, the differential privacy processing and dimensionality compression include: adding random noise conforming to the Laplace distribution to the feature vector; performing feature dimensionality reduction using the principal component analysis method; and performing normalization processing on the processed eigenvalues to accelerate model convergence.

[0012] Further, S4 includes: generating a verification request packet containing a timestamp, a client fingerprint, and a parameter hash value; calling a smart contract to verify the consistency between the parameter hash value and the locally calculated check value; after the verification passes, packing the global instruction and the verification record to generate a blockchain transaction; and writing the transaction into the immutable storage area after more than a preset proportion of verification nodes confirm it.

[0013] Further, the update effect feedback in S5 includes: deploying a lightweight verification model on the client to evaluate the change in accuracy before and after knowledge update; counting performance metrics such as knowledge retrieval response time and storage space occupancy rate; and generating an evaluation report containing the index comparison data and sending it back to the coordination server.

[0014] Further, the knowledge base update method further includes an exception handling mechanism: automatically removing the client from the list of participating nodes when it is detected that the client is offline and timed out; rolling back to the previous valid version and starting retraining when the parameter aggregation fails; and freezing the relevant client and generating a security alert when malicious parameter injection is identified.

[0015] This application also proposes a knowledge base update system based on federated learning, including: A dynamic selection module: used to dynamically screen a subset of devices participating in the current round of federated learning from online client devices according to a preset contribution degree evaluation rule; A local training module: used to perform model training on the selected client devices based on the local knowledge base difference data and generate encrypted local knowledge update parameters; A hierarchical processing module: used to input the local parameters of each client into a hierarchical knowledge distillation architecture, and sequentially perform global knowledge fusion and privacy enhancement processing to generate global knowledge update instructions; A blockchain evidence storage module: used to combine the global instruction with the client signature information to generate a verification request, and write it into the consortium chain after performing legal validity verification through a smart contract; A synchronization control module: used for each client node to obtain the verified update instruction from the consortium chain, perform local knowledge base version synchronization, and feedback the update effect.

[0016] A knowledge base update method and system based on federated learning provided by the present invention dynamically screens a subset of devices participating in the current round of federated learning from online client devices. This process ensures the quality and diversity of the participating devices, thus laying a foundation for subsequent efficient training. On the selected client devices, model training is performed based on local knowledge base difference data to generate encrypted local knowledge update parameters. By using local difference data, the amount of data transmission is reduced and the training efficiency is improved, ensuring the pertinence and effectiveness of local training. The local parameters of each client are input into a hierarchical knowledge distillation architecture, and global knowledge fusion and privacy enhancement processing are sequentially performed to generate global knowledge update instructions. This process not only realizes the efficient fusion of multi-source knowledge, improves the accuracy and richness of global knowledge update, but also further enhances data privacy protection. The global instruction is combined with the client signature information to generate a verification request, which is written into the alliance chain after performing legality verification through a smart contract. Each client node obtains the verified update instruction from the alliance chain, performs local knowledge base version synchronization and feedback on the update effect, ensuring the transparency and immutability of the update process, enhancing the credibility and traceability of the system, and at the same time promoting knowledge sharing and collaborative update among multiple clients. Description of the Drawings

[0017] Figure 1 It is a schematic flowchart of a knowledge base update method based on federated learning according to an embodiment of the present application.

[0018] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0019] In order to make the object, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0020] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0021] Referring to Figure 1 , in order to achieve the above object of the invention, the present invention provides a knowledge base update method based on federated learning, including the following steps: S1. Dynamically screen a subset of devices participating in the current round of federated learning from online client devices according to the dynamic contribution degree evaluation rule; S2. On the screened client devices, perform model training based on the local knowledge base difference data to generate encrypted local knowledge update parameters; S3. Input the local parameters of each client into the hierarchical knowledge distillation architecture, and sequentially perform global knowledge fusion and privacy enhancement processing to generate global knowledge update instructions; S4. Combine the global instructions with the client signature information to generate a verification request, and write it to the alliance chain after performing legal verification through a smart contract; S5. Each client node obtains the verified update instructions from the alliance chain, performs local knowledge base version synchronization and feedback on the update effect.

[0022] In this embodiment, the online client devices are screened through a preset dynamic contribution degree evaluation rule to select a subset of devices to participate in this round of federated learning. The dynamic contribution degree evaluation rule comprehensively considers various factors to ensure that the selected client devices can contribute high-quality data and computing power to federated learning. The screened client devices can provide better-quality data, thereby improving the effect and accuracy of model training. At the same time, unnecessary data transmission and computing are reduced, interference from low-quality data to the training process is avoided, the operating efficiency of the entire system is improved, the scope of devices participating in federated learning is restricted, the risk of data leakage is reduced, and the privacy of the client is better protected. On the screened client devices, the differential data in the local knowledge base is used to perform model training. The differential data refers to the newly added or updated part in the local knowledge base compared with the previous round of training or the global knowledge base, including newly added knowledge entities and their associated relationships, etc. By only using this differential data for training, the data processing volume and transmission volume can be reduced, thereby improving the training efficiency, reducing the amount of data to be processed, accelerating the speed of model training, enabling the client device to generate local knowledge update parameters more quickly, avoiding repeated processing of the entire knowledge base, saving the computing resources and storage space of the client device, enabling the model to quickly adapt to new knowledge and changes, and improving the practicality and accuracy of the model. The local knowledge update parameters generated by each client device are input into the hierarchical knowledge distillation architecture. First, in the global knowledge distillation layer, the multi-head attention mechanism is used to align the features of multi-source local parameters, enabling the parameters of different clients to be effectively fused in a unified feature space, realizing the efficient fusion of multi-source knowledge, improving the accuracy and richness of global knowledge update, and enabling the knowledge base to better reflect various types of information. Then, in the local knowledge enhancement layer, differential privacy processing and dimension compression are performed on the aligned feature vectors to further enhance privacy protection and optimize the data structure. Through differential privacy processing, the risk of data leakage is further reduced, ensuring the confidentiality of client data. Finally, the global knowledge features are fed back to each local model through the cross-layer feedback channel for parameter calibration to generate global knowledge update instructions. The parameter calibration process helps to improve the stability and accuracy of the model, enabling the model to better adapt to different application scenarios.Combine the generated global knowledge update instructions with the signature information of the client device to generate a verification request. The immutable feature of the blockchain ensures the credibility of the update instructions, allowing users to safely update the knowledge base based on these instructions. Use a smart contract to perform a legality check on the verification request to ensure the integrity and authenticity of the update instructions. The automatic execution and verification mechanism of the smart contract reduces human intervention and lowers the security risks caused by human errors or malicious operations. After passing the verification, write the global instructions and related information into the consortium chain. Leverage the distributed ledger feature of the blockchain to achieve the immutability and traceability of the update instructions. Each update operation is recorded on the blockchain, facilitating subsequent auditing and tracing, and helping to identify and resolve potential issues. Each client node obtains the verified update instructions from the consortium chain and executes the version synchronization operation of the local knowledge base according to the instructions to ensure that the knowledge bases of all clients are consistent and up-to-date. The knowledge bases of all clients can be synchronized to the latest version in a timely manner, ensuring the consistency and integrity of knowledge and avoiding problems caused by out-of-sync knowledge versions. After synchronization, the client can deploy a lightweight verification model to evaluate the change in accuracy before and after the knowledge update, count performance metrics such as the knowledge retrieval response time and storage space occupancy rate, and send an evaluation report containing the comparison data of these metrics back to the coordination server. Through the feedback of the evaluation report, the coordination server can understand the update effects and system performance of each client, providing a basis for further optimization and adjustment. Based on the feedback performance metrics, problems existing in the knowledge base update process can be discovered, promoting the continuous improvement and optimization of the system.

[0023] In one embodiment, the dynamic contribution degree evaluation rule in S1 includes: periodically collecting the device status data of each client, including network bandwidth, computing load, and storage space; calculating the knowledge contribution degree benchmark value of the client according to the historical contribution record; dynamically adjusting the selection ratio based on the current number of online devices and network bandwidth conditions; and sending a federated learning task package containing training configuration parameters to the selected clients.

[0024] In this embodiment, the device status data of each client is periodically collected, including network bandwidth, computing load, and storage space. These data reflect the current performance and available resources of the client device and are important bases for evaluating its ability to effectively participate in federated learning. By monitoring network bandwidth, computing load, and storage space, it can be ensured that the client devices participating in federated learning have sufficient performance and resources to execute training tasks. At the same time, based on the device status data, training tasks can be reasonably allocated to avoid training delays or failures caused by insufficient device performance and improve the utilization efficiency of resources. The knowledge contribution benchmark value of the client is calculated based on the historical contribution record. This benchmark value comprehensively considers the client's past performance in federated learning, such as data quality, parameter upload success rate, and version update timeliness, and reflects the degree of contribution of the client to federated learning. By calculating the knowledge contribution benchmark value, the client can be motivated to provide higher-quality data and more stable performance, thereby improving the performance of the entire federated learning system. Based on the historical contribution record, training resources can be allocated more fairly to ensure that clients with greater contributions can obtain more participation opportunities and further improve the training effect. The selection ratio is dynamically adjusted based on the current number of online devices and network bandwidth conditions. This process takes into account the current network load and available resources to ensure that suitable client devices can be effectively screened under different network conditions, improving the flexibility and adaptability of the system. By reasonably adjusting the selection ratio, it can be ensured that the number and quality of client devices participating in federated learning reach the best balance, thereby improving the training efficiency and effect. A federated learning task package containing training configuration parameters is sent to the selected clients, which can ensure that all participating client devices perform training according to unified specifications and requirements, improving the consistency and comparability of training. The task package contains all the configuration information and parameters required for the client to execute the federated learning task, ensuring that the client can perform training according to unified specifications and requirements. After receiving the task package, the client can quickly start the training task, reducing delays and errors caused by inconsistent configurations and further improving the training efficiency.

[0025] In one embodiment, the calculation dimensions of the knowledge contribution benchmark value include: data quality score: calculated based on the improvement amplitude of the knowledge base accuracy before and after the training task; parameter upload success rate: counting the proportion of successfully uploaded parameters in the client's historical tasks; version update timeliness: quantified according to the average time interval for the client to complete the knowledge base version update.

[0026] In this embodiment, by comprehensively considering three dimensions: data quality score, parameter upload success rate, and version update timeliness, the data quality score is calculated based on the improvement amplitude of the knowledge base accuracy before and after the training task. Specifically, by comparing the knowledge base accuracy before and after training, the contribution degree of the data to model training is evaluated. The evaluation of data quality usually includes multiple dimensions such as accuracy, integrity, consistency, uniqueness, timeliness, and availability. Through the data quality score, training resources can be reasonably allocated, high-quality data can be preferentially utilized, and training efficiency can be improved; the parameter upload success rate is calculated by counting the proportion of successfully uploaded parameters in the client's historical tasks. This indicator reflects the stability and reliability of data transmission during the federated learning process. A higher parameter upload success rate means that the client can stably participate in federated learning, reducing training interruptions caused by data transmission failures. Stable parameter upload can ensure the continuity of the training process and avoid repeated training caused by data loss; the version update timeliness is quantified according to the average time interval for the client to complete the knowledge base version update. This indicator evaluates the response speed and efficiency of the client during the knowledge base update process. Timely version update can ensure the timeliness of the client's knowledge base, enabling it to quickly adapt to new information and changes. Timely updates of all clients contribute to maintaining the knowledge consistency of the entire system and promoting collaborative work among multiple clients; this multi-dimensional evaluation method not only encourages clients to provide higher-quality data and more stable performance but also promotes the efficient operation and continuous optimization of the entire federated learning system In one embodiment, S2 includes: generating a differential dataset of the local knowledge base through a version comparison engine, where the differential dataset contains newly added knowledge entities and associated relationships; loading the federated learning basic model in a trusted execution environment (TEE) and initializing the model parameters by injecting dynamic noise; encrypting the training gradient using a homomorphic encryption algorithm to generate locally updated parameters with a digital signature; performing gradient clipping and quantization compression on the encrypted parameters, and the compression ratio is dynamically adjusted according to the client's network bandwidth.

[0027] In this embodiment, through the version comparison engine, the differences between the current version of the local knowledge base and the previous version or the global knowledge base are compared to generate a difference dataset containing newly added knowledge entities and their associated relationships. This method only focuses on the changed parts rather than the entire knowledge base, thereby reducing the data processing volume and transmission volume, accelerating the model training speed, avoiding repeated processing of the entire knowledge base, and saving the computing resources and storage space of the client device; load the federated learning base model in the trusted execution environment (TEE). The trusted execution environment (TEE) is a hardware-based security technology that can execute code and process data in an isolated environment to ensure the confidentiality and integrity of data and model parameters, preventing data leakage and tampering. Load the federated learning base model in the TEE and inject dynamic noise to initialize the model parameters to further enhance security. The injection of dynamic noise increases the robustness of the model, preventing overfitting and adversarial attacks; use the homomorphic encryption algorithm to encrypt the training gradients. Homomorphic encryption is an encryption technology that allows direct computation on encrypted data without prior decryption. In federated learning, the client uses the homomorphic encryption algorithm to encrypt the training gradients to generate locally updated parameters with digital signatures. In this way, the server can aggregate these encrypted gradients without decryption. Homomorphic encryption ensures the privacy of the training gradients during transmission and processing, preventing data leakage, and the digital signature ensures the integrity of the parameters and the credibility of the source, preventing the injection of malicious parameters; perform gradient clipping on the encrypted parameters to reduce the dynamic range of the gradients and prevent gradient explosion. Through gradient clipping and quantization compression, the data transmission volume is reduced, and the transmission efficiency is improved. At the same time, the quantization compression rate is dynamically adjusted according to the client network bandwidth to further reduce the data transmission volume and ensure efficient data transmission under different network conditions.

[0028] In one embodiment, the execution steps of the hierarchical knowledge distillation architecture in S3 include: in the global knowledge distillation layer, use the multi-head attention mechanism to align the features of multi-source local parameters; in the local knowledge enhancement layer, perform differential privacy processing and dimension compression on the aligned feature vectors; and through the cross-layer feedback channel, feed back the global knowledge features to each local model for parameter calibration.

[0029] In this embodiment, the multi-head attention mechanism performs parallel calculations through multiple attention heads. Each attention head independently performs linear transformations on the input query, key, and value, then calculates the attention weights, and finally concatenates the outputs of all attention heads and performs a linear projection to obtain the final output. In the global knowledge distillation layer, the multi-head attention mechanism is used to align the feature vectors of local parameters from different clients, enabling these parameters to be effectively fused in a unified feature space. The multi-head attention mechanism can capture the complex relationships between different features, enhancing the model's understanding and expression of multi-source knowledge. Through feature alignment, the model can better integrate knowledge from different clients, improving its applicability and accuracy in different scenarios; differential privacy processing protects data privacy by adding noise to the data, ensuring that attackers cannot accurately infer information about individual data points from the output. In the local knowledge enhancement layer, differential privacy processing and dimensionality compression are performed on the aligned feature vectors. Dimensionality compression reduces the dimension of the feature vectors, retaining the main feature information and further improving the computational efficiency. Differential privacy processing ensures the privacy of the feature vectors, preventing the leakage of sensitive information. Dimensionality compression reduces the dimension of the feature vectors, reducing the computational complexity and accelerating the model training and inference speed; the cross-layer feedback channel feeds back the fused knowledge features in the global knowledge distillation layer to each local model for calibrating the parameters of the local model. This process realizes the collaborative optimization of global knowledge and local models through a feedback mechanism, ensuring that the local models can better adapt to global knowledge updates.

[0030] In one embodiment, the differential privacy processing and dimensionality compression include: adding random noise that conforms to the Laplace distribution to the feature vectors; using the principal component analysis method for feature dimensionality reduction; performing normalization processing on the processed eigenvalues to accelerate model convergence.

[0031] In this embodiment, differential privacy processing protects data privacy by adding random noise that conforms to the Laplace distribution to the feature vector. The amount of Laplace noise added is determined by the privacy budget (epsilon) and the sensitivity of the data. After adding the Laplace noise, an attacker cannot accurately infer the information of a single data point from the output, thereby protecting the privacy of the data. By reasonably setting the privacy budget and sensitivity, the availability of the data can be maintained while protecting privacy, ensuring the effectiveness of model training; Principal Component Analysis (PCA) is a commonly used dimensionality reduction technique that maps high-dimensional data to a low-dimensional space through a linear transformation while retaining the main feature information of the data. In the local knowledge enhancement layer, PCA is applied to the feature vector with added Laplace noise for dimensionality compression. PCA calculates the covariance matrix of the feature vector and finds its principal components, and projects the data into the low-dimensional space composed of these principal components. Dimensionality compression reduces the dimension of the feature vector, reduces the computational complexity, and speeds up the model training and inference speed. PCA ensures that the compressed feature vector still retains the main feature information of the original data, thereby maintaining the representativeness of the data while reducing the dimension; perform normalization processing on the processed eigenvalue, scale the eigenvalue to a specific range (such as the interval [0, 1]). Normalization processing makes the eigenvalue distribution more uniform, avoids the problem of slow gradient descent caused by too large a difference in the eigenvalue range, thereby speeding up the model training speed, helping to reduce the fluctuations in the model training process, and improving the stability and robustness of the model.

[0032] In one embodiment, S4 includes: generating a verification request packet containing a timestamp, a client fingerprint, and a parameter hash value; invoking a smart contract to verify the consistency between the parameter hash value and the locally calculated check value; after the verification passes, packing the global instruction and the verification record to generate a blockchain transaction; when more than a preset proportion of verification nodes confirm, writing the transaction into the immutable storage area.

[0033] In this embodiment, by including a timestamp, client fingerprint, and parameter hash value, the verification request packet can completely record the key information of the request, preventing the request from being tampered with. The addition of the parameter hash value ensures the integrity of the parameters during transmission and prevents the parameters from being maliciously tampered with. The smart contract is called to verify the consistency between the parameter hash value and the locally calculated check value. A smart contract is a self-executing contract clause that can ensure the transparency and fairness of the verification process. By verifying the consistency between the parameter hash value and the locally calculated check value through the smart contract, the accuracy and credibility of the parameters are ensured. The self-executing feature of the smart contract improves the efficiency of the verification process and reduces manual intervention. After the verification passes, the global instruction and verification record are packaged to generate a blockchain transaction. The blockchain transaction records the key information of the global instruction and the verification process, ensuring the transparency and immutability of the transaction and improving the credibility of the system. When more than a preset proportion of verification nodes confirm, the transaction is written into the immutable storage area. The immutable storage area ensures the permanence and immutability of the transaction record, facilitating subsequent auditing and traceability. The confirmation by multiple verification nodes ensures the legality and credibility of the transaction.

[0034] In one embodiment, the update effect feedback in S5 includes: deploying a lightweight verification model on the client to evaluate the change in accuracy before and after knowledge update; counting performance metrics such as knowledge retrieval response time and storage space occupancy rate; and generating an evaluation report containing index comparison data and sending it back to the coordination server.

[0035] In this embodiment, a lightweight verification model is deployed on the client to evaluate the change in accuracy before and after knowledge update. The lightweight verification model is usually a simplified version of the model that can quickly evaluate the performance change of the model without consuming too much computing resources. The lightweight verification model can complete the evaluation in a short time, reducing the evaluation time and improving efficiency. Due to the lightweight model, it does not occupy too much computing resources and is suitable for running on resource-constrained client devices. Count performance metrics such as knowledge retrieval response time and storage space occupancy rate to evaluate the actual running effect after the knowledge base is updated. These metrics can reflect the retrieval efficiency and storage efficiency of the knowledge base and help identify potential performance bottlenecks. By counting these performance metrics, the deficiencies of the knowledge base in retrieval and storage can be found, providing a basis for subsequent optimization. A faster retrieval response time and a lower storage space occupancy rate can improve the user experience. Organize the above evaluation results and the counted performance metrics to generate an evaluation report containing index comparison data and send it back to the coordination server. The coordination server can summarize and analyze these reports to understand the update effect of the overall system. The coordination server can centrally manage the evaluation reports of each client, facilitating unified analysis and decision-making. By analyzing the evaluation reports, problems and improvement points in the system can be found, promoting the continuous optimization of the system.

[0036] In one embodiment, the knowledge base update method further includes an exception handling mechanism: when it is detected that a client is offline and times out, it is automatically removed from the list of participating nodes; when parameter aggregation fails, roll back to the previous valid version and start retraining; when malicious parameter injection is identified, freeze the relevant client and generate a security alert.

[0037] In this embodiment, the system continuously monitors the online status of the client. When it is detected that the client is offline and times out, it is automatically removed from the list of participating nodes. This process is achieved through heartbeat detection or regular communication checks to ensure that only active clients participate in the federated learning process. By removing offline clients in a timely manner, unnecessary waste of computing and communication resources for offline devices is avoided, ensuring that the clients participating in federated learning are all active, and reducing training interruptions or data inconsistency problems caused by offline clients; during the parameter aggregation process, if a failure occurs (such as incomplete parameters, incorrect format, etc.), the system will automatically roll back to the previous valid version of the global model and start the retraining process. This mechanism ensures the integrity and effectiveness of the global model. Rolling back to the previous valid version ensures that the global model will not be in an inconsistent or incorrect state due to aggregation failure. Through automatic retraining, the system can recover from the failure and continue with effective federated learning, improving the overall reliability; the system monitors and analyzes the parameters uploaded by the client to identify whether there is malicious parameter injection. Once malicious parameters are identified, the system will freeze the relevant client and generate a security alert. Through the security alert mechanism, the administrator can intervene in a timely manner and take further security measures to ensure the secure operation of the system.

[0038] This application also proposes a knowledge base update system based on federated learning, including: A dynamic selection module: used to dynamically screen a subset of devices participating in the current round of federated learning from online client devices according to a preset contribution degree evaluation rule; A local training module: used to perform model training on the selected client devices based on the local knowledge base difference data to generate encrypted local knowledge update parameters; A hierarchical processing module: used to input the local parameters of each client into a hierarchical knowledge distillation architecture, and sequentially perform global knowledge fusion and privacy enhancement processing to generate global knowledge update instructions; A blockchain evidence storage module: used to combine the global instruction with the client signature information to generate a verification request, and write it into the consortium chain after performing legality verification through a smart contract; A synchronization control module: used for each client node to obtain the verified update instruction from the consortium chain, perform local knowledge base version synchronization and feedback the update effect.

[0039] The operation mode of the system in this embodiment refers to the method embodiment described above, which will not be elaborated here.

[0040] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to memory, storage, database, or other media provided in this application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, there are various forms of RAM, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (RambuS) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0041] In summary, a knowledge base update method and system based on federated learning provided by the present invention dynamically screens a subset of devices participating in the current round of federated learning from online client devices. This process ensures the quality and diversity of the participating devices, thus laying a foundation for subsequent efficient training. On the screened client devices, model training is performed based on local knowledge base difference data to generate encrypted local knowledge update parameters. By using local difference data, the amount of data transmission is reduced and the training efficiency is improved, ensuring the pertinence and effectiveness of local training. The local parameters of each client are input into a hierarchical knowledge distillation architecture, and global knowledge fusion and privacy enhancement processing are sequentially performed to generate global knowledge update instructions. This process not only realizes the efficient fusion of multi-source knowledge, improves the accuracy and richness of global knowledge update, but also further enhances data privacy protection. The global instruction is combined with the client signature information to generate a verification request, which is written into the consortium chain after performing legal verification through a smart contract. Each client node obtains the verified update instruction from the consortium chain, performs local knowledge base version synchronization and feeds back the update effect, ensuring the transparency and immutability of the update process, enhancing the credibility and traceability of the system, and at the same time promoting knowledge sharing and collaborative update among multiple clients. The technical solution is applicable to various application scenarios, such as smart cities, financial services, medical fields, Internet of Things fields, etc., and has good adaptability and practicability.

[0042] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A knowledge base updating method based on federated learning, characterized in that The following steps are involved: S1. Dynamically select a subset of devices participating in this round of federated learning from online client devices according to the dynamic contribution evaluation rules; S2. On the selected client devices, perform model training based on the local knowledge base difference data to generate encrypted local knowledge update parameters; S3, input the local parameters of each client into the hierarchical knowledge distillation architecture, perform global knowledge fusion and privacy enhancement processing in sequence, and generate global knowledge update instructions; S4. Combine the global instruction with the client signature information to generate a verification request, which is then written into the alliance chain after the legality check is performed by the smart contract; S5. Each client node obtains the verified update instructions from the alliance chain, executes the local knowledge base version synchronization and provides feedback on the update effect.

2. The knowledge base updating method according to claim 1, characterized in that The dynamic contribution evaluation rules in S1 include: Periodically collect device status data of each client, including network bandwidth, computing load, and storage space; Calculate the client's knowledge contribution benchmark value based on historical contribution records; dynamically adjust the selection ratio based on the current number of online devices and network bandwidth conditions; Send a federated learning task package containing training configuration parameters to the selected client.

3. The knowledge base updating method according to claim 2, characterized in that The calculation dimensions of the knowledge contribution benchmark value include: Data quality score: calculated based on the improvement in the accuracy of the knowledge base before and after the training task; Parameter upload success rate: Statistics on the ratio of successfully uploaded parameters in the client's historical tasks; Version update timeliness: quantified based on the average time interval for clients to complete knowledge base version updates.

4. The knowledge base updating method according to claim 1, characterized in that The S2 includes: generating a difference data set of the local knowledge base through a version comparison engine, wherein the difference data set includes newly added knowledge entities and association relationships; loading the federated learning basic model in a trusted execution environment (TEE), and injecting dynamic noise to initialize the model parameters; encrypting the training gradients using a homomorphic encryption algorithm to generate local update parameters with digital signatures; performing gradient clipping and quantization compression on the encrypted parameters, and the compression rate is dynamically adjusted according to the client network bandwidth.

5. The knowledge base updating method according to claim 1, characterized in that The execution steps of the hierarchical knowledge distillation architecture in S3 include: in the global knowledge distillation layer, a multi-head attention mechanism is used to perform feature alignment on multi-source local parameters; in the local knowledge enhancement layer, differential privacy processing and dimensionality compression are performed on the aligned feature vectors; and global knowledge features are fed back to each local model through a cross-layer feedback channel for parameter calibration.

6. The knowledge base updating method according to claim 5, characterized in that The differential privacy processing and dimensionality compression include: adding random noise that conforms to the Laplace distribution to the feature vector; using the principal component analysis method to reduce the feature dimension; and performing normalization processing on the processed eigenvalues ​​to accelerate model convergence.

7. The knowledge base updating method according to claim 1, characterized in that The S4 includes: generating a verification request package including a timestamp, a client fingerprint and a parameter hash value; calling a smart contract to verify the consistency between the parameter hash value and the locally calculated verification value; after the verification passes, packaging the global instruction and the verification record to generate a blockchain transaction; when more than a preset proportion of verification nodes confirm, writing the transaction into an immutable storage area.

8. The knowledge base updating method according to claim 1, characterized in that The update effect feedback in S5 includes: deploying a lightweight verification model on the client to evaluate the accuracy change before and after the knowledge update; statistically analyzing performance indicators such as knowledge retrieval response time and storage space occupancy rate; generating an evaluation report containing indicator comparison data and transmitting it back to the coordination server.

9. The knowledge base updating method according to claims 1 to 8, characterized in that It also includes exception handling mechanisms: When the client is detected to be offline for timeout, it will be automatically removed from the participating node list; When parameter aggregation fails, roll back to the last valid version and start retraining; When malicious parameter injection is identified, the relevant client is frozen and a security alert is generated.

10. A knowledge base updating system based on federated learning, characterized in that include: Dynamic selection module: used to dynamically select a subset of devices participating in this round of federated learning from online client devices according to preset contribution evaluation rules; Local training module: used to perform model training based on the local knowledge base difference data on the selected client devices and generate encrypted local knowledge update parameters; Hierarchical processing module: used to input the local parameters of each client into the hierarchical knowledge distillation architecture, perform global knowledge fusion and privacy enhancement processing in sequence, and generate global knowledge update instructions; Blockchain evidence storage module: used to combine global instructions with client signature information to generate verification requests, and write them into the alliance chain after performing legality verification through smart contracts; Synchronization control module: used for each client node to obtain verified update instructions from the alliance chain, execute local knowledge base version synchronization and feedback the update effect.