Federal learning method and system based on multi-agent and knowledge distillation, and medium
Through the federated learning method of multi-agents and knowledge distillation, combined with blockchain and large language models, the major problems of model drift and communication overhead caused by non-independent and homogeneous distribution of data in the medical Internet of Things are solved, and efficient personalized knowledge sharing and secure federated learning are achieved.
Patent Information
- Application Number
- CN202510855313.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
In the medical IoT scenario, the existing federated learning method faces the problems of model drift, slow convergence and insufficient personalized performance caused by non-independent and homogeneous data distribution. At the same time, communication overhead is large, and there is a risk of privacy leakage on the central server.
The federated learning method based on multi-agents and knowledge distillation is adopted, and the knowledge production and learning steps are decoupled, and a large language model is introduced for personalized needs analysis and knowledge evaluation, optimize knowledge selection and integration, and reduce communication overhead.
It significantly improves the convergence speed and personalization level of the model in a heterogeneous data environment, reduces communication overhead, enhances the security and trustworthiness of the system, and adapts to the computing and network limitations of edge devices.
Smart Images

Figure CN120354255A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of federated learning, and particularly to a federated learning method, system, and medium based on multi-agent and knowledge distillation. Background Art
[0002] In the consumer-oriented medical Internet of Things (IoT) scenario, patients are increasingly relying on wearable devices, home health monitors, remote diagnosis platforms, etc. for daily health management. These terminals continuously generate a large amount of personal health data, including highly sensitive information such as heart rate, blood sugar, movement trajectories, and medical histories. What consumers are most concerned about is whether this data will be misused by unauthorized institutions once it is uploaded to the cloud or a third-party server. Especially during remote diagnosis or online consultations, patients often need to submit detailed health reports and real-time monitoring records. If the platform's security protection is not in place, there is a high risk of privacy leakage or identity theft. This concern about data security and privacy protection has become an important obstacle restricting the large-scale popularization of the medical IoT.
[0003] To balance data utilization and privacy protection, federated learning, as a distributed training paradigm, has gradually become a hot topic in the industry. Federated learning allows each participating terminal to complete model training locally and only upload model updates instead of raw data, thereby reducing the risk of exposure of sensitive information. In the medical IoT scenario, by keeping patient data on local terminals (such as wearable devices or home monitoring gateways) and periodically uploading local model parameters to the cloud or collaborating and aggregating with other devices, it is possible to achieve cross-institutional collaborative diagnosis and prediction while alleviating patients' concerns about privacy leakage.
[0004] However, the diversity and differences in the data collected by each terminal device will lead to obvious non-independent and identically distributed problems. The physiological conditions, living habits, device brands and models used, and sensor calibration methods of each user are different. When these devices independently train models locally, the model parameters highly depend on the sample distribution of the devices they belong to. If these "biased towards their own distributions" local model parameters are simply averaged, it is difficult for the global model to maintain high accuracy in any data distribution, resulting in inaccurate prediction results for certain groups (such as specific age groups or people with special living habits).
[0005] Secondly, the communication overhead of federated learning is particularly prominent in the medical IoT scenario. Terminals such as wearable devices and home monitoring gateways have limited bandwidth themselves and need to maintain long-term online monitoring. Once they need to frequently upload model updates containing millions or even tens of millions of parameters, it will occupy a large amount of network resources, leading to bandwidth congestion and transmission delays. Since these devices mostly rely on home Wi-Fi or cellular networks, network fluctuations will make large-scale parameter synchronization more unstable, and even disconnections and retransmissions may occur, further slowing down the model training progress.
[0006] Thirdly, the risks of relying on a central server for model aggregation are particularly prominent in the medical Internet of Things scenario. Medical data belongs to highly sensitive personal privacy. Once the central server is attacked or misconfigured, the health records of thousands of users may be leaked, even triggering serious consequences at the legal and regulatory levels. On the other hand, the central server's centralized processing of high-concurrency model aggregation tasks is also prone to becoming a performance bottleneck or a single point of failure source. Once it goes down, the entire collaborative training and diagnosis and treatment processes will be interrupted. Since consumers' concerns about health privacy are already very strong, the potential risks of this centralized architecture have further exacerbated users' lack of trust in medical Internet of Things services. Summary of the Invention
[0007] The technical problem to be solved by this application is to overcome the deficiencies of the prior art. This application provides a federated learning method, system and medium based on multi-agent and knowledge distillation, which realizes personalized selection, adaptive filtering and efficient sharing of knowledge in the federated learning process, effectively alleviates the client drift phenomenon caused by the non-independent and identically distributed characteristics of data, significantly improves the convergence speed, overall performance and personalization level of the model in a heterogeneous data environment, and is committed to optimizing the communication overhead between clients.
[0008] To achieve the above object, the first aspect of this application provides a federated learning method based on multi-agent and knowledge distillation, including the following steps: S1. Train the local model: Independently train the local model using the collected personal health data to generate a teacher model that can be used for knowledge distillation and a student model for receiving external knowledge distillation learning; S2. Decoupled knowledge distillation: Extract standardized distillation knowledge fragments from the teacher model trained in S1. The knowledge distillation process decouples the knowledge production step from the knowledge learning step; S3. Decentralized knowledge sharing based on blockchain: Package the standardized distillation knowledge fragments extracted in S2 and their metadata into blockchain transactions and submit them to the blockchain network, verify the legality of the transactions, and write them into the blockchain ledger to achieve decentralized distribution; S4. Multi-agent collaborative knowledge management based on large language models to update the student model: Evaluate the status of the student model through personalized demand analysis and knowledge evaluation agents, generate a personalized learning demand vector, and analyze the value of external distillation knowledge fragments to generate an evaluation result; According to the evaluation results and the personalized learning requirement vector, through the adaptive knowledge selection and fusion guidance agent, complete the screening, weighting, and fusion strategy design of the external distilled knowledge fragments; perform joint distillation operation on the selected external distilled knowledge fragments and the output of the local model, and construct a hybrid loss function that includes a real label supervision term and an external knowledge soft label term; according to the constructed hybrid loss function, guide the student model to update parameters, and realize the effective transfer of the external distilled knowledge fragments to the local model; S5. Repeat S1 - S4 until the training termination condition is met.
[0009] Optionally, the knowledge production steps in S2 include: Perform inference operation on the proxy dataset: The client uses the teacher model trained in S1 to perform inference on each sample in the proxy dataset uniformly distributed by the system, and generate the original prediction results of the output layer. The proxy dataset is a publicly released non-sensitive health dataset; Extract the model prediction logits: The client takes the original prediction vector logits of the output layer of the teacher model as the initial knowledge representation, and records the confidence level of the model for each category; Perform standardization processing: The client performs statistical standardization processing on the logits, calculates the mean and standard deviation, and normalizes the logits so that the logits have a uniform distribution characteristic on the numerical scale; Package the distilled knowledge fragments: The client organizes the standardized prediction results into a unified format of distilled knowledge fragments, that is, external distilled knowledge fragments, and adds client identification, timestamp, and current training round information to construct a structured knowledge object for knowledge sharing in S3.
[0010] Optionally, the knowledge learning steps in S2 include: Load the external distilled knowledge: The client receives and loads the set of external distilled knowledge fragments obtained from the blockchain network as the input for distillation learning; Construct a hybrid loss function: The client uses the hybrid loss function to optimize the student model. The hybrid loss function includes a local label supervision term and an external knowledge KL divergence term, and the fusion weight is specified by S4; Perform knowledge distillation learning: The client performs backpropagation and parameter update on the student model based on the constructed hybrid loss function to realize the absorption and integration of the external distilled knowledge.
[0011] Optionally, the knowledge sharing process in S3 includes: Submit the knowledge fragments to the blockchain network: The client broadcasts the constructed distilled knowledge fragments to the blockchain network as a transaction record; Execute transaction legality verification: The consensus nodes in the blockchain network verify the transaction content, checking the format of the transaction content, identity signature, and data integrity; Write to the blockchain ledger: After passing the verification, the transaction is packaged into a new block and recorded in the blockchain ledger to achieve immutable storage of knowledge objects.
[0012] Optionally, in S4, the personalized demand analysis and knowledge evaluation agent evaluates the student model status and analyzes the external knowledge value, including: Collect model performance data: The personalized demand analysis and knowledge evaluation agent monitors and statistics the prediction loss and accuracy of the local student model on the proxy dataset; Generate personalized learning demand vector: Based on the collected model performance data, the personalized demand analysis and knowledge evaluation agent uses large language model reasoning to generate a personalized learning demand vector reflecting knowledge short - boards; Retrieve external distilled knowledge: The personalized demand analysis and knowledge evaluation agent retrieves all accessible knowledge objects in the current round from the blockchain network to form a candidate knowledge object set; Execute knowledge adaptation evaluation: The personalized demand analysis and knowledge evaluation agent evaluates the value of external distilled knowledge fragments according to prediction entropy, output consistency, and historical fusion records, and outputs an adaptation degree score table.
[0013] Optionally, generating a personalized learning demand vector in S4 specifically includes: After each round of local model training, the personalized demand analysis and knowledge evaluation agent collects key indicators of the model in various categories, including The average loss for category c in step communication, denoted as: ; Where, represents the average training loss of the client on category , reflecting the "learning difficulty" of the model on this category, reflecting the overall fitting situation of the model on this category in the most recent communication rounds, is the communication round index, ranging from to ; is the communication round number at the start of the current statistical period, is the statistical window width, the number of historical communication rounds used to calculate the average value; refers to the cross - entropy loss of the student model on category in the communication round; It also includes collecting a data subset randomly selected from the proxy dataset The knowledge retention rate corresponding to the class accuracy on is expressed as: ; wherein, represents the knowledge retention rate of class and measures the degree of forgetting of the student model on class ; represents the proxy dataset uniformly issued by the system, which is a collection of public and non-private health status samples with a unified feature structure and label specification, and is used to achieve the alignment and evaluation of distilled knowledge among clients; represents a data subset randomly selected from the proxy dataset and serves as the actual sample set for the current client to perform model evaluation and knowledge retention rate calculation; represents a specific sample in the proxy dataset; then is the sample 's true label. When , it indicates that this sample belongs to class ; is the predicted label of the student model s for the input . If , it indicates that the model correctly discriminates the class of this sample; is an indicator function that takes the value 1 when the condition in the parentheses holds, otherwise 0, and is used to calculate the number of samples correctly predicted by the model for class c and the total number of samples in class c; is a very small positive number used to avoid the denominator being zero; is used to reflect the recent difficulty of class c for the model, then is used to quantify the retention of old knowledge of class c by the model, and These two indicators together constitute the basis of the personalized demand vector and are used to guide subsequent knowledge evaluation and screening, expressed as: ; wherein, represents the dimension value of class c in the personalized learning demand vector and quantifies the intensity of the external knowledge demand of the current client for class c; represents the total number of pre-defined health status classes by the system; is a hyperparameter representing the relative contribution of learning difficulty and knowledge retention rate to the demand degree. The vector is used to intuitively reflect the intensity of the external knowledge demand of the current client for each class.
[0014] Optionally, in S4, an adaptive knowledge selection and fusion guidance agent is used to formulate a knowledge adoption strategy and update the student model, including: Load the adaptation degree score table and personalized learning requirement vector: The adaptive knowledge selection and fusion guidance agent reads the adaptation degree score table and personalized learning requirement vector to construct a fused context environment; Screen knowledge fragments and set fusion weights: According to the adaptation degree and the intensity of external knowledge requirements, screen the set of knowledge objects that meet the teaching ability threshold, and assign fusion weights to each knowledge object; Construct a hybrid loss function: The adaptive knowledge selection and fusion guidance agent constructs a hybrid loss function that includes a local supervision term and an external distillation term. Each part in the distillation term has its contribution degree adjusted by the assigned fusion weight; Guide the model to perform distillation learning update: The client optimizes the parameters of the student model according to the constructed hybrid loss function to effectively absorb the external knowledge in this round.
[0015] Optionally, in S4, the adaptive knowledge selection and fusion guidance agent formulates a knowledge adoption strategy and updates the student model, specifically including: For each teacher model The given distilled knowledge , the adaptive knowledge selection and fusion guidance agent first averages according to the number of samples of each category in the proxy dataset to obtain the average logits vector of the teacher model for category c , thus, the adaptive knowledge selection and fusion guidance agent obtains the overall prediction tendency of the teacher model on each category c; Perform a softmax operation on the average logits vector for category c to obtain the corresponding probability distribution, expressed as: ; where, represents the predicted probability vector output by the teacher model on category c; is a normalization operation used to map the logits vector to a probability distribution with a sum of 1; respectively represent the normalized confidence scores of the model for each health status category; T represents vector transpose to make the result in the form of a column vector; Measure the "concentration degree" of this distribution through information entropy, expressed as: ; where, represents the information entropy of the probability distribution output by the teacher model j on category c; represents the current category index; represents the total number of pre-defined health status categories in the system; Denote the predicted probability of teacher model \(j\) for the \(k\)-th category under category \(c\); Denote the natural logarithm of the corresponding probability value; The adaptive knowledge selection and fusion guidance agent defines the confidence of teacher model \(j\) on category \(c\), denoted as: ; where, Denote the confidence. The lower the information entropy, the more concentrated the prediction, and thus the higher the confidence, indicating that the judgment of teacher model \(j\) on category \(c\) is more reliable; The adaptive knowledge selection and fusion guidance agent measures the prediction stability of teacher model \(j\) on category \(c\), including: statistically calculating the variance of the \(c\)-th dimension logits of all samples with label \(c\) in the surrogate dataset by teacher model \(j\), denoted as: ; where, Denote the output variance of the \(c\)-th dimension of logits of teacher model \(j\) on category \(c\), Denote the variance operation, Denote the average logits vector of category \(c\), Denote the component of the logits vector output by teacher model \(j\) for sample \(x\) on the \(c\)-th category, Denote that the sample comes from the surrogate dataset, Denote that the true label of this sample belongs to category \(c\); the larger the variance, the greater the fluctuation of the logits of teacher model \(j\) on category \(c\), and the less stable the knowledge; The adaptive knowledge selection and fusion guidance agent comprehensively considers the confidence and prediction stability, and calculates the teaching ability score of teacher model \(j\) on category \(c\), denoted as: ; where, and are hyperparameters for adjusting the relative importance of confidence and prediction stability, Denote the teaching ability score of teacher model \(j\) on category \(c\); Perform linear normalization on the teaching ability scores of all teacher models on the same category \(c\) to obtain .
[0016] To achieve the above object, the second aspect of this application provides a federated learning system based on multi-agent and knowledge distillation. The federated learning system includes: A local model training module, which is deployed in each client device participating in federated learning, independently trains a local model using the collected personal health data, generates a teacher model that can be used for knowledge distillation and a student model for receiving external knowledge distillation learning; The decoupled knowledge distillation module extracts standardized distillation knowledge fragments from the teacher model trained by the local model training module. The knowledge distillation process decouples the knowledge production step from the knowledge learning step; The blockchain-based knowledge sharing module encapsulates the distillation knowledge fragments and their metadata extracted by the decoupled knowledge distillation module into blockchain transactions and submits them to the blockchain network, verifies the legitimacy of the transactions, and writes them into the blockchain ledger to achieve decentralized distribution; The multi-agent-driven knowledge management module is used to evaluate the state of the student model through the personalized demand analysis and knowledge evaluation agent, generate a personalized learning demand vector, and analyze the external knowledge value to generate an evaluation result; according to the evaluation result and the personalized learning demand vector, through the adaptive knowledge selection and fusion guidance agent, complete the screening, weighting, and fusion strategy design of the external knowledge fragments; perform joint distillation operations on the selected external knowledge and the output of the local model, and construct a hybrid loss function including a true label supervision term and an external knowledge soft label term; according to the constructed hybrid loss function, guide the student model to update parameters to achieve the effective transfer of external knowledge to the local model.
[0017] To achieve the above object, the third aspect of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor is used to implement the method as described above.
[0018] After adopting the above technical solutions, the present application has the following beneficial effects compared with the prior art: The present application first conducts "category-level" personalized demand analysis, allowing each client to calculate the demand vector based on the accuracy and loss on each category, and the multi-agent locally screens the distillation knowledge fragments that best meet the needs of each category for each category to avoid the "one-size-fits-all" global distillation. Secondly, the client uploads the "knowledge" after regularizing the soft label, category annotation, and encapsulating metadata to the blockchain. All terminals can query on demand and dynamically screen and weight locally without the need for a central server (even a master node), achieving fully decentralized collaboration. Finally, only these compact distillation knowledge fragments are transmitted in each round of communication instead of the complete model parameters, and only a small batch of knowledge that highly matches the local "category-level demand" is downloaded, thus significantly reducing the communication overhead and improving the operating efficiency of resource-constrained terminals.
[0019] This application addresses issues such as model drift, slow convergence, and insufficient personalization performance in existing federated learning when dealing with non-independent and identically distributed data by introducing a large language model-driven multi-agent collaboration mechanism. The personalized learning requirements of each client model are sensed in real time by the agents and expressed in the form of dynamic learning requirement vectors, enabling each client to intelligently select the most suitable distilled knowledge according to its own state, achieve more targeted knowledge transfer and personalized improvement, and significantly enhance the model's adaptability to heterogeneous data.
[0020] Through the decoupled design of knowledge production and knowledge learning in this application, the distilled knowledge can be independently extracted, encapsulated, and shared after the model training is completed, greatly enhancing the modularity and flexibility of the federated learning architecture. At the same time, by sharing only soft labels or compressed distilled knowledge instead of complete model parameters, the communication load between clients is significantly reduced, effectively adapting to the Internet of Things environment where the computing power and network conditions of edge devices are limited, and improving the deployment feasibility of the overall system in real applications.
[0021] This application constructs a decentralized knowledge sharing mechanism based on blockchain, getting rid of the dependence on the central server in traditional federated learning and solving the performance bottleneck and single-point failure problems that may be brought by the central node. The immutability, transparency, and traceability provided by blockchain ensure the security and trustworthiness of the knowledge exchange process, enhance users' trust in data control of the system, and at the same time provide a solid foundation for the version management and reuse of knowledge.
[0022] The following further describes the specific implementation manners of this application in conjunction with the accompanying drawings in detail. Description of the Drawings
[0023] The accompanying drawings, as part of this application, are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application, but do not constitute an improper limitation to this application. Obviously, the accompanying drawings in the following description are only some embodiments, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0024] In the accompanying drawings: Figure 1 It is a schematic diagram of the steps of the federated learning method in this specific implementation manner; Figure 2 It is a schematic diagram of the training process of the federated learning method in this specific implementation manner; Figure 3 It is a schematic logical diagram of the federated learning system in this specific implementation manner; Figure 4Schematic diagram of the comparison of the test accuracy between the federated learning method and the existing algorithms when the standard image classification dataset is Fashion-MNIST and the heterogeneity parameter α is 0.1 in this specific embodiment; Figure 5 Schematic diagram of the comparison of the test accuracy between the federated learning method and the existing algorithms when the standard image classification dataset is Fashion-MNIST and the heterogeneity parameter α is 0.3 in this specific embodiment; Figure 6 Schematic diagram of the comparison of the test accuracy between the federated learning method and the existing algorithms when the standard image classification dataset is Fashion-MNIST and the heterogeneity parameter α is 0.05 in this specific embodiment; Figure 7 Schematic diagram of the comparison of the test accuracy between the federated learning method and the existing algorithms when the standard image classification dataset is CIFAR-10 and the heterogeneity parameter α is 0.1 in this specific embodiment; Figure 8 Schematic diagram of the comparison of the test accuracy between the federated learning method and the existing algorithms when the standard image classification dataset is CIFAR-10 and the heterogeneity parameter α is 0.3 in this specific embodiment; Figure 9 Schematic diagram of the comparison of the test accuracy between the federated learning method and the existing algorithms when the standard image classification dataset is CIFAR-10 and the heterogeneity parameter α is 0.05 in this specific embodiment; Figure 10 Schematic diagram of the comparison of the test accuracy between the federated learning method and the existing algorithms when the standard image classification dataset is CIFAR-100 and the heterogeneity parameter α is 0.1 in this specific embodiment; Figure 11 Schematic diagram of the comparison of the test accuracy between the federated learning method and the existing algorithms when the standard image classification dataset is CIFAR-100 and the heterogeneity parameter α is 0.3 in this specific embodiment; Figure 12 Schematic diagram of the comparison of the test accuracy between the federated learning method and the existing algorithms when the standard image classification dataset is CIFAR-100 and the heterogeneity parameter α is 0.05 in this specific embodiment. Specific Embodiment
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. The following embodiments are used to illustrate this application but are not used to limit the scope of this application.
[0026] Please refer to Figure 1 、 Figure 2 and Figure 3, Based on this, the present application provides a federated learning method based on multi-agent and knowledge distillation, including the following steps: S1. Train the local model: Independently train the local model using the collected personal health data to generate a teacher model for knowledge distillation and a student model for receiving external knowledge distillation learning. S2. Decoupled knowledge distillation: Extract standardized distillation knowledge fragments from the teacher model trained in S1. The knowledge distillation process decouples the knowledge production step from the knowledge learning step. S3. Decentralized knowledge sharing based on blockchain: Package the standardized distillation knowledge fragments extracted in S2 and their metadata into blockchain transactions and submit them to the blockchain network. Verify the legality of the transactions and write them into the blockchain ledger to achieve decentralized distribution. S4. Multi-agent collaborative knowledge management based on large language models to update the student model: Evaluate the status of the student model through personalized demand analysis and knowledge evaluation agents, generate a personalized learning demand vector, and analyze the value of external distillation knowledge to generate an evaluation result. According to the evaluation result and the personalized learning demand vector, through the adaptive knowledge selection and fusion guidance agent, complete the screening, weighting, and fusion strategy design of external distillation knowledge fragments; perform joint distillation operations on the selected external knowledge and the output of the local model, and construct a hybrid loss function including a true label supervision term and an external knowledge soft label term; according to the constructed hybrid loss function, guide the student model to update its parameters to achieve the effective transfer of external knowledge to the local model. S5. Repeat S1 - S4 until the training termination condition is met.
[0027] Specifically, collect personal health data using the health monitoring terminal on the user side, and independently train the model in the local environment of the terminal to generate a local model. This local model serves as both the teacher model for knowledge production and the student model for knowledge learning in the subsequent process, and is used to participate in knowledge distillation and model optimization in the federated learning process; among them, the health monitoring terminal on the user side includes: wearable devices, such as smart watches and health bracelets, home medical monitoring devices, such as electronic blood pressure monitors, blood glucose meters, and sleep monitors, and home health control devices with computing and communication capabilities.
[0028] Please refer to Figure 1 , Figure 2 and Figure 3, in a feasible implementation, training the local model in S1 includes: loading the local dataset and initializing the model structure: each federated learning client first loads the private dataset held locally by the client and initializes the local model structure according to the predefined model architecture, establishing the initial model parameters and weight matrices; performing the model training process: the client calls the standard optimization algorithm to perform iterative training on the local model. In a continuous number of training rounds, the model parameters are repeatedly updated based on the local data so that it can fully fit the data distribution characteristics of the client itself; determining the teacher model and the student model: after the training is completed, the client fixedly saves the current local model state as the teacher model in the current knowledge distillation process, and at the same time as the initial student model for receiving the distillation learning of external knowledge.
[0029] It should be noted that the local model training is an independent process executed on each client participating in federated learning. The purpose is to use the local private datasets held by the clients to train and optimize the parameters of the local machine learning model. This process uses standard machine learning algorithms, such as gradient descent algorithms and their variants, to iteratively update the local model weights, enabling the model to learn and fit the unique patterns and distributions of the local data, while ensuring that the original sensitive data always remains locally on the client and does not leak externally. The trained local model is not only a manifestation of the client's personalized capabilities but also serves as the "teacher" source for the subsequent knowledge distillation process.
[0030] Please refer to Figure 1 、 Figure 2 and Figure 3 , in a feasible implementation, the decoupled knowledge distillation in S2: the local model obtained by training through S1 is used as the teacher model to extract standardized distillation knowledge fragments on the home health control device with computing and communication capabilities. The knowledge distillation process decouples the knowledge production step from the knowledge learning step.
[0031] It should be noted that the decoupled knowledge distillation is a key step after the local model training on each client. Its core lies in transforming the complex knowledge learned by the local training model into one or more lighter-weight, easy-to-transmit-in-the-network and core-information-rich knowledge representation forms through knowledge distillation technology, which is called "distillation knowledge". A significant feature of this knowledge distillation process is to emphasize the decoupling of the knowledge production step from the knowledge learning step. This decoupling means that the extraction, representation, and encapsulation of knowledge, that is, the knowledge production step, and the subsequent sharing, selection, evaluation, and fusion application of knowledge, that is, the knowledge learning step, are independent of each other in mechanism, thus bringing higher flexibility, modularity, and scalability to the entire federated learning framework.
[0032] This knowledge distillation process emphasizes the "decoupling" feature. By first inferring and uploading the distilled knowledge of each client, it avoids the teacher-student dependence in traditional federated distillation methods, avoids uploading and downloading the teacher model, and instead downloads lightweight distilled knowledge, reducing communication overhead. In addition, knowledge regularization operations are performed on the logits, which not only suppresses extremely high confidence levels of overconfidence but also avoids overly flat low confidence levels, improving the learnability and stability of the distilled knowledge.
[0033] The teacher model is a machine learning model trained by each client using its local data during the local model training phase, or a snapshot of the model generated at a specific training iteration. In the decoupled knowledge distillation process, this teacher model serves as the source of knowledge, and its output is used as a "guidance signal" to train the student model. The selection and state of the teacher model are crucial for the quality of the distilled knowledge. In this embodiment, the local model of each client can act as the teacher model, thus contributing its unique knowledge learned based on local data.
[0034] Please refer to Figure 2 and Figure 3 , in an implementable embodiment, the knowledge production step in S2 includes: performing an inference operation on the proxy dataset: the client uses the teacher model trained in S1 to infer each sample in the proxy dataset uniformly distributed by the system, generating the original prediction results of the output layer. The proxy dataset is a publicly released non-sensitive health dataset; Extracting the model prediction logits: The client takes the original prediction vector logits of the output layer of the teacher model as the initial knowledge representation, recording the confidence level of the model for each category; Performing normalization processing: The client performs statistical normalization processing on the logits, calculates the mean and standard deviation, and normalizes them to have a uniform distribution characteristic on the numerical scale; Encapsulating the distilled knowledge fragment: The client organizes the normalized prediction results into a distilled knowledge fragment in a unified format, and adds client identification, timestamp, and current training round information to construct a structured knowledge object for S3 knowledge sharing.
[0035] Specifically, performing an inference operation on the proxy dataset: The proxy dataset is a publicly released non-sensitive health dataset that does not contain user privacy information and has predefined category labels and a unified feature format; Extracting the model prediction logits: Each dimension of the logits prediction vector corresponds to a set of health status categories predefined by the system, such as heart rate type, blood pressure level, sleep stage, exercise intensity level, blood oxygen risk level, etc.; Performing normalization processing: The client performs statistical normalization processing on the logits, calculates the mean and standard deviation, and normalizes them to have a uniform distribution characteristic on the numerical scale; Encapsulated distilled knowledge fragment: The client identifier is generated based on the unique identifier of the home health control device (such as MAC address, device public key, etc.), which is used to mark the knowledge source without exposing the user's identity; the timestamp is used to record the global time information when the knowledge is generated, which helps the client to judge the timeliness of the knowledge; the training round information is used to distinguish the model training stage to which the knowledge belongs, supporting subsequent version control and knowledge screening.
[0036] Please continue to refer to Figure 2 and Figure 3 In an implementable embodiment, the knowledge learning step in S2 includes: Loading external distilled knowledge: The client receives and loads the set of external distilled knowledge fragments obtained from the blockchain network as the input for distilled learning; the knowledge fragments are generated by other home health control devices based on a unified proxy dataset and uploaded to the blockchain, representing the prediction results of their local models for typical health state samples; by loading these external distilled knowledge fragments, the client can obtain the model cognitive information of other terminals in typical health scenarios without exposing the user's original health data, which is used to guide the personalized training of the local student model. Constructing a hybrid loss function: The client uses the hybrid loss function to optimize the student model. The hybrid loss function includes a local label supervision term and an external knowledge KL divergence term, and the fusion weight is specified by S4; the local label supervision term, that is, the client calculates the cross-entropy loss for the difference between the model output and the true label based on the labeled health data collected by its health monitoring terminal, such as known state samples from devices such as blood pressure monitors, heart rate belts, and blood glucose monitors, which is used to maintain the accurate fitting of the model to the individual physiological characteristics; the external knowledge KL divergence term, that is, the client calculates the KL divergence between the prediction distribution of the student model for the proxy data samples and the corresponding logits distribution in the external distilled knowledge loaded from the blockchain, which is used to guide the model to absorb the cognitive experience of other terminal devices in the same health scenario. Performing knowledge distilled learning: The client performs backpropagation and parameter update on the student model based on the constructed loss function to achieve the absorption and integration of external distilled knowledge.
[0037] Please refer to Figure 2 and Figure 3 In an implementable embodiment, S3, blockchain-based decentralized knowledge sharing: Encapsulate the distilled knowledge fragments extracted in S2 and their metadata into a blockchain transaction and submit it to the blockchain network, verify the legality of the transaction, and write it into the blockchain ledger to achieve decentralized distribution.
[0038] Specifically, the home health control device encapsulates the distilled knowledge fragments extracted in S2 and their metadata into blockchain transactions and submits them to the blockchain network. The home health control device also acts as a participating node in the blockchain network, responsible for the transaction broadcasting and consensus processes.
[0039] It should be noted that blockchain-based decentralized knowledge sharing provides a secure, transparent, traceable, and center-trust-free infrastructure for the release, storage, discovery, and acquisition of the distilled knowledge fragments generated by the aforementioned decoupled knowledge distillation. This mechanism is built on a blockchain network. When the client generates distilled knowledge fragments, these fragments themselves and their metadata, including the source client identifier, generation timestamp, etc., will be encapsulated into blockchain transactions and submitted to the chain. After being verified by the consensus mechanism, they will be packed into new blocks and permanently and immutably recorded on the chain. In this way, the secure storage of knowledge and high data availability are achieved. In addition, since there is no single centralized storage server, any access and retrieval of the knowledge already on the chain are directly carried out by querying the blockchain network itself, ensuring the decentralized nature of knowledge acquisition. Therefore, each distilled knowledge fragment can be safely, reliably, and truly decentralized shared and distributed among the clients participating in federated learning without relying on a trusted third party.
[0040] Please refer to Figure 2 and Figure 3 , in practical applications, S4. Multi-agent collaborative knowledge management based on large language models to update the student model: Deploy a knowledge management module composed of multiple agents on the home health control device. This module realizes intelligent reasoning and policy generation through remote calling of the large language model API, and is used to guide the personalized update of the student model; evaluate the state of the student model through the personalized demand analysis and knowledge evaluation agent, generate a personalized learning demand vector, and analyze the external knowledge value to generate an evaluation result; according to the evaluation result and the personalized learning demand vector, through the adaptive knowledge selection and fusion guidance agent, complete the screening, weighting, and fusion strategy design of external knowledge fragments; perform a joint distillation operation on the selected external knowledge and the output of the local model, and construct a hybrid loss function containing a real label supervision term and an external knowledge soft label term; according to the constructed hybrid loss function, guide the student model to update its parameters to achieve the effective migration of external knowledge to the local model.
[0041] It should be noted that the multi-agent in this embodiment completes personalized demand analysis, knowledge evaluation, adaptive knowledge selection, and knowledge fusion functions based on the reasoning ability of the large language model. Compared with the agent method in traditional reinforcement learning, the decision-making logic in a federated communication of this method does not depend on repeated interactions with the external environment or trial-and-error updates. Instead, by directly inputting the local model training status and candidate distilled knowledge into the large language model at one time, the demand vector, knowledge value score, and fusion weight can be directly obtained. While the method based on reinforcement learning must continuously interact with the environment, evaluate the reward obtained after a certain action, and then update the policy network to gradually find the optimal decision. And such a way of environment interaction - reward feedback - policy update requires a large number of samplings and multiple rounds of iterations to achieve convergence, which is time-consuming and has a large computational overhead. While the agent based on the LLM can quickly complete complex demand analysis and knowledge evaluation and other tasks by virtue of the reasoning ability of the LLM, greatly improving the decision-making efficiency.
[0042] In an implementable embodiment, generating the personalized learning demand vector in S4 specifically includes: after each round of local model training ends, the personalized demand analysis and knowledge evaluation agent collects the key indicators of the model in various categories, including The average loss for category c in step communication, denoted as: ; Among them, represents the average training loss of the client on category , reflecting the "learning difficulty" of the model on this category, reflecting the overall fitting situation of the model on this category in the most recent communication rounds, is the communication round index, ranging from to ; is the communication round number at the start of the current statistical period, is the statistical window width, the number of historical communication rounds used to calculate the average value; refers to the cross-entropy loss of the student model on category in the communication round; It also includes collecting the knowledge retention rate corresponding to the category accuracy on the data subset randomly sampled from the proxy dataset, denoted as: ; Among them, represents the knowledge retention rate of category , measuring the degree of forgetting of the student model on category ; Represents the proxy dataset uniformly distributed by the system, which is a collection of public and non-private health status samples with a unified feature structure and label specification, and is used to achieve the alignment and evaluation of distilled knowledge among clients; Represents a data subset randomly selected from the proxy dataset, which serves as the actual sample set for the current client to perform model evaluation and calculate the knowledge retention rate; Represents a specific sample in the proxy dataset; Then it is the sample 's true label. When , it indicates that this sample belongs to the category ; Is the predicted label of the student model s for the input . If , it indicates that the model correctly discriminates the category of this sample; Is the indicator function, which takes the value of 1 when the condition in the parentheses holds, otherwise 0, and is used to calculate the number of correctly predicted samples of category c by the model and the total number of samples of category c; Is a very small positive number, which is used to avoid the situation where the denominator is zero; Is used to reflect the recent difficulty of the model for category c, Then it is used to quantify the retention of old knowledge of the model for category c, And These two indicators together constitute the basis of the personalized demand vector, which is used to guide subsequent knowledge evaluation and screening, and is expressed as: ; Among them, Represents the dimension value of category c in the personalized learning demand vector, which quantifies the intensity of the external knowledge demand of the current client for category c; Represents the total number of predefined health status categories in the system; Is a hyperparameter, which represents the relative contribution of learning difficulty and knowledge retention rate to the demand degree. The vector Intuitively reflects the intensity of the external knowledge demand of the current client for each category.
[0043] In an implementable embodiment, in S4, the adaptive knowledge selection and fusion guidance agent formulates a knowledge adoption strategy and updates the student model, including: loading the fitness score table and the personalized learning demand vector: the adaptive knowledge selection and fusion guidance agent reads the fitness score table and the personalized learning demand vector to construct a fusion context environment; Screening knowledge fragments and setting fusion weights: According to the fitness and the intensity of external knowledge demand, screening the set of knowledge objects that meet the teaching ability threshold, and assigning fusion weights to each knowledge object; Construct a hybrid loss function: The adaptive knowledge selection and fusion guiding agent constructs a hybrid loss function that includes a local supervision term and an external distillation term. Each part of the distillation term is adjusted by the assigned fusion weight for its contribution; Guide the model to perform distillation learning update: The client optimizes the parameters of the student model according to the constructed hybrid loss function to effectively absorb the external knowledge in this round.
[0044] In an implementable embodiment, in S4, the adaptive knowledge selection and fusion guiding agent formulates a knowledge adoption strategy and updates the student model, specifically including: For each teacher model The given distilled knowledge , the adaptive knowledge selection and fusion guiding agent first averages according to the number of samples of each category in the proxy dataset to obtain the average logits vector of the teacher model for category c, , thus, the adaptive knowledge selection and fusion guiding agent obtains the overall prediction tendency of the teacher model on each category c; Perform a softmax operation on the average logits vector for category c to obtain the corresponding probability distribution, expressed as: ; where, represents the predicted probability vector output by the teacher model on category c; is a normalization operation used to map the logits vector to a probability distribution with a sum of 1; respectively represent the normalized confidence scores of the model for each health status category; T represents vector transposition to make the result in the form of a column vector; Measure the "concentration degree" of this distribution through information entropy, expressed as: ; where, represents the information entropy of the probability distribution output by the teacher model j on category c; represents the current category index; represents the total number of predefined health status categories in the system; represents the predicted probability of the teacher model j for the k-th category under category c; represents the natural logarithm of the corresponding probability value; The adaptive knowledge selection and fusion guiding agent defines the confidence of the teacher model j on category c, expressed as: ; where, represents the confidence level. The lower the information entropy, the more concentrated the prediction, and thus the higher the confidence level, indicating that the judgment of teacher model j on class c is more reliable. The adaptive knowledge selection and fusion guidance agent measures the prediction stability of teacher model j on class c, including: calculating the variance of the c -th dimension logits of all samples with label c in the surrogate dataset by teacher model j, denoted as: ; where, represents the output variance of the c -th dimension of logits of teacher model j on class c, represents the variance operation, represents the average logits vector of class c, represents the component of the logits vector output by teacher model j for sample x on the c -th class, represents that the sample comes from the surrogate dataset, represents that the true label of this sample belongs to class c; the larger the variance, the greater the fluctuation of the logits of teacher model j on class c, and the less stable the knowledge. The adaptive knowledge selection and fusion guidance agent comprehensively considers the confidence level and prediction stability, and calculates the teaching ability score of teacher model j on class c, denoted as: ; where, and are hyperparameters for adjusting the relative importance of the confidence level and prediction stability, represents the teaching ability score of teacher model j on class c; Perform linear normalization on the teaching ability scores of all teacher models on the same class c to obtain .
[0045] This embodiment is designed to solve the problem of "personalized knowledge learning in a decentralized environment and heterogeneous non - IID data scenarios": In applications such as the medical Internet of Things, each client has sensitive data with different distributions. Traditional federated learning is difficult to balance global performance and personalized needs under the conditions of no central server and limited bandwidth. The agent in this embodiment focuses on discovering the weaknesses of local models in each class and evaluating the effect of distilled knowledge from other parties on making up for these weaknesses, so as to achieve dynamic weighting and fusion at the class level and greatly alleviate client drift. In contrast, the existing reinforcement - learning - based solutions are usually used for optimizing "resource allocation", "cache scheduling" or "aggregation strategy", focusing on the dynamic allocation problems of physical - layer bandwidth, power, and cache space, and do not directly face "cross - node knowledge value evaluation and personalized distillation".
[0046] Please refer to Figure 2 and Figure 3 , based on the same inventive concept, the present application also provides a federated learning system based on multi-agent and knowledge distillation. The federated learning system includes: A local model training module, which is deployed in each client device participating in federated learning, and independently trains a local model using the collected personal health data to generate a teacher model for knowledge distillation and a student model for receiving external knowledge distillation learning; A decoupled knowledge distillation module, which extracts standardized distillation knowledge fragments from the teacher model trained by the local model training module, and decouples the knowledge production step and the knowledge learning step during the knowledge distillation process; A blockchain-based knowledge sharing module, which encapsulates the distillation knowledge fragments extracted by the decoupled knowledge distillation module and their metadata into blockchain transactions and submits them to the blockchain network, verifies the legitimacy of the transactions, and writes them into the blockchain ledger to achieve decentralized distribution; A multi-agent-driven knowledge management module, which is used to evaluate the status of the student model through personalized demand analysis and knowledge evaluation agents, generate a personalized learning demand vector, and analyze the external knowledge value to generate an evaluation result; according to the evaluation result and the personalized learning demand vector, through the adaptive knowledge selection and fusion guidance agent, complete the screening, weighting and fusion strategy design of external knowledge fragments; perform joint distillation operations on the selected external knowledge and the output of the local model, and construct a hybrid loss function including a true label supervision term and an external knowledge soft label term; according to the constructed hybrid loss function, guide the student model to update parameters to achieve the effective transfer of external knowledge to the local model.
[0047] Specifically, the local model training module is deployed in each client device participating in federated learning, and is used to independently train and update the parameters of the local model based on the local private data set held by the client. The module first loads the local data set and initializes it according to the predefined model structure; then executes the optimization process, and uses the stochastic gradient descent algorithm to iteratively update the model weights, and continuously trains for 5 epochs to make the model fully fit the current local data distribution of the client. After the training is completed, the generated local model has two functional attributes: First, it participates in the extraction process of distillation knowledge as a teacher model to generate local distillation knowledge fragments; Second, it acts as a student model to receive external distillation knowledge from other clients and completes self-update and optimization through the distillation learning mechanism. This module does not transmit any original data during the federated learning process, strictly ensuring data privacy and security, and realizing the autonomous learning ability of the model under the individual data distribution through local training, providing a basic support for the subsequent decoupled knowledge distillation module and multi-agent knowledge management module.
[0048] Specifically, the decoupled knowledge distillation module is used to extract representative knowledge content based on the inference ability of the current client's local model after each round of local model training, and convert it into distilled knowledge fragments in a unified format for subsequent knowledge sharing and fusion learning. The design of the decoupled knowledge distillation module emphasizes the complete decoupling of the knowledge generation process and the knowledge usage process, enabling the extraction, standardization, and encapsulation operations of knowledge to be executed independently of specific learning objectives, thereby enhancing the modularity and operational flexibility of the federated learning system in a decentralized architecture. First, the client uses the selected teacher model to perform per-sample inference operations on the proxy dataset uniformly distributed by the system and collects the original prediction results generated at the output layer. The prediction results are unnormalized floating-point vectors representing the model's discrimination confidence for each category. Next, the client performs standardization processing on the extracted original prediction results. Specifically, the mean and standard deviation of the vector set are calculated respectively, and each item is normalized based on this statistic to make it have a uniform distribution characteristic in the numerical scale. This step helps to eliminate the differences between the outputs of each client model and improve the generality and fusibility of the knowledge fragments when applied across clients. After that, the client organizes the standardized prediction results into structured distilled knowledge fragments and attaches the unique identifier of the client, the current system timestamp, and the information of the training round to form a complete knowledge object. The knowledge object adopts a predefined data structure to facilitate storage, verification, retrieval, and invocation operations in the subsequent blockchain module. The distilled knowledge fragments include both the normalized logits information from the output layer of the teacher model and can be extended to intermediate layer activation values or their transformed forms, with pluggability and abstract expression capabilities.
[0049] Specifically, the blockchain-based knowledge sharing module is used to host the entire process of storage, publication, and access of distilled knowledge fragments between clients, constructing a trusted knowledge circulation environment without a central server. The blockchain-based knowledge sharing module takes blockchain technology as the underlying infrastructure, combines the consensus mechanism with the immutable data structure, realizes the decentralized registration, verification, and persistence of distilled knowledge objects, and ensures the secure transmission, efficient sharing, and transparent use of knowledge in a multi-client environment. Specifically, after completing the production and encapsulation of the distilled knowledge fragments, the client submits them as knowledge objects to be published to the blockchain network. The knowledge object contains standardized prediction data, the identity identifier of the generating client, the current timestamp, and fields corresponding to its training rounds. All fields are packaged according to a predefined format to generate a blockchain transaction record. The module first receives the knowledge transaction broadcast from the client and verifies its legality. This verification process is jointly executed by the nodes participating in the blockchain consensus to ensure the integrity of the transaction data, the standardization of the format, and the validity of the identity of the generating party. After passing the verification, the transaction is packaged into the current block and permanently written into the blockchain ledger as part of the new block. The blockchain ledger is a distributed data structure jointly maintained and publicly accessible by all federated learning clients. Each client can query the distilled knowledge fragments of specific rounds, specific clients, or specific categories from the blockchain through a standard interface. In addition, the module supports high-frequency concurrent writing and reading operations to meet the knowledge sharing requirements in large-scale federated learning scenarios. Without relying on any central scheduling node, each client can independently retrieve and select the knowledge objects published by other nodes according to its own learning needs, improving the overall response speed, scalability, and fault tolerance of the federated system.
[0050] Specifically, the multi-agent-driven knowledge management module is used to realize the personalized identification, evaluation, and fusion update of external knowledge on the client side in a decentralized federated learning architecture. The multi-agent-driven knowledge management module is based on the large language model deployed locally on the client, constructs a knowledge management system composed of multiple agents with autonomous decision-making capabilities, and operates dynamically around the current learning state of the student model. The multi-agent-driven knowledge management module includes two core sub-agents: the personalized demand analysis and knowledge evaluation agent and the adaptive knowledge selection and fusion guidance agent. The two cooperate to complete the knowledge value judgment and the formulation of the distilled integration strategy.
[0051] The personalized demand analysis and knowledge evaluation agent is used to comprehensively perceive the training performance of the current student model. By analyzing the learning losses, prediction accuracies, and retention capabilities of various categories on the proxy dataset, it generates a personalized learning demand vector reflecting the knowledge gaps. This agent first conducts per-class statistics on the outputs of the local model on the proxy dataset, calculates the sliding average loss value for each category and the accuracy metric for historical samples, respectively quantifying the current learning difficulty and retention level of the model for each category. Based on the above metrics, this agent generates a learning demand vector based on the reasoning ability of the large language model to characterize the knowledge demand intensity of each category in the current round. After completing the analysis of the local learning state, the agent retrieves the currently available external distilled knowledge objects from the blockchain and evaluates their applicability and potential value in the current learning scenario according to the source categories, confidence levels, and prediction consistencies of each knowledge fragment. This evaluation process comprehensively considers the entropy value of the knowledge confidence distribution, the output variance within the category, and the historical selection weights, and conducts context-related reasoning through the large language model to form a knowledge fragment adaptation list for the current client, providing a high-quality decision-making basis for subsequent knowledge selection.
[0052] The adaptive knowledge selection and fusion guidance agent designs the screening, weighting, and fusion strategies for external knowledge fragments according to the aforementioned evaluation results and learning demand vector. This agent first sorts each knowledge fragment based on the teaching ability index in units of categories, sets a dynamic threshold, and screens the candidate knowledge with a teaching ability score higher than the threshold according to the current demand degree of the client for this category and the historical learning difficulty. Then, on the premise of meeting the teaching diversity constraint, this agent selects knowledge sources with distribution differences from the candidate set to avoid redundancy and overlap during fusion. This agent assigns a fusion weight to each selected knowledge fragment, and the weight value is generated by normalizing the teaching ability score, reflecting the actual contribution degree of this knowledge fragment to the learning objective. Subsequently, this agent performs a joint distillation operation on the selected knowledge and the local model output, constructs a hybrid loss function containing a real label supervision term and an external knowledge soft label term, where the external knowledge part is calculated by weighting with the fusion weight. This agent guides the student model to update parameters based on this loss function to achieve the effective transfer of external knowledge to the local model.
[0053] To verify the effectiveness and practicality of the federated learning system based on large language model multi-agent and knowledge distillation proposed in this application in a heterogeneous environment, this embodiment conducts systematic experiments on three standard image classification datasets (Fashion-MNIST, CIFAR-10, and CIFAR-100) and makes a comparative evaluation with six current mainstream federated learning algorithms (FedAvg, MOON, ProxyFL, FedGKD, FedHKD).
[0054] To verify the robustness and adaptability of the system under different data heterogeneity conditions, this embodiment constructs a variety of non-independent and identically distributed (Non-IID) scenarios. By introducing the Dirichlet distribution for data partitioning on three standard image classification datasets (Fashion-MNIST, CIFAR-10, and CIFAR-100), a heterogeneous environment is simulated. The experimental results are as Figures 4 to 12 shown. In this embodiment, different heterogeneity parameters α (taking values of 0.05, 0.1, and 0.3) are set. The smaller α is, the greater the difference in data distribution among clients, and the more severe the learning challenge for the system. The experimental results are as Figures 4 to 12 shown. The proposed system demonstrates excellent adaptability under various degrees of data heterogeneity. Especially in the extreme heterogeneous scenario, when α is 0.05, the system achieves accuracies of 96.66% and 90.96% on the Fashion-MNIST and CIFAR-10 datasets respectively, significantly outperforming other comparison methods. In contrast, traditional parameter averaging methods such as FedAvg and MOON experience a significant drop in accuracy under heterogeneous conditions, indicating their sensitivity to data distribution differences. In addition, from the convergence trend of the model under different α values, the system's learning curve is smoother and the convergence speed is faster. Especially in datasets with a larger number of classes and greater learning difficulty such as CIFAR-100, the system still maintains the lead. This shows that while maintaining the ability of personalized learning, this method can effectively integrate multi-source heterogeneous knowledge and alleviate the performance degradation problem faced by traditional federated learning in high-heterogeneity environments.
[0055] In addition to demonstrating good generalization ability in high-heterogeneity data environments, this system also shows outstanding advantages in terms of communication efficiency and resource usage. The evaluation of system efficiency mainly focuses on two aspects: the number of communication convergence rounds and the communication and computing overhead per unit round. The experimental results are shown in Table 1: Table 1: Comparison of communication overhead and computing cost per unit round required for different methods to reach 70% accuracy under the condition of CIFAR-10 (α = 0.1)
[0056] First, in terms of the number of communication rounds, the system can achieve effective convergence of the model within a significantly smaller number of communication rounds. Taking the CIFAR-100 dataset as an example, when the heterogeneity parameter α is 0.3, the system can approach the final performance level within approximately 50 communication rounds, while other comparison methods such as FedAvg and MOON are still in the early stage of convergence within the same number of rounds. This accelerated convergence speed directly reduces the overall communication volume of the system, which is particularly important for practical scenarios with limited bandwidth or device energy consumption.
[0057] Secondly, in terms of the communication overhead per round, the system replaces the direct exchange of large-scale model parameters in traditional federated learning with a knowledge distillation mechanism, and only needs to transmit compressed distillation knowledge fragments, significantly reducing the data transmission volume per round. In contrast, FedAvg requires the synchronization of full model parameters, and methods such as MOON and FedHKD also need to introduce additional model embeddings or high-order feature representations, significantly increasing the communication burden. The system effectively avoids this problem in its structural design, and the communication cost is more controllable.
[0058] It should be noted that the system introduces an intelligent agent mechanism based on large language models on the client side to execute knowledge evaluation and fusion strategy generation. The invocation of remote agents does increase the computing time to some extent during the local training process. However, since it can significantly reduce the total number of communication rounds required, from the overall operating efficiency perspective, while reducing the total communication volume, the system does not bring a significant increase in time cost.
[0059] Generally speaking, the system and method provided in this embodiment have obvious advantages in processing heterogeneous data, realizing personalized modeling, ensuring communication efficiency and system security, and have broad promotion value and strong engineering feasibility.
[0060] Based on the same inventive concept, the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the method as described above.
[0061] The program product of the present application for implementing the above method can be a portable compact disc read-only memory and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present application is not limited to this. In the present application, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component.
[0062] It should be noted that the computer-readable storage medium can include data signals propagated in a baseband or as part of a carrier wave, which carry the readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or component. The program code contained on the readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0063] The above are only the preferred embodiments of the present application, and there is no restriction on the present application in any form. Although the present application has been disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art of the present application can make some changes or modifications to equivalent embodiments with equivalent changes by using the technical content prompted above within the scope of the technical solution of the present application. The implementation schemes in the above embodiments can also be further combined or replaced. However, as long as the content does not deviate from the technical solution of the present application, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application still belong to the scope of the present application solution.
Claims
1. A federated learning method based on multi-agent and knowledge distillation, characterized in that It includes the following steps: S1. Train the local model: Independently train the local model using the collected personal health data to generate a teacher model for knowledge distillation and a student model for receiving external knowledge distillation learning; S2. Decoupled knowledge distillation: Extract the standardized distillation knowledge fragments from the teacher model trained in S1. The knowledge distillation process decouples the knowledge production step from the knowledge learning step; S3. Blockchain-based decentralized knowledge sharing: Package the standardized distillation knowledge fragments extracted in S2 and their metadata into blockchain transactions and submit them to the blockchain network, verify the transaction legality, and write them into the blockchain ledger to achieve decentralized distribution; S4. Multi-agent collaborative knowledge management based on large language models to update the student model: Evaluate the status of the student model through personalized demand analysis and knowledge evaluation agents, generate personalized learning demand vectors, and analyze the value of external distillation knowledge fragments to generate evaluation results; According to the evaluation results and personalized learning demand vectors, through the adaptive knowledge selection and fusion guidance agent, complete the screening, weighting, and fusion strategy design of external distillation knowledge fragments; perform joint distillation operations on the selected external distillation knowledge fragments and the output of the local model, and construct a hybrid loss function containing a true label supervision term and an external knowledge soft label term; according to the constructed hybrid loss function, guide the student model to update its parameters to achieve the effective transfer of external distillation knowledge fragments to the local model; S5. Repeat S1 - S4 until the training termination condition is met.
2. The method according to claim 1, characterized in that, The knowledge production step in S2 includes: Execute the proxy dataset inference operation: The client uses the teacher model trained in S1 to infer each sample in the proxy dataset uniformly issued by the system, and generate the original prediction results of the output layer. The proxy dataset is a publicly released non-sensitive health dataset; Extract the model prediction logits: The client takes the original prediction vector logits of the teacher model output layer as the initial knowledge representation and records the confidence level of the model for each category; Execute the standardization process: The client performs statistical standardization on the logits, calculates the mean and standard deviation, and normalizes the logits so that the logits have a uniform distribution characteristic on the numerical scale; Package the distillation knowledge fragments: The client organizes the standardized prediction results into distillation knowledge fragments in a unified format, that is, external distillation knowledge fragments, and adds client identification, timestamp, and current training round information to construct a structured knowledge object for S3 knowledge sharing.
3. The method according to claim 1, wherein The knowledge learning step in S2 includes: Load external distillation knowledge: The client receives and loads the set of external distillation knowledge fragments obtained from the blockchain network as the input for distillation learning; Construct a hybrid loss function: The client uses the hybrid loss function to optimize the student model. The hybrid loss function contains a local label supervision term and an external knowledge KL divergence term, and the fusion weight is specified by S4; Execute knowledge distillation learning: The client performs backpropagation and parameter update on the student model based on the constructed hybrid loss function to achieve the absorption and integration of external distillation knowledge.
4. The method according to claim 1, wherein The knowledge sharing process in S3 includes: Submitting Knowledge Fragments to the Blockchain Network: The client broadcasts the constructed distilled knowledge fragments as transaction records to the blockchain network; Performing Transaction Legality Verification: The consensus nodes in the blockchain network verify the transaction content, checking the format, identity signature, and data integrity of the transaction content; Writing to the Blockchain Ledger: After passing the verification, the transaction is packaged into a new block and recorded in the blockchain ledger, achieving the immutable storage of knowledge objects.
5. The method according to claim 1, characterized in that In S4, the personalized demand analysis and knowledge evaluation agent evaluates the status of the student model and analyzes the value of external knowledge, including: Collecting Model Performance Data: The personalized demand analysis and knowledge evaluation agent monitors and statistics the prediction loss and accuracy of the local student model on the proxy dataset; Generating a Personalized Learning Demand Vector: The personalized demand analysis and knowledge evaluation agent generates a personalized learning demand vector reflecting knowledge short - boards through large - language model inference based on the collected model performance data; Retrieving External Distilled Knowledge: The personalized demand analysis and knowledge evaluation agent retrieves all currently accessible knowledge objects from the blockchain network to form a candidate knowledge object set; Performing Knowledge Adaptation Evaluation: The personalized demand analysis and knowledge evaluation agent evaluates the value of external distilled knowledge fragments according to prediction entropy, output consistency, and historical fusion records, and outputs an adaptation degree score table.
6. The method according to claim 5, characterized in that, Generate a personalized learning requirement vector in S4, which specifically includes: after each round of local model training, the personalized requirement analysis and knowledge evaluation agent collects the key metrics of the model in various categories, including The average loss for category c in step communication, denoted as: ; Among them, represents the average training loss of the client on the category , reflecting the "learning difficulty" of the model on this category, and reflecting the overall fitting situation of the model on this category within the most recent communication rounds, is the communication round index, ranging from to ; is the communication round number at the start of the current statistical period, is the statistical window width, which is the number of historical communication rounds used to calculate the average value; refers to the cross-entropy loss of the student model on the category in the communication round. It also includes collecting a data subset randomly sampled from the proxy dataset The knowledge retention rate corresponding to the class accuracy on ; Among them, represents the knowledge retention rate of the category , which measures the degree of forgetting of the student model in the category ; represents the proxy dataset uniformly distributed by the system, which is a collection of public and non-private health status samples with a unified feature structure and label specification, and is used to achieve the alignment and evaluation of distilled knowledge among clients; represents a data subset randomly selected from the proxy dataset, which is used as the actual sample set for model evaluation and knowledge retention rate calculation by the current client; represents a specific sample in the proxy dataset; is the sample 's true label. When , it means that the sample belongs to the category ; is the predicted label of the student model s for the input . If , it means that the model correctly discriminates the category of the sample; is an indicator function that takes the value 1 when the condition in the parentheses holds, otherwise 0, and is used to calculate the number of correctly predicted samples of category c and the total number of samples of category c; is a very small positive number used to avoid the case of a zero denominator; to reflect the recent difficulty of the model for category c is used to quantify the retention of old knowledge of the model for category c and These two indicators together form the basis of the personalized demand vector, which is used to guide subsequent knowledge evaluation and screening, expressed as: ; Among them, represents the dimension value of category c in the personalized learning demand vector, quantifying the intensity of the current client's external knowledge demand for category c; represents the total number of predefined health status categories in the system; is a hyperparameter, representing the relative contributions of learning difficulty and knowledge retention rate to the demand degree. The vector is used to intuitively reflect the intensity of the current client's external knowledge demand in various categories.
7. The method according to claim 1, wherein In S4, the adaptive knowledge selection and fusion guidance agent formulates a knowledge adoption strategy and updates the student model, including: Loading the Adaptation Degree Score Table and the Personalized Learning Demand Vector: The adaptive knowledge selection and fusion guidance agent reads the adaptation degree score table and the personalized learning demand vector to construct a fusion context environment; Filtering Knowledge Fragments and Setting Fusion Weights: According to the adaptation degree and the intensity of external knowledge requirements, filter the set of knowledge objects that meet the teaching ability threshold, and assign a fusion weight to each knowledge object; Constructing a Hybrid Loss Function: The adaptive knowledge selection and fusion guidance agent constructs a hybrid loss function containing a local supervision term and an external distilled term, and the contribution degree of each part in the distilled term is adjusted by the assigned fusion weight; Guiding the Model to Execute Distillation Learning Update: The client optimizes the parameters of the student model according to the constructed hybrid loss function to achieve the effective absorption of external knowledge in this round.
8. The method according to claim 7, characterized in that In S4, the adaptive knowledge selection and fusion guidance agent formulates a knowledge adoption strategy and updates the student model, specifically including: For each teacher model The given distilled knowledge , the adaptive knowledge selection and fusion guiding agent first averages according to the number of samples of each category in the proxy dataset to obtain the teacher model The average logits vector for category c , thereby, the adaptive knowledge selection and fusion guiding agent obtains the teacher model The overall prediction tendency on each category c; The average logits vector for class c Perform a softmax operation to obtain the corresponding probability distribution, denoted as: ; Among them, represents the predicted probability vector output by the teacher model on class c; is a normalization operation used to map the logits vector to a probability distribution with a sum of 1; respectively represent the normalized confidence scores of the model for each health status category; T represents vector transpose to make the result in the form of a column vector; Measuring the "concentration degree" of this distribution through information entropy, expressed as: ; Among them, represents the information entropy of the probability distribution output by the teacher model j for class c; represents the current class index; represents the total number of predefined health status classes in the system; represents the prediction probability of the teacher model j for the k-th class under class c; represents the natural logarithm of the corresponding probability value; The adaptive knowledge selection and fusion guidance agent defines the confidence of teacher model j on class c, expressed as: ; Among them, represents the confidence. The lower the information entropy, the more concentrated the prediction, and thus the higher the confidence, indicating that the judgment of the teacher model j on the category c is more reliable; The adaptive knowledge selection and fusion guidance agent measures the prediction stability of teacher model j on class c, including: statistically calculating the variance of the c - th dimension logits of all samples with label c in the proxy dataset by teacher model j, expressed as: ; Among them, represents the output variance of the c-th dimension of the logits of the teacher model j on the category c, represents the variance operation, represents the average logits vector of the category c, represents the component of the logits vector output by the teacher model j for the sample x on the c-th category, indicates that the sample comes from the proxy dataset, indicates that the true label of the sample belongs to the category c; the larger the variance, the greater the fluctuation of the logits of the teacher model j on the category c, and the less stable the knowledge; The adaptive knowledge selection and fusion guidance agent comprehensively considers the confidence and prediction stability, and calculates the teaching ability score of teacher model j on class c, expressed as: ; Among them, and are hyperparameters for adjusting the relative importance of confidence and prediction stability, represents the teaching ability score of teacher model j on class c; Perform linear normalization on the teaching ability scores of all teacher models for the same category c to obtain . 9. A federated learning system based on multi-agent and knowledge distillation, characterized in that, The said federated learning system includes: The local model training module is deployed in each client device participating in federated learning, and independently trains a local model using the collected personal health data to generate a teacher model for knowledge distillation and a student model for receiving external knowledge distillation learning; The decoupled knowledge distillation module extracts standardized distillation knowledge fragments from the teacher model trained by the local model training module, and decouples the knowledge production step and the knowledge learning step during the knowledge distillation process; The blockchain-based knowledge sharing module encapsulates the distillation knowledge fragments and their metadata extracted by the decoupled knowledge distillation module into blockchain transactions and submits them to the blockchain network, verifies the legitimacy of the transactions, and writes them into the blockchain ledger to achieve decentralized distribution; The multi-agent-driven knowledge management module is used to evaluate the state of the student model through the personalized demand analysis and knowledge evaluation agent, generate a personalized learning demand vector, and analyze the external knowledge value to generate an evaluation result; according to the evaluation result and the personalized learning demand vector, through the adaptive knowledge selection and fusion guidance agent, complete the screening, weighting and fusion strategy design of the external knowledge fragments; perform joint distillation operations on the selected external knowledge and the output of the local model, and construct a hybrid loss function containing a true label supervision term and an external knowledge soft label term; according to the constructed hybrid loss function, guide the student model to update its parameters to achieve the effective transfer of external knowledge to the local model.
10. A computer-readable storage medium storing a computer program, characterized in that, When executed by a processor, the computer program is used to implement the method according to any one of claims 1-8.
Citation Information
Patent Citations
Multi-agent reinforcement learning method and system
CN113592100A
Federal learning model aggregation method based on dynamic adaptive knowledge distillation
CN116681144A
Block chain enhanced cyclic distillation guide channel decoupling framework for personalized federated learning
CN116701928A
Personalized federal learning method based on decoupling knowledge distillation
CN117152480A
Decentralization federated learning framework and method based on mutual distillation
CN117808115A
Cited By
Multi-intersection traffic signal cooperative control method driven by cross attention neural network
CN120636164A
Federal learning-based space-time diagram prompting method and system
CN120670998A
Current sensor fault isolation and signal reconstruction method, device, equipment and medium
CN121542944A
Current sensor fault isolation and signal reconstruction method, device, equipment and medium
CN121542944B
Multi-band infrared confrontation sample generation method and device, equipment and medium
CN121861184A