Triggered federal continuous learning scheduling and key data sharing method

Through the federated continuous learning method of real-time performance monitoring of edge devices and critical data screening, the problems of waste of resources and insufficient adaptability in the existing technology are solved, efficient and privacy-protected model updates and data sharing are achieved, and the system's adaptability in dynamic environments is improved.

CN120387526APending Publication Date: 2025-07-29BEIJING JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510430157.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing federated continuous learning methods lack dynamic triggering mechanisms associated with model performance, resulting in wasted resources and slow response to performance declines, and failed to prioritize the selection of devices with critical data to participate in training, making it difficult to balance protecting privacy and improving model performance, limiting the system's ability to adapt to complex dynamic environments.

Method used

Through edge devices in real time monitoring performance and uploading evaluation reports, server analysis triggers the federated learning process, selecting devices holding key data for local model training and critical data screening, server-side aggregation updates the global model, and using key data identification algorithms and privacy protection processing.

Benefits of technology

It realizes efficient triggering of federated learning when performance declines, reduce resource consumption, and prioritizes the use of high-value data to improve model adaptability, which significantly improves the system's adaptability and performance in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387526A_ABST
    Figure CN120387526A_ABST
Patent Text Reader

Abstract

The invention provides a triggering type federal continuous learning scheduling and key data sharing method. The method comprises the following steps that: edge equipment continuously performs environmental data collection and performance monitoring, and uploads a performance evaluation report to a server; the server analyzes the performance evaluation report and judges whether a federal learning process is triggered or not; the server selects an edge device holding key data to participate in federal learning; the edge device selected to participate in federated learning carries out local model training, a key data sample is screened out through a key data identification algorithm, and the edge device uploads the updated local model and the key data sample to a server; and the server side aggregates the local models uploaded by the edge devices, generates a new global model through a federal aggregation algorithm, and issues the new global model to the edge devices. According to the method, unnecessary communication and computing resource consumption is reduced, the overall efficiency of the system is improved, and the method has wide application prospects in the fields of automatic driving, intelligent medical treatment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of federated learning, and in particular, to a resource management method and apparatus in a multimedia communication system. Background Art

[0002] With the popularization of intelligent devices, federated learning has become the main paradigm for addressing data privacy and distributed machine learning requirements. Federated Continual Learning (FCL), as an emerging technology, aims to solve the adaptability problem of models in the face of continuously changing data distributions. In practical applications, the environments and scenarios faced by edge devices are constantly changing, such as autonomous vehicles encountering new road types, smartphone users generating new usage behaviors, etc., which makes the data distribution show temporal changes. Traditional federated learning does not fully consider this temporal change, and is prone to "catastrophic forgetting", that is, the learning of new knowledge leads to the forgetting of old knowledge, seriously affecting the overall performance of the model.

[0003] However, there are three key problems in existing FCL methods: First, most of them adopt a fixed-period model update strategy, updating regardless of whether the model performance is stable, resulting in a large amount of unnecessary communication and computational overhead; Second, the existing device scheduling mechanism mainly focuses on communication efficiency and resource balance, and has not incorporated the importance of the data held by the device into the core consideration factors of the scheduling decision, resulting in limited utilization efficiency of key data during the training process; Third, the pure parameter sharing method faces challenges in model adaptability when facing abnormal scenarios or rare events.

[0004] In practical applications, such as intelligent transportation systems, vehicles are driving in different regions, different weather and traffic conditions. It is difficult to quickly adapt to these changes only by model parameter updates. Selectively uploading key scenario data (such as rare traffic signs, abnormal driving behaviors, etc.) can significantly accelerate model adaptation and improve safety. Similarly, in a medical health monitoring system, when rare symptoms are detected, selectively sharing this key data can improve the early diagnosis ability of the entire system for important diseases, while avoiding privacy risks brought by large-scale patient data sharing.

[0005] Currently, the technical solutions of federated continual learning in the prior art mainly focus on the optimization of device scheduling, data sharing, and learning efficiency. In terms of device scheduling, there are solutions that propose a device selection method based on the multi-armed bandit problem, aiming to maximize the number of edge device nodes; there are also solutions that adopt a random access mechanism to allow devices to randomly select an upload channel; there are also solutions that transform device scheduling into a multi-armed bandit and matching process to optimize the latency problem in multi-task federated learning.

[0006] In terms of continuous learning and environmental adaptation, existing technologies have proposed cross-environmental federated continuous learning methods, which use a two-layer model structure and layer replacement methods to achieve rapid adaptation to new environments; there are also solutions that allow terminals to use continuously collected data for online learning and immediately integrate new data into the training process; there are also solutions that design a transmission scheduling mechanism for the vehicle networking environment to improve communication efficiency by broadcasting on service channels and allocating dedicated time slots; there are also solutions that address the challenges of federated learning in data heterogeneous environments through knowledge distillation and supervised contrast learning.

[0007] The disadvantages of the above-mentioned technical solutions for federated continuous learning in the prior art include: in terms of training triggering, these existing technical solutions lack a dynamic mechanism associated with model performance and cannot adaptively adjust the training process according to actual needs, resulting in waste of system resources or being slow to respond to performance degradation. In addition, existing methods mostly rely on preset rules or only consider the computing power and communication status of devices, and fail to incorporate the importance of the data held by devices into the core factors of scheduling decisions, unable to preferentially select devices with key data to participate in training, and reducing the adaptation efficiency to time-varying environments and time-varying data distributions.

[0008] In the data sharing mechanism, existing technologies usually adopt a pure parameter sharing mode. Although it protects data privacy, when dealing with abnormal scenarios or rare events, the model cannot obtain key information for learning, severely limiting the system's adaptability to important scenarios. Existing federated learning methods lack an effective mechanism to identify and selectively upload high-value data, and it is difficult to strike a balance between protecting privacy and improving model performance, restricting the application potential of federated continuous learning in complex dynamic environments. Summary of the Invention

[0009] Embodiments of the present invention provide a method for scheduling and key data sharing in trigger-based federated continuous learning to effectively improve the efficiency of federated continuous learning.

[0010] To achieve the above object, the present invention adopts the following technical solutions.

[0011] A method for scheduling and key data sharing in trigger-based federated continuous learning includes:

[0012] Edge devices continuously collect environmental data and monitor performance, and upload performance evaluation reports to the server;

[0013] The server analyzes the performance evaluation reports and determines whether to trigger the federated learning process based on the performance evaluation reports of the edge devices;

[0014] After triggering the federated learning process, the server selects edge devices that hold key data to participate in the federated learning;

[0015] The edge devices selected to participate in federated learning perform local model training, and filter out the key data samples that have the greatest impact on the model performance through the key data recognition algorithm. The edge devices upload the updated local model and key data samples to the server;

[0016] The server aggregates the local models uploaded by each edge device, generates a new global model through the federated aggregation algorithm, and distributes the new global model to each edge device.

[0017] Preferably, the edge devices continuously collect environmental data and monitor performance, and upload a performance evaluation report to the server, including:

[0018] The edge device collects environmental data at a certain frequency, processes the environmental data using the local model, calculates performance metrics, generates a performance evaluation report based on all performance metrics. The performance evaluation report contains the performance metric values and their change trends, detected data distribution changes, statistical characteristics of local data, the ID of the edge device, and the timestamp. The edge device uploads the performance evaluation report to the server through the network.

[0019] Preferably, the environmental data includes various parameters collected by the edge device about its surrounding environment. In the autonomous driving scenario, the environmental data includes vehicle speed, road conditions, weather conditions, and dynamic information of surrounding vehicles and pedestrians. In the remote medical scenario, the environmental data includes the vital signs and health condition data of the patient.

[0020] Preferably, the server analyzes the performance evaluation report and determines whether to trigger the federated learning process based on the performance evaluation report of the edge device, including:

[0021] The server receives and stores the performance evaluation reports submitted by each edge device, analyzes the performance evaluation reports, obtains the performance change trends of each edge device, and obtains the overall performance metrics of the system and their change rates based on the performance change trends of each edge device;

[0022] The server compares the overall performance metrics of the system with a preset threshold, and determines whether to trigger the federated learning process according to the comparison result. If the overall performance of the system drops to the preset threshold, the federated learning is triggered; if the overall performance of the system is higher than the preset threshold, and if the performance is stable and higher than the target requirement, the federated learning process is not triggered, and continue to monitor. The preset threshold is dynamically adjusted based on the importance of the application scenario and resource constraints.

[0023] Preferably, when the federated learning process is triggered, the server selects the edge devices holding the key data to participate in the federated learning, including:

[0024] After triggering the federated learning process, the server selects edge devices that hold key data to participate in federated learning. The server sends a request to participate in federated learning to the selected edge devices. This request includes training configurations and an indication of whether key data needs to be uploaded. The training configurations include federated learning strategies, learning rates, and local training rounds.

[0025] After receiving the request sent by the server, the edge device evaluates its own computing resources and data status and decides whether to participate in this round of federated learning. The edge device that accepts to participate in this round of federated learning sends a confirmation message to the server. This confirmation message includes device status and available resource information.

[0026] Based on the confirmation messages returned by each edge device, the server comprehensively considers the evaluation factors of each edge device to select the edge devices that participate in this round of federated learning. The evaluation factors of each edge device include: the potential importance of edge device data, the computing power and communication conditions of edge devices, the historical participation and contribution of edge devices, and system resource constraints.

[0027] Preferably, the selected edge devices participating in federated learning perform local model training, and use a key data identification algorithm to screen out the key data samples that have the greatest impact on model performance. The edge device uploads the updated local model and key data samples to the server, including:

[0028] The selected edge devices participating in federated learning use the environment data stored locally and perform local model training for federated learning according to the configuration parameters provided by the server. During the training process, the local model adjusts its own parameters to fit the characteristics of local data.

[0029] The edge device uses a key data identification algorithm to screen out the key data samples that have the greatest impact on model performance based on key data evaluation metrics, performs privacy protection processing on the selected key data samples, and the edge device uploads the updated local model and key data samples to the server. The key data evaluation metrics include: prediction uncertainty, gradient magnitude, difference degree from the current data distribution, and minority class samples.

[0030] Preferably, the key data refers to the data samples that have the greatest impact on model performance, including data representing rare scenarios, abnormal events, or situations where the current model performs poorly. The evaluation metrics of the key data samples include: prediction uncertainty, gradient magnitude, difference degree from the current data distribution, and minority class samples;

[0031] In the autonomous driving scenario, the key data includes rare road conditions, abnormal driving scenarios, and special traffic event data.

[0032] Preferably, the server aggregates the local models uploaded by each edge device, generates a new global model through a federated aggregation algorithm, and distributes the new global model to each edge device, including:

[0033] The server receives the local models uploaded by all edge devices participating in federated learning and executes the federated aggregation algorithm to generate a new global model;

[0034] The server processes and establishes a key database based on the key data uploaded by all edge devices participating in federated learning, stores and manages the key data uploaded by each edge device in the key database, analyzes the data characteristics in the key database, extracts valuable patterns, and fine-tunes the aggregated global model according to the valuable patterns;

[0035] The server evaluates the performance of the updated global model. If the performance improvement is not obvious, it can adjust the learning strategy for the next round and distribute the new global model to all edge devices.

[0036] As can be seen from the technical solutions provided by the embodiments of the present invention above, the present invention proposes a method for trigger-based federated continuous learning combined with key data uploading. This method triggers the federated learning process only when the performance degradation reaches a threshold by real-time monitoring of the model performance, avoiding unnecessary communication and computing consumption. At the same time, the method introduces a device scheduling mechanism and a key data identification mechanism that are aware of key data. It not only preferentially selects devices with high-value data to participate in training but also selectively uploads a small amount of high-value data in key scenarios, significantly improving the model's adaptability to new environments while protecting privacy.

[0037] Additional aspects and advantages of the present invention will be given in part in the following description, which will become apparent from the following description or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 A schematic diagram of a federated continuous learning architecture provided for an embodiment of the present invention;

[0040] Figure 2 A schematic diagram of the implementation principle of a method for scheduling and key data sharing of trigger-based federated continuous learning provided for an embodiment of the present invention;

[0041] Figure 3The processing flow chart of a method for scheduling and key data sharing of trigger-based federated continuous learning provided by an embodiment of the present invention;

[0042] Figure 4 Schematic diagram of five key information exchange steps provided by an embodiment of the present invention. Detailed implementation manners

[0043] The following details the implementation manners of the present invention. Examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The implementation manners described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0044] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any and all combinations of one or more of the associated listed items.

[0045] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless defined as here.

[0046] For ease of understanding of the embodiments of the present invention, the following will further explain with several specific embodiments as examples in conjunction with the accompanying drawings, and each embodiment does not constitute a limitation to the embodiments of the present invention.

[0047] In the embodiments of the present invention, federated learning, as a distributed learning method, allows different devices and computing nodes to collaboratively learn a model without directly exchanging raw data. In the architecture of federated learning, due to the uneven computing capabilities between devices, the uneven network connection speeds, and the imbalance in data distribution, and also due to the dynamic changes in the environment. Taking the autonomous driving scenario as an example, vehicles, as edge devices, continuously collect environmental data such as vehicle speed, road conditions (whether congested, road type, etc.), weather conditions (sunny, rainy, snowy, etc.), and dynamic information of surrounding vehicles and pedestrians. The device data is updated in real time, and these heterogeneous characteristics pose potential challenges to the training efficiency of the global model. Therefore, in the face of a dynamic environment, how to efficiently trigger model updates, intelligently schedule devices, and selectively share key data has become the core challenge for improving the system efficiency and adaptability.

[0048] A schematic diagram of a federated continuous learning architecture provided by the embodiments of the present invention is as Figure 1 shown. This architecture includes a central server and multiple edge devices. The central server is responsible for model aggregation and training scheduling, and the edge devices are responsible for local data collection, performance monitoring, and model training. The edge devices communicate with the server through the network, sharing model parameters and selectively uploading key data. Based on real-time monitoring of the changes in model performance, the system only starts the federated learning process when the performance degradation reaches a threshold. At the same time, it innovatively integrates data importance evaluation into the device scheduling decision, preferentially selecting devices holding key data to participate in the training. The selected devices not only upload model updates but also identify and screen a small number of high-value data samples processed through privacy protection and upload them to the server for model fine-tuning, thereby significantly improving the system's adaptability to environmental changes, especially rare scenarios, while protecting data privacy.

[0049] A schematic diagram of the implementation principle of a method for scheduling and key data sharing in trigger-based federated continuous learning provided by the embodiments of the present invention is as Figure 2 shown. The present invention relies on the computing capabilities of edge devices and network technologies (such as 5G, Wi-Fi, etc.). Among them, the edge devices evaluate the performance of the local model on the latest data at a certain frequency (such as every hour, every day, etc.) and generate a report containing performance metrics and data distribution characteristics and upload it to the server. The server is responsible for summarizing and analyzing the performance evaluation reports of each device and deciding whether to trigger the federated learning process. Then the server performs intelligent scheduling based on data importance and device capabilities; selects edge devices to perform local training and identify key data; the server receives the local model updates and key data, performs aggregation and fine-tuning; the updated global model is distributed to all edge devices; the system continuously monitors the performance changes to form a closed-loop optimization.

[0050] The system of the present invention is a federated continuous learning system.

[0051] The present invention can be used in scenarios with dynamically changing environments, such as the autonomous driving scenario. Vehicles, as edge devices, collect environmental data such as vehicle speed, road conditions (whether congested, road type, etc.), weather conditions (sunny, rainy, snowy, etc.), and dynamic information of surrounding vehicles and pedestrians. In the industrial Internet, various production devices and sensors, as edge devices, continuously collect data such as production line status, equipment operation parameters, and environmental temperature and humidity, so as to detect production anomalies in a timely manner, predict equipment failures, and at the same time protect sensitive production data. In remote healthcare, health monitoring devices at patients' homes, as edge devices, collect vital signs and health status data. When a change in the patient's condition or a new health pattern is detected, the model is triggered for update. Among them, the key data may include early symptoms of diseases, changes in signs after treatment plan changes, etc. These data can be selectively uploaded after strict privacy protection processing to help improve the global health monitoring model.

[0052] After training is completed, the edge device and the server can perform the following subsequent operations: The edge device continues to collect environmental data and monitor the model performance; the server continues to analyze the performance reports from the edge device. If the performance improvement is not obvious, the server can dynamically adjust the training strategy; the global model is distributed to all edge devices, including those that did not participate in the current round of training; the server maintains a key database for storing and managing the key data uploaded from the edge device.

[0053] The specific working process of a method for scheduling and key data sharing of trigger-based federated continuous learning provided by an embodiment of the present invention is as Figure 3 shown, including the following processing steps:

[0054] Step S10: The edge device continuously collects data and monitors performance, and uploads a performance evaluation report to the server.

[0055] The edge device collects environmental data at a certain frequency (such as every hour) and stores the environmental data locally. The environmental data includes various parameters collected by the edge device about its surrounding environment. For example, in the autonomous driving scenario, this will include vehicle speed, road conditions (congestion level, road type), weather conditions (sunny, rainy, snowy, etc.), and dynamic information of surrounding vehicles and pedestrians.

[0056] Edge devices regularly use the current model, that is, they use the model sent by the server before to process new data and calculate performance metrics (accuracy, F1-score, etc.). The F1-score is a comprehensive evaluation metric in classification problems. It is the weighted average of precision and recall, used to consider both the number of predicted positive examples and the actual positive examples. Among them, F1 is also called the harmonic mean of precision and recall. The calculation formula for the F1-score is: F1-score = 2 * (precision * recall) / (precision + recall). The value range of the F1-score is from 0 to 1. The closer the value is to 1, the higher the prediction accuracy of the model, and the closer the value is to 0, the lower the prediction accuracy of the model.

[0057] The calculation of these performance metrics depends on the quality and integrity of environmental data. For example, in the object detection model of autonomous driving, accurate road conditions and surrounding object information help the model identify targets more accurately, thereby improving the accuracy.

[0058] Based on the processing of environmental data, edge devices will generate a performance evaluation report, which includes performance metric values and their change trends, detected data distribution changes, statistical characteristics of local data, the ID of the edge device, and timestamps, etc. For example, in an intelligent transportation system, if the road types that vehicles travel on change frequently within a period of time, the change in data distribution will be reflected in the performance evaluation report, which is crucial for judging whether the model adapts to the new environment.

[0059] Edge devices upload the performance evaluation report to the server. By continuously collecting environmental data, edge devices ensure that the model can continuously adapt to environmental changes. Real-time calculation of performance metrics and generation of reports enable the server to timely understand the running status of the model on each edge device, provide an accurate basis for subsequent triggering decisions, help detect situations where the model performance deteriorates in a timely manner, and make preparations for optimization in advance.

[0060] Step S20: The server analyzes the performance evaluation report and determines whether to trigger the federated learning process based on the performance evaluation results of the edge devices.

[0061] The server receives and stores the performance evaluation reports submitted by each edge device. The performance evaluation report includes performance metric values and their change trends, detected data distribution changes, statistical characteristics of local data, the ID of the edge device, and timestamps, etc.

[0062] The server analyzes the performance evaluation report, obtains the performance change trends of all edge devices, and calculates the overall performance metrics of the system and their change rates. Aggregation of the performance metrics of each edge device can be performed using methods such as weighted average, and the weights can consider device importance and data volume.

[0063] Based on the performance evaluation results of the system, the server determines whether to trigger the federated learning process. If the overall system performance drops to a preset threshold (e.g., the accuracy drops by 5%), the federated learning is triggered; if the performance is stable and higher than the target requirement, the learning process is not triggered and continuous monitoring is carried out.

[0064] The above preset threshold can be dynamically adjusted based on the importance of the application scenario and resource constraints, avoiding unnecessary federated learning processes, reducing the communication and computing overhead of the system, enabling the system to flexibly control the frequency of model updates according to different application scenarios and resource conditions, improving the adaptive ability of the system, and ensuring that the model can maintain good performance in various complex environments.

[0065] Step S30: After the federated learning process is triggered, the server selects edge devices holding key data to participate in the federated learning.

[0066] Key data refers to the data samples that have the greatest impact on the model performance. This includes data representing rare scenarios, abnormal events, or situations where the current model performs poorly. For example, in the autonomous driving scenario, this key data includes rare road conditions, abnormal driving scenarios, and special traffic events, such as mountain winding roads, driving behaviors in extreme weather, and complex traffic signs. By selecting edge devices with such high-value, low-frequency but high-information data, the server can significantly improve the model's adaptability to complex environments and accelerate the model's learning of the decision-making logic for key scenarios.

[0067] The server sends a request to participate in the federated learning to the selected edge devices. This request includes training configurations (federated learning strategy, learning rate, number of local training rounds, etc.) and an indication of whether key data needs to be uploaded.

[0068] After receiving the request sent by the server, the edge device evaluates its own computing resources and data status and decides whether to participate in this round of training. The edge device that accepts the participation sends a confirmation message to the server, and this confirmation message includes device status and available resource information.

[0069] Based on the confirmation messages returned by each edge device, the server comprehensively considers the evaluation factors of each edge device to select the edge devices participating in this round of federated learning. The evaluation factors of each edge device include: the potential importance of the edge device data (analyzed based on the performance evaluation report); the computing power and communication conditions of the edge device; the historical participation situation and contribution degree of the edge device; system resource constraints (such as the maximum number of parallel training devices).

[0070] The server preferentially selects edge devices holding key data, enabling the model to obtain richer and more representative information during the training process, accelerating model convergence, enhancing the model's adaptability to new environments and rare scenarios, reducing the blindness of model training, and improving the utilization rate of training resources.

[0071] Step S40: The edge devices selected to participate in federated learning perform local model training. The key data samples that have the greatest impact on the model performance are screened out through the key data recognition algorithm. The edge devices upload the updated local model and the key data samples to the server.

[0072] The edge devices selected to participate in federated learning utilize the environment data stored locally and perform model training according to the configuration parameters provided by the server (such as learning rate, batch size, number of training rounds, etc.). During the training process, the model adjusts its own parameters to fit the characteristics of the local data.

[0073] During or after the training process, the edge devices screen out the key data samples that have the greatest impact on the model performance through the key data recognition algorithm. The evaluation of the key data samples can be based on the following metrics:

[0074] Prediction uncertainty (low prediction confidence of the model for the sample)

[0075] Gradient magnitude (the magnitude of the gradient generated by the sample)

[0076] Degree of difference from the current data distribution (the sample exhibits data drift)

[0077] Minority class samples (rare classes in the case of data imbalance)

[0078] Perform privacy protection processing on the selected key data samples. The following techniques can be adopted:

[0079] Differential privacy (adding controlled noise)

[0080] Feature perturbation (modifying non-critical features)

[0081] Data abstraction (reducing the detail accuracy)

[0082] Removing sensitive information

[0083] The device uploads to the server simultaneously:

[0084] Model update (parameter difference)

[0085] The key data samples processed by privacy protection (usually a very small part of the local data, such as less than 5%).

[0086] The edge device uploads the updated local model and key data samples to the server. Local training enables the model to be optimized for local environmental data, improving the model's adaptability locally. The key data identification algorithm is used to screen key data, which provides more valuable training information for the model and helps improve the model's generalization ability. At the same time, privacy protection is used to achieve the sharing of key data, avoiding the risks brought by data leakage and providing effective data support for model fine-tuning on the server side.

[0087] Step S50: The server aggregates the local models uploaded by each edge device, generates a new global model through the federated aggregation algorithm, and distributes the new global model to each edge device.

[0088] The server receives the local models uploaded by all edge devices participating in federated learning, executes the federated aggregation algorithm (such as FedAvg), and generates a new global model.

[0089] The server processes and establishes a key database based on the key data uploaded by all edge devices participating in federated learning, and stores and manages the key data uploaded by each edge device in the key database.

[0090] The server analyzes the data characteristics in the key database, extracts valuable patterns, and fine-tunes the aggregated global model according to the valuable patterns to enhance the adaptability to special scenarios.

[0091] The server evaluates the performance of the updated global model. If the performance improvement is not obvious, the learning strategy for the next round can be adjusted. Finally, the new global model is distributed to all edge devices, including those that did not participate in this round of training.

[0092] The server fuses the model updates of multiple edge devices through federated aggregation, integrates the data information on different devices, and improves the generalization ability of the global model. At the same time, the processing of key data and model fine-tuning enable the global model to better handle special scenarios, enhancing the practicality and robustness of the model.

[0093] The five key information exchange steps of the present invention are as Figure 4 shown:

[0094] The edge device uploads a performance evaluation report to the server.

[0095] The server analyzes the performance and sends a trigger decision and participation request to the device.

[0096] The edge device confirms whether to participate in this round of training.

[0097] The server selects existing devices and sends training configurations.

[0098] The edge device uploads local model updates and key data.

[0099] The server distributes the updated global model.

[0100] In summary, the method of the embodiment of the present invention reduces unnecessary communication and computing resource consumption based on the trigger mechanism of performance monitoring, and improves the overall efficiency of the system; the device scheduling strategy of key data perception accelerates the model's adaptation to the new environment by preferentially selecting devices with high-value data to participate in training; the key data selective upload mechanism protects data privacy while enhancing the model's processing ability for rare scenarios. At the same time, the efficiency and adaptability of federated continuous learning are comprehensively optimized, providing strong support for the application of distributed intelligent systems in dynamic environments, and having broad application prospects in fields such as autonomous driving and intelligent healthcare.

[0101] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.

[0102] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0103] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and the relevant parts can refer to the partial description of the method embodiments. The device and system embodiments described above are only illustrative, and the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0104] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for scheduling and key data sharing of trigger-based federated continuous learning, characterized in that Including: The edge device continuously collects environmental data and monitors performance, and uploads a performance evaluation report to the server; The server analyzes the performance evaluation report and determines whether to trigger the federated learning process based on the performance evaluation report of the edge device; When the federated learning process is triggered, the server selects edge devices holding key data to participate in federated learning; The edge devices selected to participate in federated learning perform local model training, screen out the key data samples that have the greatest impact on the model performance through the key data identification algorithm, and the edge devices upload the updated local model and key data samples to the server; The server aggregates the local models uploaded by each edge device, generates a new global model through the federated aggregation algorithm, and distributes the new global model to each edge device.

2. The method according to claim 1, characterized in that, The edge device continuously collects environmental data and monitors performance, and uploads a performance evaluation report to the server, including: The edge device collects environmental data at a certain frequency, processes the environmental data using the local model, calculates performance metrics, generates a performance evaluation report based on all performance metrics. The performance evaluation report contains performance metric values and their change trends, detected data distribution changes, statistical characteristics of local data, the ID of the edge device, and timestamps. The edge device uploads the performance evaluation report to the server through the network.

3. The method according to claim 2, wherein The environmental data includes various parameters collected by the edge device about its surrounding environment. In the autonomous driving scenario, the environmental data includes vehicle speed, road conditions, weather conditions, and dynamic information of surrounding vehicles and pedestrians. In the telemedicine scenario, the environmental data includes the vital signs and health condition data of patients.

4. The method according to claim 3, characterized in that, The server analyzes the performance evaluation report and determines whether to trigger the federated learning process based on the performance evaluation report of the edge device, including: The server receives and stores the performance evaluation reports submitted by each edge device, analyzes the performance evaluation reports, obtains the performance change trends of each edge device, and obtains the overall performance metrics of the system and their change rates based on the performance change trends of each edge device; The server compares the overall performance metrics of the system with a preset threshold, and determines whether to trigger the federated learning process according to the comparison result. If the overall performance of the system drops to the preset threshold, the federated learning is triggered; if the overall performance of the system is higher than the preset threshold, and if the performance is stable and higher than the target requirement, the federated learning process is not triggered, and the monitoring continues. The preset threshold is dynamically adjusted based on the importance of the application scenario and resource constraints.

5. The method according to claim 4, wherein When the federated learning process is triggered, the server selects edge devices holding key data to participate in federated learning, including: When the federated learning process is triggered, the server selects edge devices holding key data to participate in federated learning, and the server sends a request to participate in federated learning to the selected edge devices. The request includes training configuration and an indication of whether key data needs to be uploaded. The training configuration includes federated learning strategy, learning rate, and local training rounds; After the edge device receives the request sent by the server, it evaluates its own computing resources and data status, decides whether to participate in this round of federated learning, and the edge devices that accept to participate in this round of federated learning send a confirmation message to the server, and the confirmation message includes device status and available resource information; Based on the confirmation messages returned by each edge device, the server selects the edge devices participating in this round of federated learning by comprehensively considering the evaluation factors of each edge device. The evaluation factors of each edge device include: the potential importance of edge device data, the computing power and communication conditions of edge devices, the historical participation and contribution of edge devices, and system resource constraints.

6. The method according to claim 5, wherein The selected edge devices participating in federated learning perform local model training, and screen out the key data samples that have the greatest impact on model performance through a key data identification algorithm. The edge devices upload the updated local model and key data samples to the server, including: The selected edge devices participating in federated learning utilize the environmental data stored locally and perform local model training for federated learning according to the configuration parameters provided by the server. During the training process, the local model adjusts its own parameters to fit the characteristics of local data; The edge device screens out the key data samples that have the greatest impact on model performance based on the key data evaluation index through the key data identification algorithm, performs privacy protection processing on the selected key data samples, and the edge device uploads the updated local model and key data samples to the server. The key data evaluation index includes: prediction uncertainty, gradient magnitude, difference degree from the current data distribution, and minority class samples.

7. The method according to claim 6, wherein The key data refers to the data samples that have the greatest impact on model performance, including the data representing rare scenarios, abnormal events, or the situations where the current model performs poorly. The evaluation indexes of the key data samples include: prediction uncertainty, gradient magnitude, difference degree from the current data distribution, and minority class samples; In the autonomous driving scenario, the key data includes rare road conditions, abnormal driving scenarios, and special traffic event data.

8. The method according to claim 5, characterized in that The server aggregates the local models uploaded by each edge device, generates a new global model through the federated aggregation algorithm, and distributes the new global model to each edge device, including: The server receives the local models uploaded by all edge devices participating in federated learning and executes the federated aggregation algorithm to generate a new global model; The server processes and establishes a key database based on the key data uploaded by all edge devices participating in federated learning, stores and manages the key data uploaded by each edge device in the key database, analyzes the data characteristics in the key database, extracts valuable patterns, and fine-tunes the aggregated global model according to the valuable patterns; The server evaluates the performance of the updated global model. If the performance improvement is not obvious, it can adjust the learning strategy for the next round and distribute the new global model to all edge devices.

Citation Information

Cited By

  • Fault diagnosis method, system and device, electronic equipment and storage medium

    CN120448178A

  • Water body carbon flux robot observation system for algae community analysis

    CN121766615A