Personalized federal edge learning method and device, storage medium and computer equipment
Through personalized federated edge learning methods, the video data model is collected and optimized by monitoring terminals and client devices, combined with the model aggregation strategy of edge and cloud servers, the problem of poor adaptability of abnormal population behavior detection methods in the existing technology is solved, and more efficient and accurate detection effects are achieved.
Patent Information
- Application Number
- CN202510243484.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The existing behavioral detection methods for abnormal populations are difficult to meet the personalized detection needs in different monitoring scenarios, and are poor in adaptability, resulting in limited detection accuracy and unsatisfactory detection results.
A personalized federated edge learning method is proposed. By collecting video data by monitoring terminals, the client device optimizes the initial model based on the video data, and generates a local model; the edge server uses an asynchronous cross-device model aggregation strategy to aggregate the local model to generate an edge model; the cloud server regularly uses an asynchronous model aggregation strategy to globally aggregate the edge model, generates a global model, and distributes it back to each device until the iteration end condition is met.
This method can improve the accuracy and adaptability of detection, reduce the aggregation delay caused by device problems, enhance the scalability and robustness of the system, reduce dependence on the central server, and improve the real-time and efficiency of the monitoring system.
Smart Images

Figure CN120014364A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of federated learning technology, and in particular to a personalized federated edge learning method, apparatus, storage medium, and computer equipment. Background Art
[0002] In the field of public safety, abnormal crowd behavior detection (ACBD) plays a vital role. Its goal is to identify abnormal behaviors such as suspicious activities or potential security threats from surveillance videos. Based on this, with the acceleration of urbanization and the growth of public safety needs, ACBD technology plays an increasingly important role in ensuring social stability and preventing crime.
[0003] However, in actual monitoring scenarios, due to the diversity of the crowd and the complexity of the scene, the traditional ACBD method faces the challenges of data heterogeneity and personalized needs. These methods not only require a large amount of labeled data for training, but also lack the design of personalized learning models for specific crowd scenes, which makes it difficult to meet the personalized detection needs in different monitoring scenarios, and has poor adaptability, which in turn limits the detection accuracy and unsatisfactory detection effect. Summary of the invention
[0004] The purpose of this application is to solve at least one of the above-mentioned technical defects, especially the technical defect that the abnormal crowd behavior detection method in the prior art is difficult to meet the personalized detection needs in different monitoring scenarios and has poor adaptability, which in turn limits the detection accuracy and leads to unsatisfactory detection effect.
[0005] The present application provides a personalized federated edge learning method, the method comprising:
[0006] Collecting video data through a monitoring terminal, and optimizing a preset initial model based on the video data using a client device to generate a local model;
[0007] The edge server is used to obtain the local models output by multiple client devices based on the aggregation threshold, and the asynchronous cross-device model aggregation strategy is used to aggregate the local models to generate the edge model.
[0008] The cloud server is used to periodically adopt an asynchronous model aggregation strategy to globally aggregate the edge models output by multiple edge servers to generate a global model.
[0009] The global model is distributed back to each edge server and each client device as the latest initial model, and the video data is collected through the monitoring terminal and its subsequent steps are returned until the global model meets the preset iteration end condition.
[0010] Optionally, the using the client device to perform model training and optimization on a preset initial model based on the video data to generate a local model includes:
[0011] Using the client device to perform network decomposition on the preset initial model, a general feature extraction sub-network and a personalized label mapping sub-network are obtained;
[0012] Performing abnormal feature detection on the video data through the general feature extraction sub-network, and extracting multiple behavior features according to the detection results;
[0013] Perform label mapping on each behavior feature through the personalized label mapping subnetwork to obtain a behavior label corresponding to each behavior feature;
[0014] The initial model is optimized based on each behavior feature and a behavior label corresponding to each behavior feature to obtain a local model.
[0015] Optionally, the optimizing the initial model based on each behavior feature and the behavior label corresponding to each behavior feature to obtain a local model includes:
[0016] Inputting each behavior feature and the behavior label corresponding to each behavior feature into the initial model to obtain a predicted label output by the initial model based on each behavior label;
[0017] Back-propagating each predicted label using the loss function to obtain the gradient result of the initial model;
[0018] A stochastic gradient descent method and a hyperparameter control method are adopted to update the network parameters of the general feature extraction subnetwork and the personalized label mapping subnetwork respectively based on the gradient results.
[0019] Optionally, the calculation formula for the back propagation includes:
[0020]
[0021] In the formula, Represents the use of loss function The calculated gradient result; among them, represents the gradient, represents the general feature extraction subnetwork in the i-th client device Network parameters in the t-1th update iteration; represents the personalized label mapping subnetwork in the i-th client device Network parameters in the t-1th update iteration; Represents the loss function right and The partial derivative of .
[0022] Optionally, the calculation formula of the stochastic gradient descent method includes:
[0023]
[0024]
[0025] In the formula, represents the general feature extraction subnetwork in the i-th client device Network parameters in the tth update iteration; represents the learning rate; Represents the general feature extraction subnetwork The gradient of Represents the loss function The calculated gradient result; represents the personalized label mapping subnetwork in the i-th client device Network parameters in the tth update iteration; Represents the general feature extraction subnetwork gradient.
[0026] Optionally, the calculation formula of the hyperparameter control method includes:
[0027]
[0028] In the formula, represents the norm of the parameter vector; Represents the regularization term, which is used to control the personalized label mapping subnetwork The intensity of the personal adjustment.
[0029] Optionally, the calculation formula of the asynchronous cross-device model aggregation strategy includes:
[0030]
[0031]
[0032] In the formula, and Represents the general feature extraction subnetwork described in the edge model and the personalized label mapping subnetwork Network parameters in the t+1th update iteration; and represents the general feature extraction subnetwork in the i-th client device and the personalized label mapping subnetwork Network parameters in the tth update iteration; j represents the jth client device; Represents the set of client devices participating in the current aggregation, Indicates the number of parameters.
[0033] Optionally, the calculation formula of the asynchronous model aggregation strategy includes:
[0034]
[0035]
[0036] In the formula, and Respectively represent the general feature extraction subnetworks described in the global model and the personalized label mapping subnetwork Network parameters; represents the set of all edge servers participating in the aggregation, and i and j represent the i-th and j-th edge servers in the set of all edge servers participating in the aggregation, respectively.
[0037] The present application also provides a personalized federated edge learning device, including:
[0038] A model optimization module, used to collect video data through a monitoring terminal, and optimize a preset initial model based on the video data using a client device to generate a local model;
[0039] A primary aggregation module is used to obtain local models output by multiple client devices using an edge server based on an aggregation threshold, and to aggregate each local model using an asynchronous cross-device model aggregation strategy to generate an edge model;
[0040] A secondary aggregation module is used to globally aggregate edge models output by multiple edge servers using a cloud server periodically using an asynchronous model aggregation strategy to generate a global model;
[0041] The model iteration module is used to distribute the global model as the latest initial model back to each edge server and each client device, and return to the video data collected by the monitoring terminal and its subsequent steps until the global model meets the preset iteration end condition.
[0042] The present application also provides a storage medium, which stores computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the personalized federated edge learning method as described in any of the above embodiments.
[0043] The present application also provides a computer device, comprising: one or more processors, and a memory;
[0044] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the personalized federated edge learning method described in any one of the above embodiments are performed.
[0045] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0046] The personalized federated edge learning method, device, storage medium and computer equipment provided by the present application can collect video data through the monitoring terminal when detecting abnormal crowd behavior, and use the client device to optimize the preset initial model based on the video data to generate a local model, so that the model can adapt to different monitoring scenarios and improve the accuracy and adaptability of detection. Then, the edge server can be used to obtain the local models output by multiple client devices based on the aggregation threshold, and the asynchronous cross-device model aggregation strategy can be used to aggregate the local models to generate an edge model, so as to reduce the aggregation delay caused by equipment problems and improve the efficiency and flexibility of edge model aggregation. Then, the cloud server can be used to periodically use the asynchronous model aggregation strategy to globally aggregate the edge models output by multiple edge servers to generate a global model, so that the model can learn a wider range of crowd behavior patterns and characteristics, and achieve global performance improvement; finally, the global model can be distributed back to each edge server and each client device as the latest initial model, and the video data collected through the monitoring terminal and its subsequent steps are returned until the global model meets the preset iteration end condition, so as to ensure that the entire system can continue to learn and adapt to new data and changes while maintaining the real-time and efficiency of the model. In addition, the method of the present application can also reduce the dependence on the central server, reduce the overhead of data transmission, and thus enhance the scalability and robustness of the entire monitoring system. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0048] Figure 1 A schematic diagram of the architecture of a personalized federated edge learning framework provided in an embodiment of the present application;
[0049] Figure 2 A flowchart of a personalized federated edge learning method provided in an embodiment of the present application;
[0050] Figure 3A strategy flow chart of an initial model optimization process provided in an embodiment of the present application;
[0051] Figure 4 An interactive schematic diagram of a personalized federated edge learning method provided in an embodiment of the present application;
[0052] Figure 5 A strategy flow chart of a two-stage model aggregation process provided in an embodiment of the present application;
[0053] Figure 6 A schematic diagram of the structure of a personalized federated edge learning device provided in an embodiment of the present application;
[0054] Figure 7 A schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0056] In actual monitoring scenarios, due to the diversity of the crowd and the complexity of the scene, traditional ACBD methods face the challenges of data heterogeneity and personalized needs. These methods not only require a large amount of labeled data for training, but also lack the design of personalized learning models for specific crowd scenes, which makes it difficult to meet the personalized detection needs in different monitoring scenarios, and has poor adaptability, which in turn limits the detection accuracy and unsatisfactory detection effect.
[0057] Based on this, this application proposes the following technical solutions, see below for details:
[0058] In some embodiments, Figure 1 As shown, Figure 1 A schematic diagram of the architecture of a personalized federated edge learning framework provided in an embodiment of the present application; Figure 1 It can be seen that the personalized federated edge learning method of the present application can be applied to different monitoring scenarios, and the personalized federated edge learning framework it adopts is mainly composed of a cloud server, an edge server, a client device and a monitoring terminal.
[0059] Among them, monitoring terminals are deployed in public places and are responsible for collecting key video surveillance data; client devices are connected to monitoring terminals, receive video stream data, and perform preliminary model training and reasoning locally. The edge server, as the middle layer, is responsible for collecting model updates from client devices and performing model aggregation operations. The cloud server, as the central hub of the system, is responsible for creating and updating global models and coordinating the operation of the entire system.
[0060] In one embodiment, Figure 2 As shown, Figure 2 A flowchart of a personalized federated edge learning method provided in an embodiment of the present application; the present application provides a personalized federated edge learning method, which specifically includes the following:
[0061] S110: Collect video data through the monitoring terminal, and use the client device to optimize the preset initial model based on the video data to generate a local model.
[0062] In this step, when it is necessary to perform behavior detection on abnormal people in different monitoring scenarios, the present application can collect video data of the corresponding scene through the monitoring terminal in the personalized federated edge learning framework, and then use the client device to optimize the preset initial model based on the video data to generate a local model. Therefore, the model can adapt to different monitoring scenarios, thereby improving the accuracy and adaptability of detection.
[0063] Among them, the initial model refers to a general behavior detection model pre-stored in the client device. This model is trained based on a large-scale data set and has basic behavior recognition capabilities, but the recognition accuracy for specific monitoring scenarios is not high enough. Therefore, the client device needs to further optimize it for specific monitoring scenarios.
[0064] Specifically, monitoring terminals deployed in public places can collect video data of corresponding scenes in real time, ensuring that the system can comprehensively and accurately obtain behavioral data in various monitoring scenarios. After data collection is completed, the monitoring terminal can transmit the video data to the client device connected to it. The client device undertakes the core task of local intelligent analysis and can pre-process, extract features and train models for the received video data based on the preset initial model to obtain a local model.
[0065] It is understandable that by combining local environmental characteristics and historical behavior data, the client device can personalize and optimize the initial model to generate a local model that adapts to the current monitoring scenario. Compared with the traditional centralized training method, this method can better capture the behavior patterns in different scenarios and improve the accuracy of abnormal crowd behavior detection.
[0066] S120: Using the edge server to obtain local models output by multiple client devices based on an aggregation threshold, and using an asynchronous cross-device model aggregation strategy to aggregate each local model to generate an edge model.
[0067] In this step, after the local model is obtained through the client device optimization in step S110, the edge server can obtain the local models output by multiple client devices connected to it based on the aggregation threshold, and adopt an asynchronous cross-device model aggregation strategy to aggregate each local model to generate an edge model, so as to reduce the aggregation delay caused by device problems and improve the efficiency and flexibility of edge model aggregation.
[0068] The aggregation threshold refers to the condition or standard set by the edge server before executing local model aggregation to determine when to trigger the model aggregation process, thereby ensuring that the model aggregation is completed within a reasonable time frame. This can not only ensure the effectiveness of model aggregation, but also avoid system response delays caused by too long waiting time. It should be noted that the aggregation threshold of this application can be the submission threshold of the local model or the waiting time threshold, which is not limited here.
[0069] It is understandable that after completing the collection of local models, the edge server can adopt an asynchronous cross-device model aggregation strategy to efficiently fuse the local models submitted by each client device. The asynchronous cross-device model aggregation strategy here is a model aggregation method used in a federated learning edge computing environment. It allows the edge server to receive and process model updates of client devices at different time points without waiting for all client devices to submit synchronously, thereby effectively reducing aggregation delays caused by factors such as inconsistent device status, network fluctuations, or limited computing resources. Therefore, the edge server of this application not only improves the efficiency of edge model aggregation, but also enhances the flexibility of the system, enabling it to adapt to dynamically changing computing environments and communication conditions.
[0070] Furthermore, in the process of model aggregation, the edge server can use weighted average, personalized aggregation or meta-learning-based methods to perform more intelligent aggregation optimization based on the training contribution, data distribution characteristics and model performance feedback of each client device. Ultimately, the aggregated edge model will be more generalizable while retaining the personalized characteristics of each client device, enabling it to more accurately adapt to the needs of abnormal behavior detection in different monitoring scenarios, laying a solid foundation for further cloud optimization and global model updates.
[0071] S130: Using the cloud server to periodically adopt an asynchronous model aggregation strategy to globally aggregate edge models output by multiple edge servers to generate a global model.
[0072] In this step, after the edge server aggregates the edge model in step S120, the cloud server can periodically use an asynchronous model aggregation strategy to globally aggregate the edge models output by multiple edge servers connected to it to generate a global model, so that the model can learn a wider range of crowd behavior patterns and characteristics, thereby achieving improved global performance.
[0073] Specifically, after completing local model aggregation, the edge server can generate an edge model and upload it to the cloud server. The cloud server, as the central coordination node of the entire system framework, can regularly execute asynchronous model aggregation strategies and fuse models from multiple edge servers to continuously optimize the global model.
[0074] It is understandable that since the monitoring areas covered by different edge servers may be different, the data distribution, feature patterns and crowd behavior types collected will also be different. Therefore, through cloud aggregation of cloud servers, the feature information of each region can be fully integrated, so that the model can learn more extensive and diverse crowd behavior patterns, and improve the adaptability and generalization ability of the model in various monitoring scenarios.
[0075] In addition, the asynchronous model aggregation strategy adopted in this application can avoid the delay problem caused by the cloud server waiting for all edge servers to submit the model synchronously, so that the system can update the global model more efficiently. Even if some edge servers fail to submit the model in time due to network conditions, computing resources and other factors, it will not affect the continuous optimization of the global model. Therefore, the system can maintain a high detection accuracy and stability in different monitoring environments, thereby improving the performance and intelligence level of overall abnormal behavior detection.
[0076] S140: Distribute the global model as the latest initial model back to each edge server and each client device, and return to collecting video data through the monitoring terminal and subsequent steps until the global model meets the preset iteration end condition.
[0077] In this step, after the cloud server aggregates the global model through step S130, the cloud server can also distribute the global model back to each edge server and each client device as the latest initial model, so that each client device can continue to train and optimize the model to adapt to new data and changes, while maintaining the real-time and efficiency of the model. When the global model meets the preset iteration end conditions, the model iteration process can be ended.
[0078] It is understandable that through the mechanism of continuous iterative updates, this application can ensure the real-time performance of the model, so that the client device can always use the latest optimized model for reasoning and decision-making, and avoid the model from gradually becoming invalid due to changes in data distribution. At the same time, while executing model distribution, the cloud server will continue to monitor the training effect of the model, and decide whether to end the entire model iteration process based on preset iteration termination conditions, such as achieving stable detection accuracy or set training rounds, so as to ensure that while ensuring the quality of the model, the computing and communication efficiency is improved, and unnecessary resource consumption is avoided, so that the entire personalized federated edge learning framework has stronger adaptability and intelligent features while maintaining high efficiency.
[0079] It should be noted that in addition to distributing the global model back to each client device for iterative training, the cloud server can also distribute it to each edge server. To elaborate, in the personalized federated edge learning framework, the geographical location of the edge server is closer to the client device. Compared with the long-distance communication between the edge server and the cloud server, the communication delay between the edge server and the edge server is relatively small; and in the application scenario of abnormal crowd behavior detection, which has high real-time requirements, fast data transmission and model updates are crucial. Therefore, the present application can first send the local model update to the edge server for preliminary processing, which can significantly shorten the data transmission time, speed up the model iteration speed, enable the system to respond to new abnormal behavior patterns more promptly, improve the timeliness and practicality of the system, and enhance the monitoring and early warning capabilities of public safety incidents.
[0080] In addition, in a large-scale monitoring network, there may be a large number of client devices that perform task calculations and data transmission at the same time. If all data flows directly to the cloud server, it will put tremendous pressure on the network bandwidth and the computing resources of the cloud server, and easily lead to data congestion and processing delays. The edge server is located between the client device and the cloud server. As an intermediate layer, it can effectively divert and pre-process data and tasks. It can temporarily store and filter data and model updates from client devices, and process and forward them according to certain rules and priorities, thereby reducing the burden on the cloud server, optimizing the data transmission and processing process of the entire system, and ensuring the efficient and stable operation of the system.
[0081] In the above embodiment, when detecting abnormal crowd behavior, video data can be collected through the monitoring terminal, and the client device can be used to optimize the preset initial model based on the video data to generate a local model, so that the model can adapt to different monitoring scenarios and improve the accuracy and adaptability of detection. Then, the edge server can be used to obtain the local models output by multiple client devices based on the aggregation threshold, and the asynchronous cross-device model aggregation strategy can be used to aggregate each local model to generate an edge model, so as to reduce the aggregation delay caused by equipment problems and improve the efficiency and flexibility of edge model aggregation. Then, the cloud server can be used to periodically adopt an asynchronous model aggregation strategy to globally aggregate the edge models output by multiple edge servers to generate a global model, so that the model can learn a wider range of crowd behavior patterns and characteristics, and achieve global performance improvement; finally, the global model can be distributed back to each edge server and each client device as the latest initial model, and the video data collected through the monitoring terminal and its subsequent steps are returned until the global model meets the preset iteration end condition, so as to ensure that the entire system can continue to learn and adapt to new data and changes while maintaining the real-time and efficiency of the model. In addition, the method of the present application can also reduce the dependence on the central server, reduce the overhead of data transmission, and thus enhance the scalability and robustness of the entire monitoring system.
[0082] In one embodiment, the process of using the client device to perform model training and optimization on the preset initial model based on the video data to generate the local model in step S110 may include:
[0083] S111: Using the client device to perform network decomposition on the preset initial model to obtain a general feature extraction subnetwork and a personalized label mapping subnetwork.
[0084] S112: Perform abnormal feature detection on the video data through the general feature extraction sub-network, and extract multiple behavior features based on the detection results.
[0085] S113: Perform label mapping on each behavior feature through a personalized label mapping sub-network to obtain a behavior label corresponding to each behavior feature.
[0086] S114: Optimize the initial model based on each behavior feature and the behavior label corresponding to each behavior feature to obtain a local model.
[0087] In this embodiment, when optimizing the initial model, the client device can first perform network decomposition on the preset initial model to obtain a general feature extraction subnetwork and a personalized label mapping subnetwork, and then perform abnormal feature detection on the video data through the general feature extraction subnetwork, extract multiple behavior features based on the detection results, and perform label mapping on each behavior feature through the personalized label mapping subnetwork to obtain a behavior label corresponding to each behavior feature. Finally, the initial model can be optimized based on each behavior feature and the behavior label corresponding to each behavior feature to obtain a local model.
[0088] Specifically, the general feature extraction subnetwork and the personalized label mapping subnetwork each assume different functions and work together to achieve model optimization and personalized adaptation. Among them, the general feature extraction subnetwork is mainly responsible for extracting general features related to abnormal behavior detection from the input video data. These features are shared across multiple monitoring scenes and can capture common abnormal behavior patterns in different scenes. Through this subnetwork, the system can extract low-dimensional, representative features from high-dimensional video data. These features can reflect key information in the video, such as motion trajectory, posture changes, etc., and provide a basis for subsequent abnormal behavior detection. The personalized label mapping subnetwork is mainly responsible for mapping the features extracted by the general feature extraction subnetwork to specific behavior labels. This subnetwork is trained according to the characteristics of the local data set, can adapt to the needs of specific monitoring scenes, and realize personalized abnormal behavior recognition. Through this subnetwork, the system can associate the extracted general features with specific behavior labels, such as "running", "gathering", "fighting", etc., so as to achieve accurate recognition of abnormal behaviors in specific scenes.
[0089] In a specific embodiment, the general feature extraction subnetwork and the personalized label mapping subnetwork obtained after the initial model decomposition can be expressed as follows:
[0090]
[0091]
[0092] In the formula, and denote the common feature extraction subnetwork in the local model on the i-th client device. and personalized label mapping subnetwork Parameters; and Represent the general feature extraction subnetwork in the initial model and personalized label mapping subnetwork The global parameters of .
[0093] Among them, the general feature extraction sub-network It can be expressed as a mapping function that transforms high-dimensional input video data Mapping to a lower dimensional feature space , thereby extracting the common features in the monitoring scene. The mapping function here is specifically expressed as follows:
[0094]
[0095] Where d represents the dimension of the video data; t represents the dimension of the hidden features.
[0096] Personalized label mapping subnetwork It can also be expressed as a mapping function, which transforms the feature space output by the general feature extraction subnetwork into Mapping to behavior label space , so as to assign corresponding behavior labels to each feature. The mapping function here is specifically expressed as follows:
[0097]
[0098] In the formula, y represents the dimension of the label feature.
[0099] Based on this, the client device can combine all the extracted behavior features and their corresponding behavior labels to further optimize the initial model, so that the model can more accurately adapt to the characteristics of the current monitoring environment, and finally obtain an optimized local model. Through such an optimization process, the client device can not only improve the recognition accuracy of the model in the local scene, but also enhance the adaptability of the model in different monitoring environments, providing a better model foundation for subsequent edge model aggregation and global model optimization.
[0100] In one embodiment, the process of optimizing the initial model based on each behavior feature and the behavior label corresponding to each behavior feature to obtain the local model in step S114 may include:
[0101] S1141: Input each behavior feature and the behavior label corresponding to each behavior feature into the initial model to obtain a predicted label output by the initial model based on each behavior label.
[0102] S1142: Use the loss function to back-propagate each predicted label to obtain the gradient result of the initial model.
[0103] S1143: Using the stochastic gradient descent method and the hyperparameter control method, the network parameters of the general feature extraction subnetwork and the personalized label mapping subnetwork are updated based on the gradient results.
[0104] In this embodiment, the client device can input each behavior feature and the behavior label corresponding to each behavior feature into the initial model to obtain the predicted label output by the initial model based on each behavior label, and then use the loss function to backpropagate each predicted label to obtain the gradient result of the initial model, and use the stochastic gradient descent method and the hyperparameter control method to update the network parameters of the general feature extraction subnetwork and the personalized label mapping subnetwork based on the gradient results.
[0105] Among them, stochastic gradient descent is a gradient descent algorithm used to optimize neural network parameters. Compared with standard gradient descent (Batch Gradient Descent, BGD), it only uses one or a small batch of samples (mini-batch) to calculate gradients and update parameters in each iteration, instead of using the entire data set for calculation, thereby reducing the amount of calculation, increasing the training speed, and being able to jump out of the local optimal solution faster, improving the generalization ability of the model. The hyperparameter control method refers to the reasonable adjustment of key hyperparameters that affect the training effect during the model training process to optimize the training efficiency and improve the final performance of the model; the hyperparameters here can include learning rate, batch size, momentum, etc., which are not restricted here.
[0106] Specifically, the initial model can calculate each behavior label based on the input behavior features and output the corresponding prediction label, thereby providing the model with understanding and classification of the current data. After obtaining the prediction label, the client device can use the preset loss function to calculate the error between each predicted label and its corresponding true label to measure the accuracy of the model on the current training data. Subsequently, the error information can be passed to the parameters of each layer of the model through the back propagation algorithm, so that the client device can calculate the gradient information of the initial model and clarify the direction in which the model needs to be optimized. On this basis, the client device can use the stochastic gradient descent method, combined with the hyperparameter control strategy, to iteratively update the network parameters of the general feature extraction subnetwork and the personalized label mapping subnetwork to gradually optimize the model performance, so that the model can adapt to local monitoring data while improving its ability to understand complex behavior patterns, thereby ultimately obtaining a more accurate and efficient local model.
[0107] Indicatively, Figure 3 As shown, Figure 3 A strategy flow chart of an initial model optimization process provided in an embodiment of the present application; Figure 3In the system, the monitoring terminal is responsible for collecting data and transmitting the data to the data participants, i.e., the client devices, so that the client devices can divide the data into multiple batches (Batch 1, Batch 2, ..., Batch k) for processing, and then the data flows in the system through forward transmission (FP) and backward transmission (BP). Among them, forward transmission (FP) is used to transmit data from the monitoring terminal to the data processing node, and backward transmission (BP) is used to transmit the processing results or model updates from the data processing node to other nodes, i.e., the edge server, so that the edge server receives the model updates from multiple client devices and aggregates or further optimizes them.
[0108] all in all, Figure 3 The data collection, transmission, processing and model training process of a monitoring terminal is specifically introduced. The data flows in the system through forward and reverse transmission. After multiple batches of processing and model training, a trained model is finally generated and the output results are output. The output results here include soft labels (probability distribution) and true values (hard labels), which are mainly used to evaluate model performance.
[0109] In one embodiment, the calculation formula for back propagation in step S1242 may include:
[0110]
[0111] In the formula, Represents the use of loss function The calculated gradient result; among them, represents the gradient, represents the general feature extraction subnetwork in the i-th client device Network parameters in the t-1th update iteration; represents the personalized label mapping subnetwork in the i-th client device Network parameters in the t-1th update iteration; Represents the loss function right and The partial derivative of .
[0112] In this embodiment, back propagation can guide the model on how to adjust parameters to reduce losses. Its formula is mainly used for gradient calculation when the client device optimizes the model locally, ensuring that the model can learn specific abnormal behavior patterns in monitoring scenarios through personalized training. It plays a core role in personalized federated edge learning, enabling the model to maintain high detection accuracy while adapting to different monitoring environments, and ultimately improve the generalization ability of the overall system through the aggregation of edge servers and cloud servers.
[0113] In one embodiment, the calculation formula of the stochastic gradient descent method in step S1143 may include:
[0114]
[0115]
[0116] In the formula, represents the general feature extraction subnetwork in the i-th client device Network parameters in the tth update iteration; represents the learning rate; represents the general feature extraction subnetwork The gradient of Represents the loss function The calculated gradient result; represents the personalized label mapping subnetwork in the i-th client device Network parameters in the tth update iteration; represents the general feature extraction subnetwork gradient.
[0117] In this embodiment, the calculation formula of the stochastic gradient descent method is mainly used for local model training of the client device. The general feature extraction subnetwork and the personalized label mapping subnetwork are continuously optimized through stochastic gradient descent, thereby improving the model's ability to detect abnormal behaviors and adapting to the personalized needs of different monitoring scenarios. Finally, the optimized local model is uploaded to the edge server for asynchronous aggregation, further improving the overall abnormal behavior detection capability of the system.
[0118] In one embodiment, the calculation formula of the hyperparameter control method in step S1243 may include:
[0119]
[0120] In the formula, represents the norm of the parameter vector; Represents the regularization term, which is used to control the personalized label mapping subnetwork The intensity of the personal adjustment.
[0121] In this embodiment, in order to balance the absorption of global knowledge and the retention of local data features, this application introduces a hyperparameter To adjust the degree of learning from the global model, this hyperparameter allows the client device to adjust the degree of personalization according to the characteristics of the local data, balancing the personalization adjustment and stability of the model during local training, ensuring that the optimized model will not undergo drastic parameter changes while adapting to new data, thereby improving the generalization ability and stability of the model.
[0122] Specifically, the calculation formula of the hyperparameter control method is mainly used for the personalized label mapping subnetwork It uses regularization terms to control the parameter update amplitude, thereby balancing the model's personalized adaptability and training stability. ,This application can ensure that the model can learn new behavioral features without affecting its generalization ability and convergence due to ,significant parameter adjustments.
[0123] In one embodiment, the calculation formula of the asynchronous cross-device model aggregation strategy in step S120 may include:
[0124]
[0125]
[0126] In the formula, and Represents the general feature extraction subnetwork in the edge model and personalized label mapping subnetwork Network parameters in the t+1th update iteration; and represents the general feature extraction subnetwork in the i-th client device and personalized label mapping subnetwork Network parameters in the tth update iteration; j represents the jth client device; Represents the set of client devices participating in the current aggregation, Indicates the number of parameters.
[0127] In this embodiment, the formula is mainly used for the edge server to aggregate the local models of the client devices to generate a new edge model, and an asynchronous aggregation mechanism is adopted so that different client devices can submit local models at different times without affecting the aggregation process of the overall model. This strategy fuses the models of the client devices through a weighted average method to improve the accuracy and adaptability of the edge model, while reducing the aggregation waiting time caused by offline or delayed client devices, improving the update efficiency of the edge model, and adapting to the data distribution characteristics of different client devices.
[0128] In one embodiment, the calculation formula of the asynchronous model aggregation strategy in step S130 may include:
[0129]
[0130]
[0131] In the formula, and Respectively represent the general feature extraction subnetwork in the global model and personalized label mapping subnetwork Network parameters; represents the set of all edge servers participating in the aggregation, and i and j represent the i-th and j-th edge servers in the set of all edge servers participating in the aggregation, respectively.
[0132] In this embodiment, the formula is mainly used by the cloud server to asynchronously aggregate the edge server model and finally generate a global model to improve the generalization ability and stability of the model. Here, through the weighted average strategy, the cloud server can integrate the model parameters of different edge servers, so that the global model can learn a wider range of behavior patterns and adapt to the data distribution in different environments. At the same time, it improves the generalization ability and robustness of the model, ensuring that the model can adapt to the data characteristics in different edge environments and achieve more accurate prediction and decision-making.
[0133] In order to better explain the personalized federated edge learning method of this application, Figure 4 and Figure 5 To further illustrate, schematically, Figure 4 and Figure 5 As shown, Figure 4 An interactive schematic diagram of a personalized federated edge learning method provided in an embodiment of the present application; Figure 5 A strategy flow chart of a two-stage model aggregation process provided in an embodiment of the present application.
[0134] Figure 4 In the process, first, multiple client devices collect and transmit data through monitoring devices, and then each performs local model training to generate a local model. After the training is completed, the client device uploads the local model to the corresponding edge server. After the edge server receives the local models from multiple client devices, it aggregates the edge models and generates an edge model. Then, the edge server uploads the aggregated edge model to the cloud server. The cloud server is responsible for global model aggregation and further integrates the edge models from multiple edge servers to generate a global model. After the global model aggregation is completed, the cloud server distributes the updated global model back to each edge server and client device through the FEL task to achieve model sharing, and then conduct the next round of local model training and update. The whole process is iterated continuously until the model reaches the expected performance indicators or goes through the predetermined number of training rounds.
[0135] Figure 5In the system, the edge server acts as a core node, responsible for data interaction and model update with multiple client devices. First, the edge server obtains the global model from the cloud server or other upstream nodes and distributes it to each client device. After receiving the model, the client device performs data processing and model training locally to generate an updated local model. After the training is completed, the client device uploads the updated model to the edge server. The edge server receives model updates from multiple client devices, performs aggregation operations, and generates a new edge model. The aggregated edge model can be further uploaded to the cloud server for global model update, or directly used for local reasoning and task execution. The entire process is coordinated by the edge server, achieving efficient collaboration of distributed model training and update, ensuring the real-time and scalability of the system.
[0136] The personalized federated edge learning device provided in an embodiment of the present application is described below. The personalized federated edge learning device described below and the personalized federated edge learning method described above can be referenced to each other.
[0137] In one embodiment, Figure 6 As shown, Figure 6 A structural diagram of a personalized federated edge learning device provided in an embodiment of the present application; the present application also provides a personalized federated edge learning device, including a model optimization module 210, a primary aggregation module 220, a secondary aggregation module 230 and a model iteration module 240, specifically including the following:
[0138] The model optimization module 210 is used to collect video data through the monitoring terminal, and use the client device to optimize the preset initial model based on the video data to generate a local model.
[0139] The primary aggregation module 220 is used to obtain local models output by multiple client devices using the edge server based on an aggregation threshold, and aggregate each local model using an asynchronous cross-device model aggregation strategy to generate an edge model.
[0140] The secondary aggregation module 230 is used to use the cloud server to periodically adopt an asynchronous model aggregation strategy to globally aggregate the edge models output by multiple edge servers to generate a global model.
[0141] The model iteration module 240 is used to distribute the global model as the latest initial model back to each edge server and each client device, and return to the video data collection through the monitoring terminal and its subsequent steps until the global model meets the preset iteration end condition.
[0142] In the above embodiment, when detecting abnormal crowd behavior, video data can be collected through the monitoring terminal, and the client device can be used to optimize the preset initial model based on the video data to generate a local model, so that the model can adapt to different monitoring scenarios and improve the accuracy and adaptability of detection. Then, the edge server can be used to obtain the local models output by multiple client devices based on the aggregation threshold, and the asynchronous cross-device model aggregation strategy can be used to aggregate each local model to generate an edge model, so as to reduce the aggregation delay caused by equipment problems and improve the efficiency and flexibility of edge model aggregation. Then, the cloud server can be used to periodically adopt an asynchronous model aggregation strategy to globally aggregate the edge models output by multiple edge servers to generate a global model, so that the model can learn a wider range of crowd behavior patterns and characteristics, and achieve global performance improvement; finally, the global model can be distributed back to each edge server and each client device as the latest initial model, and the video data collected through the monitoring terminal and its subsequent steps are returned until the global model meets the preset iteration end condition, so as to ensure that the entire system can continue to learn and adapt to new data and changes while maintaining the real-time and efficiency of the model. In addition, the method of the present application can also reduce the dependence on the central server, reduce the overhead of data transmission, and thus enhance the scalability and robustness of the entire monitoring system.
[0143] In one embodiment, the model optimization module 210 may include:
[0144] The network decomposition submodule is used to use the client device to perform network decomposition on the preset initial model to obtain a general feature extraction subnetwork and a personalized label mapping subnetwork.
[0145] The feature extraction submodule is used to detect abnormal features of video data through a general feature extraction subnetwork, and extract multiple behavioral features based on the detection results.
[0146] The label mapping submodule is used to perform label mapping on each behavior feature through a personalized label mapping subnetwork to obtain a behavior label corresponding to each behavior feature.
[0147] The model optimization submodule is used to optimize the initial model based on various behavioral features and the behavioral labels corresponding to each behavioral feature to obtain a local model.
[0148] In one embodiment, the model optimization submodule may include:
[0149] The label prediction unit is used to input each behavior feature and the behavior label corresponding to each behavior feature into the initial model to obtain the predicted label output by the initial model based on each behavior label.
[0150] The gradient calculation unit is used to back-propagate each predicted label using the loss function to obtain the gradient result of the initial model.
[0151] The parameter updating unit is used to adopt the stochastic gradient descent method and the hyperparameter control method to update the network parameters of the general feature extraction subnetwork and the personalized label mapping subnetwork based on the gradient results.
[0152] In one embodiment, the gradient calculation unit may include:
[0153]
[0154] In the formula, Represents the use of loss function The calculated gradient result; among them, represents the gradient, represents the general feature extraction subnetwork in the i-th client device Network parameters in the t-1th update iteration; represents the personalized label mapping subnetwork in the i-th client device Network parameters in the t-1th update iteration; Represents the loss function right and The partial derivative of .
[0155] In one embodiment, the parameter updating unit may include:
[0156]
[0157]
[0158] In the formula, represents the general feature extraction subnetwork in the i-th client device Network parameters in the tth update iteration; represents the learning rate; represents the general feature extraction subnetwork The gradient of Represents the loss function The calculated gradient result; represents the personalized label mapping subnetwork in the i-th client device Network parameters in the tth update iteration; represents the general feature extraction subnetwork gradient.
[0159] In one embodiment, the parameter updating unit may further include:
[0160]
[0161] In the formula, represents the norm of the parameter vector; Represents the regularization term, which is used to control the personalized label mapping subnetwork The intensity of the personal adjustment.
[0162] In one embodiment, the primary aggregation module 220 may include:
[0163]
[0164]
[0165] In the formula, and Represents the general feature extraction subnetwork in the edge model and personalized label mapping subnetwork Network parameters in the t+1th update iteration; and represents the general feature extraction subnetwork in the i-th client device and personalized label mapping subnetwork Network parameters in the tth update iteration; j represents the jth client device; Represents the set of client devices participating in the current aggregation, Indicates the number of parameters.
[0166] In one embodiment, the secondary aggregation module 230 may include:
[0167]
[0168]
[0169] In the formula, and Respectively represent the general feature extraction subnetwork in the global model and personalized label mapping subnetwork Network parameters; represents the set of all edge servers participating in the aggregation, and i and j represent the i-th and j-th edge servers in the set of all edge servers participating in the aggregation, respectively.
[0170] In one embodiment, the present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the personalized federated edge learning method as described in any of the above embodiments.
[0171] In one embodiment, the present application also provides a computer device having computer-readable instructions stored therein. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the personalized federated edge learning method as described in any one of the above embodiments.
[0172] Indicatively, if Figure 7 As shown, Figure 7 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. The computer device 300 may be provided as a server. Figure 7 The computer device 300 includes a processing component 302, which further includes one or more processors, and a memory resource represented by a memory 301 for storing instructions executable by the processing component 302, such as an application. The application stored in the memory 301 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 302 is configured to execute instructions to perform the personalized federated edge learning method of any of the above embodiments.
[0173] The computer device 300 may further include a power supply component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 may operate based on an operating system stored in the memory 301, such as Windows Server TM, Mac OS X TM, Unix TM, Linux TM, Free BSD TM, or the like.
[0174] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0175] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0176] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.
[0177] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A personalized federated edge learning method, characterized in that: The method comprises: Collecting video data through a monitoring terminal, and optimizing a preset initial model based on the video data using a client device to generate a local model; The edge server is used to obtain the local models output by multiple client devices based on the aggregation threshold, and the asynchronous cross-device model aggregation strategy is used to aggregate the local models to generate the edge model. The cloud server is used to periodically adopt an asynchronous model aggregation strategy to globally aggregate the edge models output by multiple edge servers to generate a global model. The global model is distributed back to each edge server and each client device as the latest initial model, and the video data is collected through the monitoring terminal and its subsequent steps are returned until the global model meets the preset iteration end condition.
2. The personalized federated edge learning method according to claim 1, characterized in that: The method of using the client device to perform model training and optimization on a preset initial model based on the video data to generate a local model includes: Using the client device to perform network decomposition on the preset initial model, a general feature extraction sub-network and a personalized label mapping sub-network are obtained; Performing abnormal feature detection on the video data through the general feature extraction sub-network, and extracting multiple behavior features according to the detection results; Perform label mapping on each behavior feature through the personalized label mapping subnetwork to obtain a behavior label corresponding to each behavior feature; The initial model is optimized based on each behavior feature and a behavior label corresponding to each behavior feature to obtain a local model.
3. The personalized federated edge learning method according to claim 2, characterized in that: The optimizing the initial model based on each behavior feature and the behavior label corresponding to each behavior feature to obtain a local model includes: Inputting each behavior feature and the behavior label corresponding to each behavior feature into the initial model to obtain a predicted label output by the initial model based on each behavior label; Back-propagating each predicted label using the loss function to obtain the gradient result of the initial model; A stochastic gradient descent method and a hyperparameter control method are adopted to update the network parameters of the general feature extraction subnetwork and the personalized label mapping subnetwork respectively based on the gradient results.
4. The personalized federated edge learning method according to claim 3, characterized in that: The calculation formula of the back propagation includes: In the formula, Represents the use of loss function The calculated gradient result; among them, represents the gradient, represents the general feature extraction subnetwork in the i-th client device Network parameters in the t-1th update iteration; represents the personalized label mapping subnetwork in the i-th client device Network parameters in the t-1th update iteration; Represents the loss function right and The partial derivative of .
5. The personalized federated edge learning method according to claim 3, characterized in that: The calculation formula of the stochastic gradient descent method includes: In the formula, represents the general feature extraction subnetwork in the i-th client device Network parameters in the tth update iteration; represents the learning rate; Represents the general feature extraction subnetwork The gradient of Represents the loss function The calculated gradient result; represents the personalized label mapping subnetwork in the i-th client device Network parameters in the tth update iteration; Represents the general feature extraction subnetwork gradient.
6. The personalized federated edge learning method according to claim 3, characterized in that: The calculation formula of the hyperparameter control method includes: In the formula, represents the norm of the parameter vector; Represents the regularization term, which is used to control the personalized label mapping subnetwork The intensity of the personal adjustment.
7. The personalized federated edge learning method according to claim 1, characterized in that: The calculation formula of the asynchronous cross-device model aggregation strategy includes: In the formula, and Represents the general feature extraction subnetwork described in the edge model and the personalized label mapping subnetwork Network parameters in the t+1th update iteration; and represents the general feature extraction subnetwork in the i-th client device and the personalized label mapping subnetwork Network parameters in the tth update iteration; j represents the jth client device; Represents the set of client devices participating in the current aggregation, Indicates the number of parameters.
8. The personalized federated edge learning method according to claim 1, characterized in that: The calculation formula of the asynchronous model aggregation strategy includes: In the formula, and Respectively represent the general feature extraction subnetworks described in the global model and the personalized label mapping subnetwork Network parameters; represents the set of all edge servers participating in the aggregation, and i and j represent the i-th and j-th edge servers in the set of all edge servers participating in the aggregation, respectively.
9. A personalized federated edge learning device, characterized in that: include: A model optimization module, used to collect video data through a monitoring terminal, and optimize a preset initial model based on the video data using a client device to generate a local model; A primary aggregation module is used to obtain local models output by multiple client devices using an edge server based on an aggregation threshold, and to aggregate each local model using an asynchronous cross-device model aggregation strategy to generate an edge model; A secondary aggregation module is used to globally aggregate edge models output by multiple edge servers using a cloud server periodically using an asynchronous model aggregation strategy to generate a global model; The model iteration module is used to distribute the global model as the latest initial model back to each edge server and each client device, and return to the video data collected by the monitoring terminal and its subsequent steps until the global model meets the preset iteration end condition.
10. A storage medium, characterized in that: The storage medium stores computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the personalized federated edge learning method as described in any one of claims 1 to 8.
11. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, which, when executed by the one or more processors, perform the steps of the personalized federated edge learning method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Asynchronous aggregation and privacy protection method in resource limited federated edge learning
CN116911382A
Client selection and personalized privacy protection method in asynchronous federated edge learning
CN117252253A
Personalized federated learning intrusion detection method for edge isomerism
CN117375947A
Federal learning client resource heterogeneous method and system for edge intelligence
CN118690873A
Edge anomaly detection method based on global-personalized collaborative federated learning
CN119202839A