Self-adaptive distribution line component detection system based on federal learning and knowledge distillation
Through federal learning and knowledge distillation technology, combined with adaptive feature fusion and active learning, efficient and accurate detection of distribution line components is achieved, solving the shortcomings of traditional manual acceptance methods, adapting to complex environments and providing continuous optimization capabilities.
Patent Information
- Application Number
- CN202510430555.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
The traditional manual acceptance method is inefficient, incomplete coverage, insufficient professional quality, difficult to cope with complex environments, chaotic data management, high cost and poor real-time performance, and cannot meet the intelligent and automated distribution line component inspection needs.
Adaptive distribution line component detection system based on federated learning and knowledge distillation is adopted, and distributed training and optimization of the model is realized through federated learning module, knowledge distillation module, adaptive feature fusion module and active learning module, real-time detection is performed using lightweight student models, and valuable samples are selected through active learning mechanisms for labeling and update.
It realizes efficient, accurate and adaptive power distribution line component detection under the premise of protecting data privacy, improves detection efficiency and accuracy, reduces labor costs, adapts to complex environments and provides continuous optimization capabilities.
Smart Images

Figure CN120354882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distribution line component detection, and in particular to an adaptive distribution line component detection system based on federated learning and knowledge distillation. Background Art
[0002] With the continuous development of the distribution system, the detection of distribution line components has become increasingly important. Traditional construction acceptance methods rely on manual participation. In the construction sites of distribution networks with complex terrains and wide - area distributions, the following problems exist in manual acceptance:
[0003] Low efficiency: Manual acceptance requires a large amount of manpower and time, and it is difficult to meet the needs of large - scale distribution line component detection.
[0004] Incomplete coverage: In construction sites with complex terrains and wide - area distributions, it is difficult for manual acceptance to comprehensively cover all line components, and some key parts are easily missed.
[0005] Insufficient professional quality: The professional levels of different acceptance personnel vary, which may lead to inadequate quality supervision and potential safety hazards.
[0006] Difficulty in coping with complex environments: In complex environments such as bad weather and insufficient light, the accuracy and reliability of manual acceptance will be greatly affected.
[0007] Chaotic data management: The data recording and management methods of manual acceptance are relatively backward, making it difficult to conduct effective data analysis and traceability.
[0008] High cost: Manual acceptance requires a large amount of manpower and material resources, increasing the cost of distribution line maintenance.
[0009] Poor real - time performance: Manual acceptance cannot provide real - time feedback on detection results, making it difficult to detect and handle problems in a timely manner.
[0010] The existence of these problems has prompted people to seek an intelligent and automated distribution line component detection solution to improve detection efficiency, accuracy, and reliability, and reduce labor costs and safety risks. Summary of the Invention
[0011] In view of this, an embodiment of the present invention provides an adaptive distribution line component detection system based on federated learning and knowledge distillation to at least partially solve the above problems.
[0012] An adaptive distribution line component detection system based on federated learning and knowledge distillation according to an embodiment of the present invention includes: a central server and edge devices, wherein a federated learning module, a knowledge distillation module, an adaptive feature fusion module, and an active learning module are deployed on the central server and the edge devices; the federated learning module is used for each participant on the edge device to train an object detection model and regularly send the updated object detection model to the central server, and the central server uses a secure aggregation algorithm to merge and update the object detection model to obtain a global model, and distributes the global model back to each participant; the knowledge distillation module is used for training a complex teacher model based on the global model and compressing the knowledge of the teacher model into a lightweight student model; the adaptive feature fusion module is used for the edge device to dynamically adjust the feature extraction and fusion strategy based on the lightweight student model and the current environmental conditions to perform real-time detection on the distribution line components; the active learning module is used for the adaptive distribution line component detection system to select the most valuable samples through uncertainty estimation and sample screening algorithms, request manual annotation as new samples, and use the new samples for updating the object detection model.
[0013] Optionally, the detection method of the adaptive distribution line component detection system includes the following steps: Step 1, initialization: each participant performs data preprocessing and initial model training locally; Step 2, federated learning: each participant regularly sends the updated object detection model to the central server, and the server uses a secure aggregation algorithm to merge and update to obtain a global model, and distributes the global model back to each participant; Step 3, knowledge distillation: training a complex teacher model based on the global model, and then compressing the knowledge of the complex teacher model into a lightweight student model; Step 4, edge deployment: deploying the distilled lightweight student model to the edge device; Step 5, adaptive feature fusion: the edge device dynamically adjusts the feature extraction and fusion strategy according to the current environmental conditions; Step 6, real-time detection: the edge device uses the distilled lightweight student model to perform real-time detection on the distribution line components; Step 7, active learning: the adaptive distribution line component detection system identifies samples that are difficult to process, requests manual annotation as new samples, and uses the new samples for updating the object detection model; Step 8, continuous optimization: regularly collecting the feedback of the edge device, updating the global model, and repeating Steps 2 to 8 to form a closed-loop optimization;
[0014] Optionally, the federated learning module adopts the MFedAvg algorithm, and the MFedAvg algorithm screens and aggregates edge nodes through model distance to exclude abnormal nodes to ensure the process of federated learning, which specifically includes client local training, central server model aggregation, and node status detection links.
[0015] Optionally, the knowledge distillation module adopts a multi-teacher distillation strategy, including the following steps: training multiple complex teacher models; representing the knowledge of the teacher models through the soft labels output by the teacher models; training a small student model using the soft labels provided by the teacher models, and the student model learns how to make decisions during the process of imitating the teacher models; optimizing the knowledge distillation process by adjusting the temperature parameter, so that the student model can better absorb the knowledge of the teacher models.
[0016] Optionally, the process of optimizing the knowledge distillation by adjusting the temperature parameter has the following formula:
[0017]
[0018] where, q i represents the probability of class i, and its value range is [0, 1], Z i represents the logits output by the model in the linear layer, that is, the original score before applying the Softmax function, and T represents the temperature parameter used to adjust the logits.
[0019] Optionally, the adaptive feature fusion module includes an AWF module and a cross-scale cross-stage network. Embed the AWF module and the cross-scale cross-stage network into the residual blocks of the YOLO backbone network, and automatically adjust the feature extraction strategy according to the input image quality and lighting conditions to perform adaptive weighting on features at different levels and scales.
[0020] Optionally, the AWF module consists of three parts: compression, extraction, and assignment. Among them, the compression part uses 1×1 convolution on the two input feature maps to compress the number of channels to a constant T and refine the feature information. The compressed feature map is used as an intermediate feature map to extract weight information; the extraction part first concatenates the two intermediate feature maps at the channel level to obtain a combined feature map with 2T channels, and then uses 1×1 convolution to further compress the number of channels of the combined feature map, and at the same time extracts the spatial weights of the input feature maps. After compressing the channels, a weight feature map with 2 channels is obtained. The two channels of the weight feature map respectively contain the weight information of the two input feature maps. Since these generated weights are obtained by convolving the same feature map, there is a dependency relationship between them, which is used to control the weighted fusion of the two input feature maps. Then, use the following formula to map the value of the weight between 0 and 1:
[0021]
[0022] In the formula, α i,j and β i,j respectively represent the parameter values at (i, j) in the two channels, and Represent two finally obtained spatial weights; after obtaining the spatial weights, perform an allocation operation, match the obtained spatial weights with the initially input feature maps, and by multiplying these weights with the corresponding input feature maps, the dependency relationships between the weights can be passed to the input feature maps, thereby establishing correlations between the input feature maps; add these correlated input feature maps to form the final output feature map, thus realizing the adaptive weighted fusion of features.
[0023] Optionally, the calculation formula of the output feature map is as follows:
[0024] W = Conv(Concat[Conv(C1), Conv(C2)])
[0025] E = W[0] × C1 + W[1] × C2
[0026] Wherein, C1 and C2 represent input feature maps, W represents a weight feature map, E represents an output feature map, Conv(·) represents a convolution operation, Concat[·] represents a concatenation operation along the channels, W[0] is the first channel of the weight feature map, containing the weight information of the first input feature map, and W[1] is the second channel of the weight feature map, containing the weight information of the second feature map.
[0027] Optionally, the active learning module obtains multi-scale features by using the feature extraction network of the object detection model, fuses the multi-scale features as the feature representation of a single unlabeled image, provides pseudo-labels through the K-Means clustering method, combines uncertainty estimation, sorts the uncertainties of the unlabeled images according to the size under different pseudo-label classifications, selects samples for annotation according to the annotation budget, updates the training set, and retrains the object detection model.
[0028] Optionally, the active learning module uses a calculation method based on information entropy to estimate the uncertainty of samples, and the formula is:
[0029]
[0030] Wherein, θ represents the parameters of a trained deep learning model, and P θ (y i (|x) represents the probability value of each classification of the sample x, represents the information entropy of each sample x.
[0031] In summary, the system of the present invention can, on the premise of protecting data privacy, make full use of distributed data resources to achieve efficient, accurate, and adaptive detection of distribution line components, not only solve the current technical challenges, but also provide an extensible and sustainable solution for the future development of smart grids. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0033] Figure 1 It is a structural block diagram of the adaptive distribution line component detection system based on federated learning and knowledge distillation of the present invention.
[0034] Figure 2 For Figure 1 The corresponding schematic diagram of the system architecture of the present invention.
[0035] Figure 3 It is the overall architecture design diagram of the MFedAvg algorithm.
[0036] Figure 4 It is the local training flow chart of the client.
[0037] Figure 5 It is the schematic diagram of model distance calculation.
[0038] Figure 6 It is the schematic diagram of distance-weighted aggregation.
[0039] Figure 7 It is the schematic diagram of the idea of redundant node exclusion.
[0040] Figure 8 It is the principle diagram of knowledge distillation.
[0041] Figure 9 It is the structural diagram of the AWF module.
[0042] Figure 10 It is the structural diagram of the residual block.
[0043] Figure 11 It is the schematic diagram of the cross-scale and cross-stage feature fusion network.
[0044] Figure 12 It is the structural diagram of the improved feature fusion network imPANet.
[0045] Figure 13 It is the structural diagram of the YOLO-AWF network.
[0046] Figure 14 It is the flow chart of the active learning algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] To have a clearer understanding of the technical features, objectives, and effects of the embodiments of the present invention, the specific implementation manners of the embodiments of the present invention will now be described with reference to the accompanying drawings.
[0048] In this document, "exemplarily" means "serving as an instance, example, or illustration", and any illustration or implementation manner described as "exemplarily" in this document should not be construed as a more preferred or more advantageous technical solution.
[0049] For ease of understanding, before describing the specific embodiments of the present invention in detail, the professional terms related to the present invention will be explained first:
[0050] (1) Federated learning: It is an emerging machine learning method that allows multiple devices or computing nodes to perform model training without sharing the original data. The training of the model is carried out on local devices, and only the updated parameters of the model are aggregated to the central server for aggregation.
[0051] (2) Secure aggregation algorithm: It is a technology that aims to protect the privacy of user data while ensuring the accuracy and security of the aggregation result. This algorithm is usually applied in a federated learning environment where multiple participants jointly participate in model training without sharing their local data. The secure aggregation algorithm uses encryption technology, secure computing protocols, etc. to ensure the privacy of user data during the aggregation process, and at the same time verifies the correctness of the aggregation result to prevent malicious attacks or data tampering.
[0052] (3) Knowledge distillation: It is a method of model compression. By constructing a lightweight small model and using the supervision information of a larger model with better performance to train this small model, in order to achieve better performance and accuracy.
[0053] (4) Feature extraction network: Through image analysis and transformation, it extracts the required features, such as corners, edges, etc. This process involves transforming a set of measured values of a certain pattern to highlight the representative features of this pattern. Feature extraction is not only about finding and extracting specific information in an image, but also includes transforming and processing this information for better use in subsequent image recognition, classification, or understanding tasks.
[0054] (5) Feature fusion algorithm: It is a method of combining multiple different types of features together to improve the performance of the model. Its goal is to effectively combine different feature information to extract more representative and discriminative feature representations to improve the prediction ability of the model.
[0055] (6) Active learning: A machine learning or artificial intelligence method that actively selects the most valuable samples for annotation. Its purpose is to use as few high-quality sample annotations as possible to enable the model to achieve the best possible performance. It can improve the gain of samples and annotations, maximize the performance of the model under the premise of limited annotation budgets, and is a solution to improve data efficiency from the perspective of samples. Therefore, it is applied to tasks with high annotation costs and difficult annotation.
[0056] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present invention.
[0057] The following further illustrates the specific implementation of the embodiments of the present invention in conjunction with the accompanying drawings of the embodiments of the present invention.
[0058] See Figure 1 , an adaptive distribution line component detection system based on federated learning and knowledge distillation provided by the present invention includes:
[0059] A central server 1 and edge devices 2, where a federated learning module 110, a knowledge distillation module 120, an adaptive feature fusion module 210, and an active learning module 220 are deployed on the central server 1 and the edge devices 2;
[0060] The federated learning module 110 is used for each participant on the edge device 2 to train an object detection model and regularly send the updated object detection model to the central server 1. The central server 1 uses a secure aggregation algorithm to merge and update the object detection model to obtain a global model and distribute the global model back to each participant;
[0061] The knowledge distillation module 120 is used to train a complex teacher model based on the global model and compress the knowledge of the teacher model into a lightweight student model;
[0062] The adaptive feature fusion module 210 is used for the edge device to dynamically adjust the feature extraction and fusion strategy based on the lightweight student model and the current environmental conditions to perform real-time detection of distribution line components;
[0063] The active learning module 220 is used for the adaptive distribution line component detection system to select the most valuable samples for manual annotation as new samples through uncertainty estimation and sample screening algorithms and use the new samples for updating the object detection model.
[0064] Optionally, the detection method of the adaptive distribution line component detection system includes the following steps:
[0065] Step 1, Initialization: Each participant performs data preprocessing and initial model training locally.
[0066] Step 2, Federated Learning: Each participant periodically sends the updated target detection model to the central server. The server uses a secure aggregation algorithm to merge the updates to obtain a global model and distributes the global model back to each participant.
[0067] Step 3, Knowledge Distillation: Train a complex teacher model based on the global model, and then compress the knowledge of the complex teacher model into a lightweight student model.
[0068] Step 4, Edge Deployment: Deploy the distilled lightweight student model to the edge device.
[0069] Step 5, Adaptive Feature Fusion: The edge device dynamically adjusts the feature extraction and fusion strategy according to the current environmental conditions.
[0070] Step 6, Real-time Detection: The edge device uses the distilled lightweight student model to perform real-time detection on the distribution line components.
[0071] Step 7, Active Learning: The adaptive distribution line component detection system identifies difficult-to-process samples, requests manual annotation as new samples, and uses the new samples for updating the target detection model.
[0072] Step 8, Continuous Optimization: Regularly collect the feedback from the edge device, update the global model, and repeat Steps 2 to 8 to form a closed-loop optimization.
[0073] Optionally, the federated learning module adopts the MFedAvg algorithm. The MFedAvg algorithm screens and aggregates edge nodes through model distance, excludes abnormal nodes, and guarantees the process of federated learning, which specifically includes the client local training, central server model aggregation, and node status detection links.
[0074] Optionally, the knowledge distillation module adopts a multi-teacher distillation strategy, including the following steps: Train multiple complex teacher models; Represent the knowledge of the teacher models through the soft labels output by the teacher models; Use the soft labels provided by the teacher models to train a small student model, and the student model learns how to make decisions in the process of imitating the teacher models; Optimize the process of knowledge distillation by adjusting the temperature parameter so that the student model can better absorb the knowledge of the teacher models.
[0075] Optionally, the process of optimizing knowledge distillation by adjusting temperature parameters is formulated as follows:
[0076]
[0077] where q i represents the probability of class i, with a value range of [0, 1], Z i represents the logits output by the model in the linear layer, i.e., the raw scores before applying the Softmax function, and T represents the temperature parameter used to adjust the logits.
[0078] Optionally, the adaptive feature fusion module includes an AWF module and a cross-scale cross-stage network. The AWF module and the cross-scale cross-stage network are embedded in the residual blocks of the YOLO backbone network, and the feature extraction strategy is automatically adjusted according to the input image quality and lighting conditions to adaptively weight features at different levels and scales.
[0079] Optionally, the AWF module consists of three parts: compression, extraction, and assignment. In the compression part, 1×1 convolutions are used on the two input feature maps to compress the number of channels to a constant T and refine the feature information. The compressed feature maps are used as intermediate feature maps for extracting weight information. In the extraction part, first, the two intermediate feature maps are concatenated channel-wise to obtain a combined feature map with 2T channels. Then, 1×1 convolutions are used to further compress the number of channels of the combined feature map while extracting the spatial weights of the input feature maps. After compressing the channels, a weight feature map with 2 channels is obtained. The two channels of the weight feature map respectively contain the weight information of the two input feature maps. Since these generated weights are obtained by convolving the same feature map, there is a dependency relationship between them, which is used to control the weighted fusion of the two input feature maps. Then, the following formula is used to map the values of the weights between 0 and 1:
[0080]
[0081] In the formula, α i,j and β i,j respectively represent the parameter values at (i, j) in the two channels, and represent the two finally obtained spatial weights. After obtaining the spatial weights, the assignment operation is performed. The obtained spatial weights are matched with the originally input feature maps. By multiplying these weights with the corresponding input feature maps, the dependency relationship between the weights can be passed to the input feature maps, thereby establishing a correlation between the input feature maps. The input feature maps with established correlations are added together to form the final output feature map, thus realizing the adaptive weighted fusion of features.
[0082] Optionally, the calculation formula of the output feature map is as follows:
[0083] W = Conv(Concat[Conv(C1), Conv(C2)])
[0084] E = W[0] × C1 + W[1] × C2
[0085] Among them, C1 and C2 represent the input feature maps, W represents the weight feature map, E represents the output feature map, Conv(·) represents the convolution operation, Concat[·] represents the concatenation operation along the channels, W[0] is the first channel of the weight feature map, which contains the weight information of the first input feature map, and W[1] is the second channel of the weight feature map, which contains the weight information of the second feature map.
[0086] Optionally, the active learning module obtains multi-scale features by using the feature extraction network of the object detection model, fuses the multi-scale features as the feature representation of a single unlabeled image, provides pseudo-labels through the K-Means clustering method, combines uncertainty estimation, sorts the uncertainties of the unlabeled images according to their magnitudes under different pseudo-label classifications, selects samples for annotation according to the annotation budget, updates the training set, and retrains the object detection model.
[0087] Optionally, the active learning module uses a calculation method based on information entropy to estimate the uncertainty of samples, and the formula is:
[0088]
[0089] Among them, θ represents the parameters of a trained deep learning model, P θ (y i |x) represents the probability value of each classification of sample x, represents the information entropy of each sample x.
[0090] In summary, the federated learning module based on the improved federated averaging algorithm (MFedAvg) in the present invention aggregates the model updates of each participant using a weighted average strategy, and the weights are dynamically adjusted according to the data quality and quantity. This method not only protects data privacy but also improves the overall performance of the model. The knowledge distillation module adopts a multi-teacher distillation strategy. This module first trains multiple "teacher" models specifically for different types of distribution line components, and then compresses the knowledge into a lightweight "student" model through a multi-task learning framework. This method not only transfers soft labels but also includes the matching of feature maps, ensuring that the student model can capture richer knowledge. The present invention also includes an adaptive feature weighting module, namely the AWF module, and an adaptive feature fusion module for cross-scale and cross-stage networks, which can automatically adjust the feature extraction strategy according to factors such as the quality and lighting conditions of the input image, and adaptively weight features at different levels and scales. The present invention also includes an active learning module based on pseudo-labels. Starting from the perspectives of sample uncertainty and diversity, this module uses the feature extraction network of the object detection model to obtain multi-scale features, and fuses them as the feature representation of unlabeled images. It also provides pseudo-labels through a clustering method, and combines uncertainty estimation to select the most valuable samples for annotation. The present invention organically combines advanced technologies such as federated learning, knowledge distillation, adaptive feature fusion, and active learning to form a complete, efficient, and sustainable optimization distribution line component detection system.
[0091] Specifically, the solution of the present invention is further described according to the following examples:
[0092] The adaptive distribution line component detection system based on federated learning and knowledge distillation in the present invention aims to achieve efficient, accurate, and safe detection of distribution line components. It mainly includes a federated learning module, a knowledge distillation module, an adaptive feature fusion module, and an active learning module.
[0093] Federated learning module: The system adopts the improved federated averaging algorithm (MFedAvg). Each participant trains a model locally using private data and only transmits the model parameters or gradient updates to the central server. The server aggregates the updates using a weighted average strategy, and the weights are dynamically adjusted according to the data quality and quantity of each participant.
[0094] Knowledge distillation module: This module adopts a multi-teacher distillation strategy. First, multiple "teacher" models specifically for different types of distribution line components are trained based on the global model. Then, a lightweight "student" model is designed to learn from multiple teacher models simultaneously through a multi-task learning framework. The distillation process not only transfers soft labels (the output probabilities of the teacher models) but also includes the matching of feature maps to ensure that the student model can capture richer knowledge.
[0095] Adaptive Feature Fusion Module: This module designs an Adaptive Feature Weighting Module (AWF) and a cross-scale and cross-stage network, which are embedded in the YOLO residual block. It can automatically adjust the feature extraction strategy according to factors such as the quality of the input image and lighting conditions, and adaptively weight features at different levels and scales.
[0096] Active Learning Module: From the perspectives of sample uncertainty and diversity, a pseudo-label based active learning method is proposed for the characteristics of object detection algorithms. The feature extraction network of the object detection model is used to obtain multi-scale features, and the multi-scale features are fused as the feature representation of a single unlabeled image.
[0097] The system of the present invention can make full use of distributed data resources under the premise of protecting data privacy, and achieve efficient, accurate and adaptive detection of distribution line components. It not only solves the current technical challenges, but also provides an extensible and sustainable solution for the future development of smart grids.
[0098] See Figure 2 , the detection method of the adaptive distribution line component detection system of the present invention includes the following steps:
[0099] Step 1. Initialization: Each participant performs data preprocessing and initial model training locally.
[0100] Step 2. Federated Learning: The participants periodically send model updates to the central server, and the server uses a secure aggregation algorithm to merge the updates and distribute the global model back to each participant.
[0101] Step 3. Knowledge Distillation: Train a complex "teacher" model based on the global model, and then compress its knowledge into a lightweight "student" model.
[0102] Step 4. Edge Deployment: Deploy the distilled lightweight model to edge devices.
[0103] Step 5. Adaptive Feature Fusion: The edge device dynamically adjusts the feature extraction and fusion strategy according to the current environmental conditions.
[0104] Step 6. Real-time Detection: The edge device uses the optimized model to perform real-time detection of distribution line components.
[0105] Step 7. Active Learning: The system identifies difficult-to-process samples, requests manual annotation, and uses the new samples for model update.
[0106] Step 8. Continuous Optimization: Regularly collect feedback from edge devices, update the global model, and repeat steps 2-8 to form a closed-loop optimization.
[0107] The main feature of Step 2 is the application of the aggregation and screening mechanism of edge nodes based on model distance (MFedAvg). In this algorithm, the central server will use an auxiliary dataset to train a reference model, and aggregate the global model based on the distance between the local models of each edge computing node (client) and the reference model. In addition, the MFedAvg algorithm uses a node status detection mechanism to intelligently exclude abnormal nodes that generate redundant data during the federated learning process, thus protecting the federated learning process. Finally, when the number of abnormal nodes is too large, MFedAvg will consider that the auxiliary dataset is no longer applicable to this federated learning process and will re-obtain a new auxiliary dataset.
[0108] See the overall architecture model design diagram of MFedAvg in Figure 3 , this model architecture divides the edge computing scenario into three levels: the central server, edge computing nodes, and end users. Different edge computing nodes will collect specific business data from different users. Since the user overlap of different edge computing nodes is relatively low, and the number of edge nodes and user terminals is large, the data collected by different edge nodes is very likely to be non-independent and identically distributed. This model architecture takes into account the possible existence of malicious attack nodes in edge computing nodes to generate a large amount of redundant data. After obtaining the latest global model from the central server, the edge computing node performs local model training on the local data, and then the edge computing node uploads its respective local model to the central server. The central server itself also maintains an auxiliary model in each round, and aggregates each local model by weighting according to the distance between each local model and the auxiliary model, so as to obtain the global model. For malicious redundant nodes, the central server also identifies malicious nodes through the model distance and excludes the node in subsequent federated learning to avoid its impact on the overall accuracy and convergence speed.
[0109] The MFedAvg algorithm has several key technical points such as local training of the client (computing node), weighted model aggregation by the central server, and node status detection. Each key point is elaborated as follows:
[0110] Client local training: During the process of federated learning, each edge computing node (client) that needs to participate in the training needs to perform local training on the local dataset in each round of the federated learning process. The process of local training is essentially a deep learning process, and the stochastic gradient descent (SGD) method is used to find the model parameters that minimize the local loss. The flowchart of the client local training of the MFedAvg algorithm is as Figure 4 shown, and the process of client local training can be elaborated as follows:
[0111] (1) In the process of local training on the client side, the client node needs to split the local dataset into several data batches of equal size to obtain the set B;
[0112] (2) For each small batch of data b ∈ B in the set B, the method of stochastic gradient descent is used to adjust the model parameters.
[0113] (3) Steps (1) and (2) need to be repeated for several rounds, that is, each client needs to perform local training for E rounds;
[0114] (4) After the local training on the client side is completed, the client needs to upload the currently obtained local model of the client to the central server for subsequent aggregation operations.
[0115] Central server model aggregation: The central server itself also maintains an auxiliary independent and identically distributed dataset internally (this dataset can be purchased from a small number of clients or downloaded from an open-source public platform), and in each round of federated learning, an auxiliary model needs to be obtained based on the auxiliary set. After the central server collects the local models of each client node, it will determine how much weight should be assigned to each client when aggregating all local models by referring to the model distance between each local model and the auxiliary model.
[0116] The MFedAvg algorithm needs to determine the weight for aggregating the models of each client based on the distance between the local model of the client and the auxiliary model in this round of federated learning.
[0117] As Figure 5 shown, for two models with determined model parameters, since the shapes of the two models are the same, the neurons in the two models can uniquely determine a one-to-one correspondence. Therefore, only by calculating the overall Euclidean distance based on the neuron parameters at the corresponding positions can represent the distance between the two models.
[0118] Based on the model distance, the central server assigns different weights to the models of different clients and obtains the final global model of the current round. For a certain round of federated learning, the process of weighted aggregation by the central server is as Figure 6 shown and can be described as:
[0119] (1) In each round of federated learning, the central server needs to combine the global model W of the previous round all with the auxiliary IID dataset to train and obtain the auxiliary model W of the current round example .
[0120] (2) At the same time, the central server needs to collect the local models {W of each participating client (assuming there are N client nodes in total) i} where \(i = 1, 2, \ldots, N\). After the central server collects all local models, it will calculate the distance between each client's local model and the auxiliary model using a formula to obtain a distance set \(\{d\) i}, where \(i = 1, 2, \ldots, N\).
[0121] (3) Subsequently, for a certain client's local model \(d\) i , the weight calculation formula during its aggregation is as follows:
[0122]
[0123] In this way, it can be ensured that the client model closer to the auxiliary model has a larger weight proportion during aggregation, so that the final model is closer to the ideal model.
[0124] (4) Finally, the central server calculates the global model of the current round according to the weight mapping formula. The formula is as follows:
[0125]
[0126] Ultimately, according to the above process, the central server can obtain the global model of the current round. If there is a next round of federated learning training, this global model will be pulled down by each client node and a new round of local training and model aggregation will be started.
[0127] Node status detection: Since there may be some malicious redundant nodes during the edge computing process to generate a large amount of dirty data to interfere with the federated learning training process, MFedAvg designs a redundant node elimination idea for this situation. During the federated learning training process, it monitors the status of each client node in real time. If a certain client node is considered an abnormal node, then in the subsequent training process, this client node will be excluded and no longer participate in the training. The overall idea of this redundant node elimination is as Figure 7 shown. The idea of the central server to identify and exclude redundant nodes can be summarized as follows:
[0128] (1) If the distance \(d\) i between a certain computing node \(W\) example and \(W\) i exceeds a specific threshold \(\rho\), or the test accuracy of this computing node drops significantly compared to the previously extracted computing accuracy, then this computing node will be put into the abnormal node set \(C\).
[0129] (2) When the central server of the subsequent federated learning collects data from client nodes, if a certain client node is already in the abnormal node set, then this client node will be directly removed to avoid reducing the overall federated learning accuracy due to its data redundancy.
[0130] MFedAvg uses the above-mentioned outlier node aggregation method based on model distance to identify and detect outlier nodes with data redundancy, so as to ensure that the quality of the nodes that perform local training and provide local models during the federated learning process is relatively high, thereby improving the accuracy and convergence speed of the overall model.
[0131] The process of the MFedAvg algorithm is described from the perspectives of client computing nodes and the central server.
[0132] (1) Client Node (Client)
[0133] In a certain round of federated learning training, all participating client nodes need to download and obtain the global model W of the previous round from the central server. all ;
[0134] Subsequently, the client starts several rounds of local training from the current model parameters and obtains the local model W. i ;
[0135] Finally, the client uploads the local model to the central server.
[0136] (2) Central Server (Server)
[0137] The central server first needs to select N client nodes for local training.
[0138] While the client nodes are performing local training, the central server also needs to start using the auxiliary dataset to train and obtain the reference model W from W. all ; example ;
[0139] When all participating clients have completed local training and uploaded their respective local models, the central server needs to calculate the model distances between all local models W and the reference model W respectively, and obtain {d i , i = 1, 2,... N; if the distance of a certain client model exceeds the threshold of node status detection, it needs to be marked as an outlier and does not participate in subsequent federated learning training. example}, i = 1, 2,... N; if the distance of a certain client model exceeds the threshold of node status detection, it needs to be marked as an outlier and does not participate in subsequent federated learning training. i}, i = 1, 2,... N; if the distance of a certain client model exceeds the threshold of node status detection, it needs to be marked as an outlier and does not participate in subsequent federated learning training.
[0140] Subsequently, the central server uses 1 / d i as the weight reference, so that the client models closer to the auxiliary model have a greater weight during aggregation, and finally weighted aggregation is performed on all client local models to obtain a new round of global model.
[0141] The knowledge distillation technique used in Step 3 has the following characteristics: Knowledge distillation (KD) is a technique widely applied to model lightweighting. It compresses the knowledge of a pre-trained complex model (also known as the "teacher network") into a smaller target model (also known as the "student network"). The goal of knowledge distillation is to extract knowledge from a teacher model with better performance and more complex structure to guide the learning process of a lightweight student model, thereby improving the performance of the student model without changing its structure. The principle of knowledge distillation training is as Figure 8 shown. The steps of knowledge distillation include:
[0142] First, train multiple large and performant teacher models. These models usually have high accuracy and strong generalization ability.
[0143] Then, the knowledge of the teacher model can be represented by the probability distribution (soft labels) of its output. These soft labels can provide more information compared to hard labels (i.e., true labels).
[0144] Second, use the soft labels provided by the teacher model to train a small student model. The student model learns how to make decisions while mimicking the teacher model.
[0145] Finally, optimize the knowledge distillation process by adjusting methods such as temperature parameters, enabling the student model to better absorb the knowledge of the teacher model.
[0146] In knowledge distillation, the temperature parameter T is introduced into the Softmax function to adjust the "softness" of the output probability distribution. By increasing the parameter T, the probability distribution can be made smoother and more uniform, making it easier to learn and mimic the behavior of the teacher model. The formula is:
[0147]
[0148] where q i represents the probability of class i, with a value range of [0,1], and Z i represents the logits output by the model in the linear layer, i.e., the raw scores before applying the Softmax function. T is used to adjust these logits to obtain a more sharp or smooth probability distribution. On the one hand, a higher temperature value will increase the smoothness of the probability distribution, resulting in closer prediction results for different classes. On the other hand, a lower temperature value will make the probability distribution more sharp, which may cause the student model to fail to fully learn the useful information of the teacher network, thus affecting its generalization ability.
[0149] The feature of Step 5 is the design of an Adaptive Feature Weighting module AWF and a cross-scale and cross-stage network, which are embedded in the YOLO residual block.
[0150] The Adaptive Feature Weighting module, namely the AWF module: The residual block is a basic component module of the backbone network, and its performance determines the performance of the backbone network to a certain extent. Currently, in the commonly used residual blocks, the identity mapping of the input features is directly added to the output features of the stacked layers. This method treats the information of these two features equally. In fact, there are differences in the semantic levels of these two features, and the spatial information they contain is also different. Therefore, from the perspective of feature fusion analysis, the direct addition method used in the residual block does not fully utilize the information contained in these two features.
[0151] In order to more efficiently utilize the residual blocks in the backbone network and improve the ability of the backbone network to extract features, the Adaptive Weighted Fusion (AWF) module is proposed and the AWF module is embedded in the residual block to achieve the adaptive weighted fusion of features in the residual block, so that the residual block no longer treats the two feature maps equally when fusing features, but is more refined and flexible, strengthening the regions containing effective information while weakening the regions containing invalid information. The structure of the AWF module is as Figure 9 shown, Figure 9 In (a) is the AWF using independent convolution, and (b) is the AWF using shared convolution.
[0152] The AWF module is divided into two structures, which are composed of three parts: compression, extraction, and assignment. The difference between these two structures lies in whether independent convolution or shared convolution is used in the compression stage. The compression part uses 1×1 convolution on the two input feature maps to compress the number of channels to a constant T and refine the feature information (the constant T can be determined manually or adaptively according to the number of input channels). The compressed feature map is used as the intermediate feature map for extracting weight information: The extraction part first concatenates the two intermediate feature maps at the channel level to obtain a combined feature map with 2T channels. This combined feature map contains the important information of the two input feature maps. Then, 1×1 convolution is used to further compress the number of channels of the combined feature map, and at the same time, the spatial weights of the input feature maps are extracted. After compressing the channels, a weight feature map with 2 channels is obtained. The two channels of the weight feature map respectively contain the weight information of the two input feature maps. Since these generated weights are obtained by convolving the same feature map, there is a certain dependence between them, which is used to control the weighted fusion of the two input feature maps. Then, the value of the weight is mapped between 0 and 1 using the following formula.
[0153]
[0154] In the formula, α i,j and β i,jrespectively represent the parameter values at (i, j) in the two channels, and represent the two spatial weights finally obtained; after obtaining the spatial weights, an assignment operation is performed. The obtained spatial weights are matched with the initially input feature maps. By multiplying these weights with the corresponding input feature maps, the dependency relationships between the weights can be passed to the input feature maps, thereby establishing correlations between the input feature maps. The input feature maps that have established correlations with each other are added together to form the final output feature map, thus realizing the adaptive weighted fusion of features. The calculation formula for the output feature map is as follows:
[0155] W = Conv(Concat[Conv(C1), Conv(C2)])
[0156] E = W[0] × C1 + W[1] × C2
[0157] where C1 and C2 represent the input feature maps, W represents the weight feature map, E represents the output feature map, Conv(·) represents the convolution operation, Concat[·] represents the concatenation operation along the channels, W[0] is the first channel of the weight feature map, containing the weight information of the first input feature map, and W[1] is the second channel of the weight feature map, containing the weight information of the second feature map.
[0158] The present invention embeds the AWF module into the residual block of the YOLO backbone network, as Figure 10 shown, Figure 10 where (a) in it is the original residual block and (b) is the improved residual block. When the improved residual block outputs features, it inputs the output of the stacked layer and the identity mapping of the input features into the AWF module together instead of directly adding them. In the AWF module, weights are extracted and assigned to the two feature maps to establish the correlations between them. Compared with the original residual block, the feature information extracted by the improved residual block is more beneficial to subsequent tasks.
[0159] Cross-scale and cross-stage network: The GiraeDet network proposed in 2022 uses a mechanism that combines skip connections and cross-scale connections in its neck, enabling the model to fully perform high-level and low-level information interaction. The skip connections it uses are the mutual connections of features at the same scale in different propagation stages. The feature fusion network used by the benchmark model YOLO of the present invention is based on FPN and PANet and only performs cross-scale feature fusion. To address this problem, a new fusion method, imPANet, is designed. Cross-stage fusion is added to the cross-scale fusion structure, enabling the features to supplement the original information of this feature layer after obtaining the information of other scale feature layers, as Figure 11 shown.
[0160] In the algorithm of the present invention, the cross-scale fusion path retains the original method of YOLO, which first samples, then cascades, and then convolves. The cross-stage fusion path follows the way of skip connection. This way has a shorter distance in backpropagation, and the feature dimension does not change. The features extracted by the backbone network can be directly added to the same-level features after feature fusion. In addition, the way of skip connection does not introduce additional parameters, and the impact on the computational overhead can be ignored. See Figure 12 is the specific structure of the improved feature fusion network imPANet. The improved structure performs downsampling after cross-stage connection, so that after the low-level features are fused with the original information input to the neck network, they are then propagated as a whole to the high-level features.
[0161] YOLO-AWF object detection algorithm: The overall structure of YOLO-AWF is as Figure 13 shown, and it is mainly composed of three parts: a backbone network, a neck network, and a prediction network. YOLO-AWF enhances the feature extraction ability and feature fusion ability of the network respectively on the basis of the YOLO algorithm. Among them, an adaptive weighted fusion module is added to the backbone network to improve the learning ability of the residual block, and the cross-scale and cross-stage fusion network proposed by the present invention is used in the neck to make up for the information loss in the feature fusion process, so that the neck network can make more full use of the feature information.
[0162] When YOLO-AWF makes predictions, first, an image with a size of 416×416 is input into the backbone network for feature extraction. The backbone network is composed of 5 large residual blocks, and some large residual blocks are embedded with an adaptive weighted fusion module. Each large residual block contains 2 branches. One branch stacks residual blocks, and the other branch serves as a residual edge spanning the entire large residual block. The number of residual blocks contained in each large residual block is 1, 2, 8, 8, 4 respectively. Among them, the features output by the 3rd, 4th, and 5th large residual blocks will be sent to the neck network for feature fusion. The dimensions of these 3 features are 52×52×256, 26×26×512, and 13×13×1024 respectively. Among them, the feature with a size of 13x13 will first be sent into the SPP structure for maximum pooling with 4 different strides to expand the receptive field of the feature. The neck network then fuses the 3 features of different scales input. First, the rich semantic information contained in the high-level features is propagated layer by layer to the low-level features through the top-down path, and then the rich spatial information contained in the low-level features is propagated layer by layer to the high-level features through the bottom-up path. The dimension of the fused features is the same as the feature dimension when input to the neck network. Then, the cross-stage fusion path is used to add the features before and after fusion at the same scale. Finally, the 3 features output by the neck network are used as effective features and input into the prediction network to identify and regress the category and position of the target in the image. The prediction network used is the prediction network of YOLO.
[0163] The main feature of step 7 is to propose an active learning method based on pseudo-labels from the perspectives of sample uncertainty and diversity, aiming at the characteristics of the object detection algorithm. The multi-scale features are obtained by using the feature extraction network of the object detection model, and the multi-scale features are fused as the feature representation of a single unlabeled image. At the same time, a clustering method is used to cluster the fused image features to provide a pseudo-label as the classification result of the image. On this basis, the trained object detection model is used to identify the unlabeled image, and the information entropy of the classification confidence is calculated as the uncertainty of each image. Combining the pseudo-labels and uncertainties, the uncertainties of the unlabeled images are sorted according to their magnitudes under different pseudo-label classifications. The total number of annotation budgets for each round of active learning is B, the number of pseudo-label classifications is k, and b unlabeled images are selected from each pseudo-label classification and added to the annotation set, and b×k <= B. In this way, the uncertainty and representativeness of the samples are taken into account as much as possible, so that the selected sample set can fit the distribution of the real data set.
[0164] The specific algorithm process of the active learning in step 7 is as Figure 14 shown. The unlabeled samples extract features through the trained object detection model to output the uncertainty of each unlabeled sample, as well as the pseudo-labels formed by clustering the image features output by the feature extraction network. The pseudo-labels and uncertainties jointly determine whether the unlabeled samples are worthy of being labeled. In each round of active learning, the labeled samples will update the training set, and the object detection model will be retrained before the next round of active learning, and then the above operations will continue until the annotation budget is exhausted or the performance of the object detection model meets the requirements.
[0165] In order to make the classification of the samples selected by active learning as balanced as possible, the K-Means clustering method is used to pre-classify the unlabeled samples by applying pseudo-labels, that is, the unlabeled data is feature-extracted through the trained object detection network model, and then clustered according to the features, so that each sample obtains an initial classification. Since there are large and small objects in the object detection images, the present invention takes into account the small objects corresponding to the large-scale feature maps and the large objects corresponding to the small-scale feature maps in the feature extraction stage. Each feature map is first interpolated and then fused to ensure that the features at each level can be considered as much as possible. Also, since the classification details in the data set cannot be clearly known in advance, only the number of classification targets in the data set can be roughly understood as N. The present invention sets the number of pre-classifications as M, and M > N, and over-clustering is used to find the differences between the samples.
[0166] The uncertainty is calculated using a calculation method based on information entropy. The characteristic of this method is that it can measure the determination results of all classes of a model for a certain sample x in the case of multi-classification. When the entropy is larger, it indicates that the prediction of the model is more unstable, that is, the uncertainty of this sample is higher. When the entropy is smaller, it indicates that the prediction of the model is more stable, that is, the uncertainty of this sample is lower. The information formula is as follows:
[0167]
[0168] Among them, θ represents the parameters of a trained deep learning model, and P θ (y i |x) represents the probability value of each classification of the sample x, represents the information entropy of each sample x.
[0169] To better utilize the advantages of information entropy, the information entropy is calculated for each pixel position of each intermediate feature map output by the YOLO model. The information entropy of the entire image is the sum of the information entropies of all positions, which is used as the uncertainty of the image. Then, according to the annotation budget B and the number of pre-classifications M, the number of samples b to be sampled for each classification is calculated as If B and M cannot be divided evenly, the remaining sampling quantity is then allocated to all classifications until the annotation budget B is satisfied or the unannotated dataset is exhausted.
[0170] In summary, the system of the present invention adopts a distributed architecture, consisting of a central server and multiple edge nodes, and the core components include a federated learning module, a knowledge distillation module, an adaptive feature fusion module, and an active learning module. The system operation begins with each participant performing data preprocessing and initial model training locally. Then, through the federated learning framework, the participants regularly send model updates to the central server securely. The server uses a secure aggregation algorithm to merge these updates and distributes the optimized global model back to each party. Next, the system trains multiple complex "teacher" models based on the global model to capture detailed distribution line component features, and then uses knowledge distillation technology to effectively compress its knowledge into a lightweight "student" model. This distilled lightweight model is then deployed to each edge device, enabling high-performance detection to be achieved on devices with limited computing resources. In actual operation, the edge device dynamically adjusts the feature extraction and fusion strategies according to the current environmental conditions. This adaptive feature enables the system to cope with complex and changeable distribution line environments. The edge device uses the optimized model to perform real-time detection of distribution line components to ensure rapid response and timely decision-making. In order to continuously improve system performance, an active learning mechanism is introduced. The system can identify difficult samples, select the most valuable samples for manual labeling through uncertainty estimation and sample screening algorithms, and then use these newly labeled samples to update the model, forming a closed loop of continuous learning and optimization. The system also regularly collects feedback from edge devices, updates the global model, and repeats the entire process to ensure that the model always remains advanced and adaptable.
[0171] Compared with the prior art, the present invention has the following advantages:
[0172] (1) Protecting data privacy: The present invention uses an improved federated averaging algorithm (MFedAvg) for federated learning. Each participant only needs to share model updates instead of original data, which effectively protects sensitive information. This method is particularly suitable for the power distribution industry because different power companies or regions may have their own data protection requirements. Through federated learning, all parties can jointly train the model without directly sharing the original data, greatly reducing the risk of data leakage.
[0173] (2) Improve model performance: The present invention uses a federated learning framework to fully utilize distributed data resources for training, significantly improving the generalization ability and accuracy of the model. In particular, by aggregating updates from all parties through a weighted average strategy, the weights are dynamically adjusted according to the quality and quantity of the data, ensuring that high-quality data contributes more to the model, thereby improving the overall model performance.
[0174] (3) Implement edge intelligence: Through the knowledge distillation module with a multi-teacher distillation strategy, the present invention compresses a complex model into a lightweight model. This method not only transfers soft labels but also includes the matching of feature maps, ensuring that the student model can capture richer knowledge. This enables high-performance detection to be achieved on edge devices with limited computing resources, greatly improving the practicality and deployment flexibility of the system.
[0175] (4) Strong adaptability: The adaptive feature fusion module of the present invention includes an adaptive feature weighting (AWF) module and a cross-scale cross-stage network, which can automatically adjust the feature extraction strategy according to factors such as the quality of the input image and lighting conditions. This design enables the system to adapt to complex and variable distribution line environments, improving the accuracy and robustness of detection.
[0176] (5) Continuous optimization: The present invention introduces an active learning method based on pseudo-labels. The system can identify difficult-to-process samples and, through uncertainty estimation and sample screening algorithms, select the most valuable samples for manual annotation. This mechanism ensures that the system can continuously learn new patterns and features, continuously improving the detection performance.
[0177] It should be understood that the present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0178] The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0179] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0180] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Smalltalk, C++, Python, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.
[0181] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0182] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0183] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0184] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As is well known to those skilled in the art, implementations by hardware, by software, and by a combination of software and hardware are equivalent.
[0185] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.
Claims
1. An adaptive distribution line component detection system based on federated learning and knowledge distillation, characterized in that, Including: A central server and edge devices, where the federated learning module, knowledge distillation module, adaptive feature fusion module, and active learning module are deployed on the central server and the edge devices; The federated learning module is used for each participant on the edge device to train an object detection model and regularly send the updated object detection model to the central server. The central server uses a secure aggregation algorithm to merge and update the object detection model to obtain a global model, and distributes the global model back to each participant; The knowledge distillation module is used to train a complex teacher model based on the global model and compress the knowledge of the teacher model into a lightweight student model; The adaptive feature fusion module is used for the edge device to dynamically adjust the feature extraction and fusion strategy based on the lightweight student model and the current environmental conditions to perform real-time detection of distribution line components; The active learning module is used for the adaptive distribution line component detection system to select the most valuable samples through uncertainty estimation and sample screening algorithms, request manual annotation as new samples, and use the new samples for updating the object detection model.
2. The system according to claim 1, wherein The detection method of the adaptive distribution line component detection system includes the following steps: Step 1. Initialization: Each participant performs data preprocessing and initial model training locally; Step 2. Federated learning: Each participant regularly sends the updated object detection model to the central server. The server uses a secure aggregation algorithm to merge and update to obtain a global model, and distributes the global model back to each participant; Step 3. Knowledge distillation: Train a complex teacher model based on the global model, and then compress the knowledge of the complex teacher model into a lightweight student model; Step 4. Edge deployment: Deploy the distilled lightweight student model to the edge device; Step 5. Adaptive feature fusion: The edge device dynamically adjusts the feature extraction and fusion strategy according to the current environmental conditions; Step 6. Real-time detection: The edge device uses the distilled lightweight student model to perform real-time detection of distribution line components; Step 7. Active learning: The adaptive distribution line component detection system identifies samples that are difficult to process, requests manual annotation as new samples, and uses the new samples for updating the object detection model; Step 8. Continuous optimization: Regularly collect feedback from the edge device, update the global model, and repeat steps 2 to 8 to form a closed-loop optimization; 3. The system according to claim 2, wherein The federated learning module adopts the MFedAvg algorithm. The MFedAvg algorithm aggregates edge nodes by model distance screening, excludes abnormal nodes, and guarantees the process of federated learning, specifically including the client local training, central server model aggregation, and node status detection links.
4. The system according to claim 2, wherein The knowledge distillation module adopts a multi-teacher distillation strategy, including the following steps: Train multiple complex teacher models; Represent the knowledge of the teacher model through the soft labels output by the teacher model; Use the soft labels provided by the teacher model to train a small student model. The student model learns how to make decisions in the process of mimicking the teacher model. Optimize the knowledge distillation process by adjusting the temperature parameter, enabling the student model to better absorb the knowledge of the teacher model.
5. The system according to claim 4, characterized in that, The process of optimizing knowledge distillation by adjusting the temperature parameter has the following formula: where q i represents the probability of class i, with a value range of [0, 1], and Z i represents the logits output by the model in the linear layer, i.e., the raw scores before applying the Softmax function, and T represents the temperature parameter used to adjust the logits.
6. The system according to claim 1, wherein The adaptive feature fusion module includes an AWF module and a cross-scale cross-stage network. Embed the AWF module and the cross-scale cross-stage network into the residual blocks of the YOLO backbone network, and automatically adjust the feature extraction strategy according to the input image quality and lighting conditions to adaptively weight features at different levels and scales.
7. The system according to claim 1, wherein The AWF module consists of three parts: compression, extraction, and assignment. In the compression part, use 1×1 convolution on the two input feature maps to compress the number of channels to a constant T and refine the feature information. The compressed feature map is used as an intermediate feature map for extracting weight information. In the extraction part, first concatenate the two intermediate feature maps at the channel level to obtain a combined feature map with 2T channels. Then use 1×1 convolution to further compress the number of channels of the combined feature map, and at the same time extract the spatial weights of the input feature maps. After compressing the channels, a weight feature map with 2 channels is obtained. The two channels of the weight feature map respectively contain the weight information of the two input feature maps. Since these generated weights are obtained by convolving the same feature map, there is a dependency relationship between them, which is used to control the weighted fusion of the two input feature maps. Then use the following formula to map the values of the weights between 0 and 1: where α i,j and β i,j represent the parameter values at (i, j) in the two channels respectively, and represent the two finally obtained spatial weights; After obtaining the spatial weights, perform the assignment operation, match the obtained spatial weights with the initially input feature maps. By multiplying these weights with the corresponding input feature maps, the dependency relationship between the weights can be passed to the input feature maps, thereby establishing a correlation between the input feature maps. Add these input feature maps with established correlations to form the final output feature map, thus realizing the adaptive weighted fusion of features.
8. The system according to claim 7, wherein The calculation formula of the output feature map is as follows: W = Conv(Concat[Conv(C1), Conv(C2)]) E = W[0] × C1 + W[1] × C2 Where, C1 and C2 represent the input feature maps, W represents the weight feature map, E represents the output feature map, Conv(·) represents the convolution operation, Concat[·] represents the concatenation operation along the channel, W[0] is the first channel of the weight feature map, containing the weight information of the first input feature map, and W[1] is the second channel of the weight feature map, containing the weight information of the second feature map.
9. The system according to claim 1, characterized in that, The active learning module uses the feature extraction network of the object detection model to obtain multi-scale features, fuses the multi-scale features as the feature representation of a single unlabeled image, provides pseudo-labels through the K-Means clustering method, combines uncertainty estimation, sorts the uncertainties of the unlabeled images according to size under different pseudo-label classifications, selects samples for annotation according to the annotation budget, updates the training set, and retrains the object detection model.
10. The system according to claim 9, wherein The active learning module uses a calculation method based on information entropy for uncertainty estimation of samples, and the formula is: Among them, θ represents the parameters of a trained deep learning model, and P θ (y i |x) represents the probability value of each classification of the sample x, represents the information entropy of each sample x.