Electric power AR emergency image recognition method based on end-side complementation and edge collaboration

By adopting the power AR emergency image recognition method based on end-side complementarity and edge coordination in power communication emergency scenarios, using the EdgeYOLO and CloudYOLO models, the traditional methods have been solved in real-time and accuracy, and efficient and accurate image recognition is achieved.

CN119942389APending Publication Date: 2025-05-06国网宁夏电力有限公司信息通信公司 +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411975865.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In complex power communication emergency scenarios, traditional image recognition methods are difficult to meet real-time and accuracy requirements, especially when network bandwidth is limited and resources are limited.

Method used

Using a power AR emergency image recognition method based on end-side complementarity and edge collaboration, by building an identification platform including wearable AR devices, edge servers and cloud servers, the EdgeYOLO model is used to perform real-time image recognition on edge servers, and in-depth calculation and model optimization are performed in the cloud through the CloudYOLO model.

Benefits of technology

It improves the real-time and accuracy of image recognition in complex emergency scenarios, reduces the data transmission volume and calculation complexity, and enhances the generalization ability and overall performance of edge models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942389A_ABST
    Figure CN119942389A_ABST
Patent Text Reader

Abstract

The invention provides an electric power AR emergency image recognition method based on end-side complementation and edge collaboration, and relates to the technical field of image recognition, and the method comprises the steps: constructing an electric power AR emergency image recognition platform; the platform comprises a wearable AR device, an edge server and a cloud server. The wearable AR device performs multi-source fusion on the collected historical image data in the electric power emergency scene; the edge server deploys an EdgeYOLO model and is used for performing frequency domain fusion and optimization on the image features after multi-source fusion based on Fourier federal learning; the cloud server deploys a Cloud YOLO model which is used for depth calculation and model optimization; the Backbone part of the Cloud YOLO model is deployed on the edge server as a part of the EdgeYOLO model, and is used for assisting the EdgeYOLO to carry out federated learning; the wearable AR device collects image data in the power emergency scene in real time; and operating the EdgeYOLO model to identify the real-time image data, and outputting an identification result. According to the scheme, the real-time performance and the accuracy of image recognition in a complex emergency scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to an electric power AR emergency image recognition method based on end-side complementarity and edge collaboration. Background Art

[0002] With the rapid development of information technology, augmented reality (AR) wearable devices have shown significant application value in power communication emergency scenarios. The power emergency environment is usually complex and changeable. There may be areas with limited signals, and the equipment is required to respond quickly under limited resources to help workers efficiently identify target objects and accurately perceive the surrounding environment in emergencies. However, traditional image recognition methods rely on high-performance computing and stable networks, and it is difficult to meet the real-time and accuracy requirements in emergency scenarios. Especially in power emergency environments, limited network bandwidth and complex environmental characteristics further increase the task load and data transmission pressure of AR devices.

[0003] At present, deep learning-based object detection algorithms have become mainstream due to their advantages in accuracy and speed, but such models are computationally complex and resource-intensive, and it is still challenging to deploy them directly on AR devices. Although the traditional cloud computing model has powerful computing power, the process of transmitting data to a remote end, processing it, and returning it will cause delays in emergency scenarios, making it difficult to meet real-time response requirements and posing a potential risk of communication interruption.

[0004] Although cloud-edge-end integrated collaborative computing provides new ideas for AR device recognition enhancement in emergency scenarios, current methods still have shortcomings in practical applications. For example, the limited computing power of the end device leads to limited accuracy and real-time performance of preliminary detection. Edge nodes are difficult to fully utilize the complementary information of multi-source data under high load conditions, and deep computing in the cloud relies on the stability of the network and is susceptible to network fluctuations. These problems are particularly prominent in complex power communication emergency scenarios, making it difficult to meet the actual needs of rapid response and high-precision recognition. Summary of the invention

[0005] In view of this, in order to address the above shortcomings, it is necessary to propose an electric power AR emergency image recognition method based on end-side complementarity and edge collaboration to improve the real-time and accuracy of image recognition in complex emergency scenarios.

[0006] In a first aspect, the present invention provides a power AR emergency image recognition method based on end-side complementarity and edge collaboration, comprising:

[0007] Construct an electric power AR emergency image recognition platform; wherein the recognition platform is deployed with an image recognition model, and the recognition platform includes: a wearable AR device, an edge server and a cloud server; each cloud server is correspondingly provided with a number of edge servers, and each edge server is correspondingly provided with a number of wearable AR devices; when constructing the recognition platform, the wearable AR device is used to perform multi-source fusion of historical image data collected in electric power emergency scenes; the edge server deploys an EdgeYOLO model, which is used to perform frequency domain fusion and optimization of image features after multi-source fusion based on Fourier federated learning; the cloud server deploys a CloudYOLO model, which is used to perform deep calculation and model optimization; and the Backbone part of the CloudYOLO model is deployed on the edge server as a part of the EdgeYOLO model, which is used to assist EdgeYOLO in federated learning;

[0008] Real-time image data in power emergency scenarios can be collected through wearable AR devices;

[0009] The edge server runs the EdgeYOLO model to recognize the real-time image data and outputs the recognition result.

[0010] Preferably, the wearable AR device is used to collect image data in the power scene through a camera and a sensor, and perform image denoising, image edge extraction and edge image fusion on the image data; wherein the image data includes device status, location information and environmental characteristics.

[0011] Preferably, the image denoising process includes:

[0012] The original image F collected by the wearable AR device is processed by at least two mathematical morphological filters of different sizes and shapes to obtain a denoised image F. y ;

[0013] The process of image edge extraction includes:

[0014] Using structural elements C in different directions ni , use the following calculation formula to calculate the denoised image F y Edge features are extracted in different directions to obtain edge images in each direction:

[0015]

[0016] In the formula, C ni represents the mathematical morphological filter corresponding to the nth size and shape in the i-th direction, E ni Characterizes the directional edge image corresponding to the nth size and shape in the i-th direction, is used to represent matrix addition, and Θ is used to represent matrix subtraction;

[0017] The image fusion process includes:

[0018] Based on the following calculation formula, the directional edge images obtained in each direction are weighted and summed using the weight coefficient to obtain a complete edge image:

[0019]

[0020] In the formula, E n (F) Edge image corresponding to the mathematical morphological filter used to characterize the nth size and shape, q i Used to represent the weight coefficient corresponding to the i-th direction.

[0021] Preferably, the processing of the original image by using mathematical morphological filters of at least two sizes and shapes comprises:

[0022] The mathematical morphological filters B1 and C1 are connected in series in descending order of size, and a morphological operation is performed on the original image F to obtain a denoised image F1;

[0023] The mathematical morphological filters B2 and C2 are connected in series in descending order of size, and a morphological operation is performed on the original image F to obtain a denoised image F2;

[0024] Using the following calculation formula, images F1 and F2 are connected in parallel according to the weight coefficient related to the peak signal-to-noise ratio to obtain the denoised image F y :

[0025] F y =q1F1+q2F2

[0026] in,

[0027] q1, q2∈[0,1], and q1+q2=1.

[0028] Preferably, the frequency domain fusion and optimization of the image features after multi-source fusion based on Fourier federated learning includes:

[0029] The convolutional model θ k The convolutional layer parameters Convert to a two-dimensional matrix Where O and C are the number of output channels and input channels respectively, s1 and s2 are the spatial shapes of the convolution kernel;

[0030] For the transformed two-dimensional matrix w′ k Perform fast Fourier transform to obtain the amplitude graph F A and phase diagram FP :

[0031] The kth local model is aggregated in the frequency domain space through the low-frequency mask, and we get

[0032] Based on the following calculation formula, the amplitude is mapped using the inverse Fourier transform and the phase map F P Convert to parameter form

[0033]

[0034] In the formula, F -1 Used to characterize the inverse Fourier transform.

[0035] Preferably, the two-dimensional matrix w′ is calculated using the following formula: k Perform a fast Fourier transform:

[0036]

[0037] In the formula, m and n are given parameters;

[0038] Based on the following calculation formula, the kth local model is aggregated in the frequency domain space through the low-frequency mask:

[0039]

[0040] Where G is the low-frequency mask, Z is the indicator function, and g∈(0, 0.5) represents the low-frequency threshold.

[0041] Preferably, the initial stage of deploying the EdgeYOLO model on the edge server includes:

[0042] Use historical image data from power emergency scenarios to train CloudYOLO on a cloud server;

[0043] Extract the first m feature layers corresponding to CloudYOLO's Backbone to assist EdgeYOLO training;

[0044] With the help of CloudYOLO, using W mc Replace the corresponding Backbone parameters in EdgeYOLO; where W c is the trained CloudYOLO parameter, W mc The parameters of its Backbone part, W e are the parameters of EdgeYOLO;

[0045] Training parameter W e-mc ; Among them, W e-mc For characterization of W removalmc The remaining parameters of EdgeYOLO;

[0046] When the edge server receives the parameters W of the first m feature layers mc After that, n higher layers are randomly initialized and combined with W mc Connect to form EdgeYOLO;

[0047] Freeze parameter W mc , and use local data to e-mc Train and get the final EdgeYOLO model.

[0048] Preferably, for a given training data The loss function used when training EdgeYOLO is expressed as follows:

[0049]

[0050] In the formula, f e (W e-mc ) is the total loss function used when training EdgeYOLO, H is used to represent the loss function of training to the i-th picture, σ is used to represent the activation function, σ(f(x i ; W e-mc ) is the calculation function of the model in the forward propagation process, which is used to calculate the input x i and the current parameter W e-mc Generate output.

[0051] Preferably, after outputting the recognition result each time, the method further includes:

[0052] The edge server uploads real-time image data and recognition results to the cloud server;

[0053] The cloud server accumulates the uploaded data and forms sample training data for updating the model based on the real-time image data and the recognition result;

[0054] According to the preset period, or when the accumulated sample training data reaches the preset value, the cloud server sends the sample training data for updating to the edge server;

[0055] The edge server updates EdgeYOLO using the sample training data sent by the cloud server.

[0056] In a second aspect, the present invention further provides a computing device, including a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, any method described in the first aspect is implemented.

[0057] It can be seen from the above technical scheme that in the electric power AR emergency image recognition method based on end-side complementarity and edge system provided by the present invention, firstly, an electric power AR emergency image recognition platform is constructed, and then the wearable AR device of the electric power AR emergency image recognition platform collects real-time image data in the electric power emergency scene in real time, and then the edge server runs the EdgeYOLO model to recognize the real-time image data to obtain the recognition result. The constructed electric power AR emergency image recognition platform includes a wearable AR device, an edge server and a cloud server. The wearable AR device can perform multi-source fusion of the collected historical image data in the electric power emergency scene, and the edge server deploys the EdgeYOLO model, which can perform frequency domain fusion and optimization of the image features after multi-source fusion based on Fourier federated learning, and the cloud server deploys the CloudYOLO model, which can perform deep calculation and model optimization. Moreover, the Backbone part of the CloudYOLO model is deployed on the edge server as part of the EdgeYOLO model, which can assist EdgeYOLO in federated learning. It can be seen that the wearable AR device on the end side of this solution can perform preliminary data processing, information fusion and multivariate data complementation, integrate the advantageous information of multi-source images, strengthen the integrity of device feature expression and environmental information, and provide efficient data support for complex emergency scenarios. In addition, the EdgeYOLO model is deployed on the edge side, which can aggregate features in the frequency domain, share global knowledge at low frequencies, and optimize personalized features at high frequencies, thereby improving recognition accuracy and collaborative efficiency. The EdgeYOLO model completes feature fusion locally, which also reduces the amount of uploaded data and the risk of privacy leakage, providing reliable support for efficient response and accurate decision-making in complex environments. In addition, by deploying the Backbone part of the CloudYOLO model on the edge server as part of the EdgeYOLO model to assist EdgeYOLO training, the generalization ability and overall performance of the edge model can be significantly improved, avoiding the high computational cost of training the model from scratch, which can not only save computing resources, but also improve recognition accuracy when data is limited, thereby better adapting to the needs of complex emergency environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 A flowchart of a power AR emergency image recognition method based on end-side complementarity and edge collaboration is provided in an embodiment of the present invention.

[0059] Figure 2 A schematic diagram of the architecture of an electric power AR emergency image recognition platform provided in an embodiment of the present invention.

[0060] Figure 3 Schematic diagram of the detection accuracy of EdgeYOLO-r and EdgeYOLO-s with different numbers of training images on the COCO dataset.

[0061] Figure 4 The following is a comparison chart of response time and upload time for two different deployment solutions. DETAILED DESCRIPTION

[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0063] like Figure 1 As shown, the present invention provides a power AR emergency image recognition method based on end-side complementarity and edge collaboration, which may include the following steps:

[0064] Step 101: constructing an electric power AR emergency image recognition platform; wherein the recognition platform is deployed with an image recognition model, and the recognition platform includes: a wearable AR device, an edge server and a cloud server; each cloud server is correspondingly provided with a number of edge servers, and each edge server is correspondingly provided with a number of wearable AR devices; when constructing the recognition platform, the wearable AR device is used to perform multi-source fusion of historical image data collected in the electric power emergency scene; the edge server deploys an EdgeYOLO model, which is used to perform frequency domain fusion and optimization of image features after multi-source fusion based on Fourier federated learning; the cloud server deploys a CloudYOLO model, which is used to perform deep calculation and model optimization; and the Backbone part of the CloudYOLO model is deployed on the edge server as a part of the EdgeYOLO model, which is used to assist EdgeYOLO in performing federated learning;

[0065] Step 102: Collect real-time image data in the power emergency scene through a wearable AR device in real time;

[0066] Step 103: The edge server runs the EdgeYOLO model to recognize the real-time image data and outputs the recognition result.

[0067] In this embodiment, an integrated cloud-edge-end (cloud server, edge server, AR device end) architecture is adopted. The end side uses a wearable AR device to collect environmental data using multimodality, and combines multi-scale decomposition and reconstruction technology to complete image preprocessing and multi-source data complementation; the EdgeYOLO model is deployed on the edge side, and the image features are fused and optimized in the frequency domain based on the Fourier federated learning method, low-frequency global knowledge is shared, and high-frequency individual features are optimized, and dynamic collaboration between multiple devices is supported, and recognition accuracy and reasoning speed are improved through dynamic collaboration; the cloud side undertakes deep computing and global model optimization through the CloudYOLO model, providing computing power support and task expansion capabilities for the edge side. In this way, through the collaborative work of the cloud, edge and end, the real-time and accuracy of image recognition in complex emergency scenarios are significantly improved, providing efficient and reliable technical guarantees for power emergency.

[0068] like Figure 2 As shown in the figure, the constructed power AR emergency image recognition platform consists of three layers, namely wearable AR devices, edge servers and cloud servers. It is mainly used to achieve the following:

[0069] 1. Multi-source image fusion method for wearable AR devices on the edge

[0070] Wearable AR devices collect image data in power emergency scenarios through built-in cameras and sensors, such as equipment status, location information, and environmental features. As end-side devices, they are responsible for preliminary data processing, information fusion, and multi-dimensional data complementation. In order to address the viewing angle limitations and noise sensitivity of single-source data, a multi-scale multi-source image fusion technology based on mathematical morphology is used to reduce noise through structural element filters of different sizes and shapes, retain multi-directional edge features, and perform registration and weighted fusion of multimodal data. This method integrates the advantageous information of multi-source images, strengthens the expression of equipment features and the integrity of environmental information, and provides efficient data support for complex emergency scenarios.

[0071] In power emergency scenarios, AR wearable devices, as the core tool for end-side data collection, can collect multimodal image data in real time in complex environments. However, due to problems such as insufficient light, noise interference, and limited viewing angle in the collection environment, the collected images often need to be processed in multiple steps such as denoising, enhancement, and fusion. To this end, a multi-scale and multi-source image fusion method based on mathematical morphology is proposed to improve the quality and processing efficiency of images collected by AR wearable devices.

[0072] The original images collected by AR wearable devices may contain high noise components. Using mathematical morphology filters for noise reduction is an important means to enhance image quality. In combination with the multimodal characteristics of AR devices, it is possible to consider using two structural elements of different sizes and shapes to denoise the image, specifically 3×3 structural elements and 5×5 structural elements. Mathematical morphology methods are used to obtain image edge features in order to obtain all image edges. The logical combination of various structural elements is used so that the structural elements can include all the filter windows of the image, and the edges of the image in different directions can be detected, and then the edges in different directions are fused to obtain the complete edge of the image. The specific steps may include:

[0073] (1) Image denoising

[0074] The original image F collected by the wearable AR device is subjected to denoising through a multi-scale, multi-structural element mathematical morphological filter to obtain a denoised image F. y .

[0075] In this embodiment, in order to minimize the impact of noise on image processing, the present invention performs noise reduction processing by designing filters of different structural elements, and optimizes the image by a combination of series connection followed by parallel connection. Among them, the weight coefficient of the parallel processing is calculated and determined by the peak signal-to-noise ratio of the image. In structural elements of uniform shape, such as B1 and C1, B2 and C2, a series filter is constructed in order from large to small in size to realize morphological operations, and the denoised images F1 and F2 are obtained respectively. Then, F1 and image F2 are connected in parallel, and the respective weight coefficients q1 and q2 are determined based on the peak signal-to-noise ratio, and then the image F is obtained by weighted summation. y Specifically, when the original image is processed by mathematical morphological filters of at least two sizes and shapes, it can be achieved in the following manner:

[0076] The mathematical morphological filters B1 and C1 are connected in series in descending order of size, and a morphological operation is performed on the original image F to obtain a denoised image F1;

[0077] The mathematical morphological filters B2 and C2 are connected in series in descending order of size, and a morphological operation is performed on the original image F to obtain a denoised image F2;

[0078] Using the following calculation formula, images F1 and F2 are connected in parallel according to the weight coefficient related to the peak signal-to-noise ratio to obtain the denoised image F y :

[0079] F y =q1F1+q2F2

[0080] in,

[0081] q1, q2∈[0,1], and q1+q2=1.

[0082] (2) Image edge extraction

[0083] According to the edge features of power equipment in different directions, structural elements C in different directions are used ni Perform edge extraction to obtain edge images in all directions. The edge extraction formula can be expressed as:

[0084]

[0085] In the formula, C ni represents the mathematical morphological filter corresponding to the nth size and shape in the i-th direction, E ni Characterizes the directional edge image corresponding to the nth size and shape in the i-th direction, is used to represent matrix addition, and Θ is used to represent matrix subtraction.

[0086] In this embodiment, in the aspect of image edge feature extraction, a mathematical morphological method is used to design a multi-scale and multi-directional structural element combination so that the structural element covers the entire filter window of the image, thereby being able to detect multi-directional edge features. Based on the multi-scale morphological filter, by decomposing the 3×3 square structural element Get 8 structural elements in different directions:

[0087] And use these structural elements to reconstruct the complete edge information. According to the relationship between the structural elements, we can get: B3 = B 31 ∪B 32 ∪B 33 ∪B 34 ∪B 35 ∪B 36 ∪B 37 ∪B 38 .

[0088] Similarly, based on a 5×5 square structural element Decompose to get 8-directional structural elements:

[0089] In this way, the edge features of the image can be completely extracted. According to the relationship between the structural element decomposition, we can get: C3 = C 31 ∪C 32 ∪C 33 ∪C 34 ∪C 35∪C 36 ∪C 37 ∪C 38 .

[0090] (3) Image Fusion

[0091] The directional edge images obtained in each direction are weighted and summed using the weight coefficient to obtain a complete edge image. The formula for weighted summation is as follows:

[0092]

[0093] In the formula, E n (F) Edge image corresponding to the mathematical morphological filter used to characterize the nth size and shape, q i It is used to represent the weight coefficient corresponding to the i-th direction. After the image is filtered, the noise can be ignored and the weight coefficients in different directions can be calculated as the same value.

[0094] In power emergency scenarios, the data collected by wearable AR devices is directional. To this end, this embodiment combines the direction weight adjustment strategy in the edge fusion process to enhance the edge information in a specific direction, further improving the accuracy of anomaly detection.

[0095] 2. Edge aggregation method based on Fourier federated learning

[0096] When there are multiple edge nodes in the power emergency scenario, a federated learning strategy is needed to train the model on the local data on different edge nodes to achieve efficient aggregation and update of model parameters. In the power emergency environment, the task scenarios are usually complex and changeable, including equipment damage caused by natural disasters, fault location in communication interruption areas, and real-time status monitoring of distributed power equipment. These tasks require the recognition system to respond quickly and provide personalized support according to the actual needs of different nodes.

[0097] To meet these requirements, the edge nodes in EdgeYOLO upload local model parameters to CloudYOLO. After receiving these parameters, CloudYOLO uses fast Fourier transform to convert the parameters to frequency domain space to optimize aggregation efficiency. The low-frequency part of the model parameters after fast Fourier transform represents basic knowledge, which is suitable for sharing common information across nodes, while the high-frequency part retains the personalized characteristics of each node in a specific power emergency scenario.

[0098] Specifically, firstly, the convolution model θ k The convolutional layer parameters Convert to a two-dimensional matrix Where O and C are the number of output channels and input channels respectively, s1 and s2 are the spatial shapes of the convolution kernel;

[0099] Then, the two-dimensional matrix w′ is calculated using the following formula k Perform fast Fourier transform to obtain the amplitude graph F A and phase diagram F P :

[0100]

[0101] In the formula, m and n are given parameters;

[0102] Furthermore, in order to extract the low-frequency components for aggregation, consider using a low-frequency mask G, the value of which is 0 except for the central area. The average low-frequency component is passed, the high-frequency component is retained, and the k-th local model is aggregated in the frequency domain space to obtain The specific expressions are as follows:

[0103]

[0104] Where G is the low-frequency mask, Z is the indicator function, and g∈(0, 0.5) represents the low-frequency threshold.

[0105] Finally, the amplitude is mapped to and the phase map F P Convert to parameter form

[0106]

[0107] In the formula, F -1 Used to characterize the inverse Fourier transform.

[0108] According to the above process, CloudYOLO customizes a corresponding personalized local model for each EdgeYOLO. The edge node then combines the personalized local model with the data collected by the AR devices within its coverage area for edge-coordinated local federated learning, and finally obtains a personalized local model.

[0109] 3. Image recognition algorithm for cloud-edge-device collaboration

[0110] The goal of the present invention is to use edge servers to complete all recognition operations, thereby providing on-site recognition services in emergency environments. This means that after capturing data, the wearable AR device directly uploads the data to the nearest edge server without any preprocessing. This design can effectively reduce the computational burden of wearable devices and extend their battery life. In addition, since the physical distance between the end side and the edge side server is relatively close, the data transmission delay is small, thereby improving the overall response speed and meeting the real-time requirements in emergency scenarios.

[0111] However, due to the limited amount of data collected by a single edge server, directly using EdgeYOLO may lead to insufficient recognition accuracy and prone to overfitting problems during training. To solve this problem, this solution considers using cloud servers to assist edge servers. Specifically, first, a large amount of data is used to train the CloudYOLO model in the cloud, and its Backbone part is split out and deployed on the edge server as part of EdgeYOLO. CloudYOLO is deployed in the cloud, while EdgeYOLO is deployed on the edge server, mainly for fast extraction of basic features. As a more complex model, CloudYOLO has higher accuracy but higher computational overhead. In order to shorten the inference time in emergency recognition scenarios, EdgeYOLO uses the Backbone part of YOLO for efficient feature extraction at the edge, thereby reducing the complexity of edge computing. By sharing the feature extraction part of the first m layers of CloudYOLO to assist EdgeYOLO in training, the generalization ability and overall performance of the edge model can be significantly improved, avoiding the high computational cost of training the model from scratch. With this cloud-edge collaborative approach, EdgeYOLO can not only save computing resources, but also improve recognition accuracy when data is limited, thereby better adapting to the needs of complex emergency environments.

[0112] In addition, in actual scenarios, the client side will continue to upload data to the edge server, and it is proposed to use these uploaded data to assist in retraining EdgeYOLO to further improve its performance. Specifically, the auxiliary process of CloudYOLO can be divided into the initialization phase and the update phase.

[0113] 3.1 Model initialization phase

[0114] In the initialization phase, we first train CloudYOLO on the cloud server using historical image data from power emergency scenarios. Then we extract the first m feature layers of CloudYOLO (i.e., the Backbone part) to assist the training of EdgeYOLO. c is the trained CloudYOLO parameter, W mc The parameters of its Backbone part, W e are the parameters of EdgeYOLO. With the assistance of CloudYOLO, using W mc Replace the corresponding Backbone parameters in EdgeYOLO, so that only W needs to be trained e-mc , where W e-mc Indicates removal of W mc The remaining parameters of EdgeYOLO. Therefore, given the training data EdgeYOLO is trained to optimize the following loss function:

[0115]

[0116] In the formula, f e (W e-mc ) is the total loss function used when training EdgeYOLO, H is used to represent the loss function of training to the i-th picture, σ is used to represent the activation function, σ(f(x i ; W e-mc ) is the calculation function of the model in the forward propagation process, which is used to calculate the input x i and the current parameter W e-mc Generate output.

[0117] When the edge server receives the first m feature layers (i.e., W mc ), randomly initialize n higher layers (n is a positive integer) and compare them with W mc Then, by freezing W mc And fine-tune W e-mc To train EdgeYOLO, as shown above. It should be noted that the freezing operation means that the frozen parameters will not change when training the neural network model. By sharing the first m feature layers, EdgeYOLO not only improves the recognition accuracy, but also saves a lot of computing resources. This is because EdgeYOLO inherits the knowledge of CloudYOLO and only needs to train n higher layers. It should be noted that W e ' -mc Represents the updated value. The symbol "∪" represents the connection of two parameter sets. For example, S1 represents the first m layers of parameters of the neural network model, S2 represents the remaining parameters of the neural network model, and S1∪S2 represents the parameters of the entire neural network model.

[0118] 3.2 Model update phase

[0119] Through the initialization phase, EdgeYOLO overcomes the overfitting problem by sharing the feature layer of CloudYOLO. In actual scenarios, the client may continue to upload data to the edge server. In mobile applications based on deep learning models, training data is crucial to improving application performance. Based on this, these uploaded data are used to further assist in training EdgeYOLO. However, this process is challenging because these data are usually unlabeled. Therefore, consider using CloudYOLO for auxiliary labeling and then updating EdgeYOLO.

[0120] Specifically, CloudYOLO is a high-precision deep convolutional neural network, such as ResNet, which has achieved an accuracy of 96.43% on the ImageNet dataset. Therefore, CloudYOLO can be used to predict the labels of uploaded data. In the present invention, it is assumed that CloudYOLO can always accurately predict labels. After receiving the uploaded data, the edge server first preprocesses the data, such as using target detection and target segmentation techniques to obtain segmented targets. This is because target detection and target segmentation can remove information irrelevant to the target, reduce the amount of data transmission between the edge server and the cloud server, and do not affect the integrity of the target data. Target detection, target segmentation, and target recognition are different convolutional neural network models, and the main focus is on target detection. It should be noted that target recognition and segmentation techniques have been fully studied in many studies.

[0121] Next, the edge server runs EdgeYOLO to obtain the recognition results and returns them to the end side. It is worth noting that the edge server saves the segmented target data. When the core network load is low, the edge server uploads these segmented targets to the cloud server. After receiving the targets, the cloud server predicts their labels using CloudYOLO and sends these labels back to the edge server to store the corresponding target data. In this way, EdgeYOLO can use these labeled target data for retraining. Similar to the initialization stage, EdgeYOLO freezes its first m feature layers when retraining and fine-tunes the following n layers.

[0122] CloudYOLO can continuously predict labels and send them to the edge server where the corresponding data is stored. Therefore, the training process of EdgeYOLO is continuous. The update trigger condition can be triggered by setting a time interval or by the accumulated recognition data volume.

[0123] The effect of this scheme is further explained below with reference to specific experiments.

[0124] The power AR emergency image recognition platform built by this solution verifies the accuracy of EdgeYOLO in target detection tasks and its rapid response capability, and then shows the experimental results on the COCO dataset to evaluate the effectiveness of the framework in providing on-site recognition services in emergency environments. The experimental results show that with the support of CloudYOLO, EdgeYOLO significantly improves recognition accuracy and reduces response time.

[0125] First, CloudYOLO is trained using the training set, where EdgeYOLO-r represents a shallow convolutional neural network model trained from scratch, and EdgeYOLO-s represents a shallow YOLO model trained based on the CloudYOLO shared layer. EdgeYOLO-s performs significantly better than EdgeYOLO-r in terms of detection accuracy, especially when there are fewer training images. Figure 3 The detection accuracy results of EdgeYOLO on the COCO dataset are shown. It can be seen that compared with EdgeYOLO-r, EdgeYOLO-s has a 7.3% improvement in detection accuracy when the number of training images is 500. The main reason is that when the number of training images is small, it is difficult for EdgeYOLO-r to converge quickly and learn the distribution characteristics of the data from scratch, while CloudYOLO uses a large number of images for training in the cloud, which can better grasp the data distribution. By inheriting the knowledge of these shared layers, EdgeYOLO-s can also learn data features more effectively, thereby achieving higher detection accuracy.

[0126] also, Figure 3 It shows that when the number of training images is 100, the detection accuracy of EdgeYOLO-r is low, mainly due to the overfitting problem caused by insufficient training data. By using the shared layer of CloudYOLO, EdgeYOLO-s has a higher accuracy because it can learn the distribution of data with the knowledge of the deep model instead of learning from scratch. This shows that with the assistance of the deep model, EdgeYOLO can effectively avoid overfitting.

[0127] Continuously increasing the number of training images can further improve the accuracy of EdgeYOLO. Figure 3 It shows that when the training images increase from 100 to 4500, the accuracy of EdgeYOLO-s increases by 55.7%, which shows that with the help of the labeled data provided by CloudYOLO, EdgeYOLO can significantly improve the detection performance. In contrast, EdgeYOLO-r is relatively low due to the lack of labeled data. For example, when the training images are 100, the detection accuracy of EdgeYOLO-r is 10.2%. When the amount of training data increases from 100 to 4500, the detection accuracy of EdgeYOLO-r increases by 47.6%, which shows that a large amount of training data helps to better capture the distribution characteristics of the data. Overall, the experimental results show that for practical applications based on deep learning, it is crucial to collect enough trainable data.

[0128] In addition, in order to verify the advantage of EdgeYOLO in response speed, two deployment schemes are designed for verification to evaluate the recognition efficiency of the edge and cloud in power emergency scenarios, and to verify the significant advantage of the edge in response speed when using EdgeYOLO-s for target detection. The specific scheme is as follows:

[0129] (1) Deploy EdgeYOLO-s on the edge server

[0130] EdgeYOLO-s is deployed on edge servers to achieve localized processing, reduce latency caused by data transmission, and improve the ability to respond quickly to emergency events. Edge servers can perform reasoning close to the data source, thereby providing real-time detection results more efficiently.

[0131] (2) Deploy EdgeYOLO-s on a cloud server

[0132] A standard YOLOv8 target detection model was deployed on a cloud server as a control group. Although the cloud model has stronger computing power by centrally processing image data, its response time is often affected by the bandwidth and latency of network transmission.

[0133] Figure 4 The response time and upload time of two different deployment schemes are shown. As can be seen from the figure, compared with CloudYOLO, EdgeYOLO-s has an average response time reduction of 55.1% when the image size is 256k. This significant performance improvement is due to two reasons: first, the geographical location of the edge server deployment is closer to the end side, and the data transmission delay is greatly reduced; second, the network structure of EdgeYOLO-s is more lightweight, making the processing speed faster. By sinking computing resources to the edge, EdgeYOLO-s can more effectively respond to the real-time detection needs of emergency scenarios and significantly shorten the overall response time.

[0134] Uploading images to edge servers also significantly reduces upload time. Figure 4 It shows that compared with uploading images to the cloud server, uploading images to the edge server reduces the average upload time by 53.4% ​​when the image data volume is 512K. The main reason is that the edge server is closer to the end side, and usually only one hop is required between the mobile device and the edge server, which greatly shortens the path and time of data transmission. The experimental results show that using edge servers to provide cognitive services is of great significance in emergency scenarios, which helps to achieve faster response speed and higher service quality. Therefore, edge computing shows obvious advantages in reducing upload delays and provides strong support for providing real-time cognitive services.

[0135] In summary, the power AR emergency image recognition method based on end-side complementarity and edge collaboration provided by the present invention can at least have the following beneficial effects:

[0136] (1) The present invention uses a multi-scale and multi-directional structural element morphological filter to enhance and detect edges of multi-source images, and combines spatial transformation to achieve image registration. Specifically, a multi-scale and multi-source image fusion technology based on mathematical morphology is used to reduce noise through structural element filters of different sizes and shapes, retain multi-directional edge features, and perform registration and weighted fusion of multimodal data. This method integrates the advantageous information of multi-source images, significantly improves image quality and fusion efficiency, and provides efficient and reliable technical support for abnormal detection and emergency decision-making of power equipment.

[0137] (2) The present invention transforms the image model parameters uploaded by the edge node into the frequency domain through fast Fourier transform, uses the average aggregation of the low-frequency part to share the global common knowledge, and retains the high-frequency part to support personalized feature optimization, thereby achieving a balance between knowledge sharing and personalized models. For the low-frequency extraction of frequency domain parameters, a low-frequency mask and weighting strategy are designed to ensure aggregation accuracy and computational efficiency. Finally, the model parameters are restored and optimized through inverse Fourier transform to achieve the generation of a personalized model. This method reduces the amount of data transmission between edge nodes and improves the real-time, collaborative and adaptable performance of the model in complex scenarios.

[0138] (3) The goal of the present invention is to use the edge server to complete all recognition operations, thereby providing on-site recognition services in emergency environments. This means that after capturing data, the wearable AR device directly uploads the data to the nearest edge server without any preprocessing. This design can effectively reduce the computing burden of wearable devices and extend their battery life. In addition, since the physical distance between the end side and the edge server is relatively close, the data transmission delay is relatively small, thereby improving the overall response speed and meeting the real-time requirements in emergency scenarios.

[0139] (4) This solution effectively reduces the computing pressure of the device and extends its working time by transferring the computing load from the wearable AR device to the edge server.

[0140] (5) Through the collaborative work of CloudYOLO and EdgeYOLO and continuous data upload, the recognition accuracy of EdgeYOLO is significantly improved.

[0141] The present specification also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute a method in any one of the embodiments in the specification.

[0142] The present specification also provides a computing device, including a memory and a processor, wherein executable codes are stored in the memory, and when the processor executes the executable codes, a method in any embodiment of the present specification is implemented.

[0143] Since the device embodiments provided by the present invention are based on the same inventive concept as the method embodiments of this specification, the specific contents can be found in the description of the method embodiments of this specification and will not be repeated here.

[0144] The modules or units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. The above disclosure is only the preferred embodiment of the present invention, and of course it cannot be used to limit the scope of the rights of the present invention. Those skilled in the art can understand that all or part of the processes of the above embodiment are implemented, and the equivalent changes made according to the claims of the present invention still fall within the scope of the invention.

Claims

1. A method for electric power AR emergency image recognition based on end-side complementarity and edge collaboration, characterized in that: include: Construct an electric power AR emergency image recognition platform; wherein the recognition platform is deployed with an image recognition model, and the recognition platform includes: a wearable AR device, an edge server and a cloud server; each cloud server is correspondingly provided with a number of edge servers, and each edge server is correspondingly provided with a number of wearable AR devices; when constructing the recognition platform, the wearable AR device is used to perform multi-source fusion of historical image data collected in electric power emergency scenes; the edge server deploys an EdgeYOLO model, which is used to perform frequency domain fusion and optimization of image features after multi-source fusion based on Fourier federated learning; the cloud server deploys a CloudYOLO model, which is used to perform deep calculation and model optimization; and the Backbone part of the CloudYOLO model is deployed on the edge server as a part of the EdgeYOLO model, which is used to assist EdgeYOLO in federated learning; Real-time image data in power emergency scenarios can be collected through wearable AR devices; The edge server runs the EdgeYOLO model to recognize the real-time image data and outputs the recognition result.

2. The power AR emergency image recognition method based on end-side complementarity and edge collaboration according to claim 1 is characterized in that: The wearable AR device is used to collect image data in power scenes through cameras and sensors, and perform image denoising, image edge extraction and edge image fusion on the image data; wherein the image data includes device status, location information and environmental characteristics.

3. The power AR emergency image recognition method based on end-side complementarity and edge collaboration according to claim 2 is characterized in that: The image denoising process includes: The original image F collected by the wearable AR device is processed by at least two mathematical morphological filters of different sizes and shapes to obtain a denoised image F. y ; The process of image edge extraction includes: Using structural elements C in different directions ni , use the following calculation formula to calculate the denoised image F y Edge features are extracted in different directions to obtain edge images in each direction: In the formula, C ni represents the mathematical morphological filter corresponding to the nth size and shape in the i-th direction, E ni Characterizes the directional edge image corresponding to the nth size and shape in the i-th direction, is used to represent matrix addition, and Θ is used to represent matrix subtraction; The image fusion process includes: Based on the following calculation formula, the directional edge images obtained in each direction are weighted and summed using the weight coefficient to obtain a complete edge image: In the formula, E n (F) Edge image corresponding to the mathematical morphological filter used to characterize the nth size and shape, q i Used to represent the weight coefficient corresponding to the i-th direction.

4. The power AR emergency image recognition method based on end-side complementarity and edge collaboration according to claim 3 is characterized in that: The processing of the original image by using mathematical morphological filters of at least two sizes and shapes comprises: The mathematical morphological filters B1 and C1 are connected in series in descending order of size, and a morphological operation is performed on the original image F to obtain a denoised image F1; The mathematical morphological filters B2 and C2 are connected in series in order of size from large to small, and a morphological operation is performed on the original image F to obtain a denoised image F2; Using the following calculation formula, images F1 and F2 are connected in parallel according to the weight coefficient related to the peak signal-to-noise ratio to obtain the denoised image F y : F y =q1F1+q2F2 in, q1, q2∈[0,1], and q1+q2=1.

5. The power AR emergency image recognition method based on end-side complementarity and edge collaboration according to claim 1 is characterized in that: The frequency domain fusion and optimization of the image features after multi-source fusion based on Fourier federated learning includes: The convolutional model θ k The convolutional layer parameters Convert to a two-dimensional matrix Where O and C are the number of output channels and input channels respectively, s1 and s2 are the spatial shapes of the convolution kernel; For the transformed two-dimensional matrix w′ k Perform fast Fourier transform to obtain the amplitude graph F A and phase diagram F P : The kth local model is aggregated in the frequency domain space through the low-frequency mask, and we get Based on the following calculation formula, the amplitude is mapped using the inverse Fourier transform and the phase map F P Convert to parameter form In the formula, F -1 Used to characterize the inverse Fourier transform.

6. The power AR emergency image recognition method based on end-side complementarity and edge collaboration according to claim 5 is characterized in that: Use the following calculation formula to calculate the two-dimensional matrix w′ k Perform a fast Fourier transform: In the formula, m and n are given parameters; Based on the following calculation formula, the kth local model is aggregated in the frequency domain space through the low-frequency mask: Where G is the low-frequency mask, Z is the indicator function, and g∈(0, 0.5) represents the low-frequency threshold.

7. The power AR emergency image recognition method based on end-side complementarity and edge collaboration according to claim 1 is characterized in that: The initial stage of deploying the EdgeYOLO model on the edge server includes: Use historical image data from power emergency scenarios to train CloudYOLO on a cloud server; Extract the first m feature layers corresponding to CloudYOLO's Backbone to assist EdgeYOLO training; With the help of CloudYOLO, using W mc Replace the corresponding Backbone parameters in EdgeYOLO; where W c is the trained CloudYOLO parameter, W mc The parameters of its Backbone part, W e are the parameters of EdgeYOLO; Training parameter W e-mc ; Among them, W e-mc For characterization of W removal mc The remaining parameters of EdgeYOLO; When the edge server receives the parameters W of the first m feature layers mc After that, n higher layers are randomly initialized and combined with W mc Connect to form EdgeYOLO; Freeze parameter W mc , and use local data to W e-mc Train and get the final EdgeYOLO model.

8. The power AR emergency image recognition method based on end-side complementarity and edge collaboration according to claim 7 is characterized in that: For a given training data The loss function used when training EdgeYOLO is expressed as follows: In the formula, f e (W e-mc ) is the total loss function used when training EdgeYOLO, H is used to represent the loss function of training to the i-th picture, σ is used to represent the activation function, σ(f(x i ; W e-mc ) is the calculation function of the model in the forward propagation process, which is used to calculate the input x i and the current parameter W e-mc Generate output.

9. The power AR emergency image recognition method based on end-side complementarity and edge collaboration according to claim 1 is characterized in that: After each recognition result is output, it further includes: The edge server uploads real-time image data and recognition results to the cloud server; The cloud server accumulates the uploaded data and forms sample training data for updating the model based on the real-time image data and the recognition result; According to the preset period, or when the accumulated sample training data reaches the preset value, the cloud server sends the sample training data for updating to the edge server; The edge server updates EdgeYOLO using the sample training data sent by the cloud server.

10. A computing device comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Power supply cooperative control method and system based on sensing edge calculation

    CN120497924A

  • Self-driving automobile environment sensing method based on federal learning and automobile and road cloud cooperation

    CN120543994A