Target detection method for image collected by AR wearable device based on improved YOLOv8
By improving the YOLOv8 model, combining dynamic convolutional network, effective channel attention mechanism and convolutional attention mechanism, the problem of low recognition accuracy of AR wearable devices in power communication field operation and maintenance is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411800055.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-06
AI Technical Summary
In the on-site operation and maintenance of power communications, the recognition accuracy of power equipment is poor in the face of complex environments and variable equipment and lighting conditions.
Using the object detection method based on improved YOLOv8, an improved YOLOv8 model is constructed by acquiring and labeling historical image data. The model includes dynamic convolutional network module DC_C2f, effective channel attention mechanism module ECA and convolutional attention mechanism module CBAM to improve the robustness and recognition accuracy of the model.
It significantly improves the power equipment recognition accuracy of AR wearable devices in complex environments and variable conditions, and enhances the model's robustness and recognition capabilities for different conditions.
Smart Images

Figure CN119942059A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a target detection method for images collected by an AR wearable device based on an improved YOLOv8. Background Art
[0002] Augmented reality (AR) technology is a technology that combines virtual information with the real environment. It captures the user's actual environment through a camera, processes this information in real time, and superimposes virtual elements on the real scene, providing real-time interaction and immersive experience, allowing users to obtain value-added information in the actual environment, thereby improving decision-making capabilities and operational efficiency. In power maintenance scenarios, AR wearable devices such as smart glasses can help technicians obtain key information and receive collaborative instructions and guidance from back-end personnel during repair and maintenance, enabling technicians to identify problems and take action more quickly in complex environments, thereby improving work efficiency.
[0003] Object detection technology provides strong support for AR technology. Existing object detection methods include region-based convolutional neural networks, single-stage detectors such as SSD and YOLO, etc., each of which has its own advantages and disadvantages. Region-based methods are usually more accurate, but the computational complexity is large and the speed is slow, making it difficult to meet the needs of real-time applications. Relatively speaking, the YOLO (You Only Look Once) series of detectors have attracted widespread attention for their fast and efficient characteristics, especially the latest YOLOv8 model has significantly improved detection accuracy, speed and robustness, and can better handle complex backgrounds, low-light conditions and objects of various scales, which makes YOLOv8 an ideal choice for AR technology.
[0004] Although in terms of AR precise recognition, the recognition accuracy of existing AR recognition methods can meet the needs of on-site operation and maintenance to a certain extent under good environmental conditions and obvious equipment features. However, in the actual on-site operation and maintenance of power communications, there are many facilities on the operation and maintenance site, and the characteristics of the same cluster of equipment are slightly different. In addition, there are variable equipment and environmental conditions such as equipment wear, occlusion, and unsatisfactory lighting conditions, which leads to poor recognition accuracy of existing AR wearable device target detection methods for power equipment. Summary of the invention
[0005] In view of this, in order to address the above shortcomings, it is necessary to propose a target detection method for images collected by AR wearable devices based on improved YOLOv8, so as to improve the recognition accuracy of AR wearable devices for power equipment in actual environments.
[0006] In a first aspect, the present invention provides a method for detecting an object in an image collected by an AR wearable device based on an improved YOLOv8, comprising:
[0007] Obtain historical image data of AR wearable devices;
[0008] Annotating the historical image data to obtain sample training data;
[0009] An improved YOLOv8 model is constructed; wherein the improved YOLOv8 model includes a backbone network, a neck network and a head network; the improvements of the improved YOLOv8 model compared to the traditional YOLOv8 model include: the convolutional network module C2f in the backbone network of the traditional YOLOv8 model is replaced by a dynamic convolutional network module DC_C2f, an effective channel attention mechanism module ECA is added to the shallow layer of the backbone network, and a convolutional attention mechanism module CBAM is added to the deep layer of the backbone network and before the detection head of the head network;
[0010] Using the sample training data to train the improved YOLOv8 model;
[0011] The image data to be tested collected in real time by the AR wearable device is input into the trained improved YOLOv8 model to identify and detect power equipment targets.
[0012] Preferably, the dynamic convolutional network module DC_C2f is used to extract features from the input data using convolution kernels of variable sizes to capture temporal and spatial variation characteristics in the data.
[0013] Preferably, the dynamic convolutional network module DC_C2f uses the following calculation formula to extract features from the input data:
[0014]
[0015] Among them, x and y are the input feature map and the feature map enhanced by the dynamic convolutional network module, respectively. represents the weighted sum vector of K convolution kernels, and its transposed vector is represents the weight parameters of K convolution kernels, represents the bias weighted sum vector of K convolution kernels, g represents the ReLU activation function, π k (x) represents the weight coefficient obtained by the attention module, which is adaptively generated by nonlinearly aggregating k parallel convolutions through the attention mechanism.
[0016] Preferably, the effective channel attention mechanism module ECA is added between the first DC_C2f module and the third Conv module, and is used to output channel weight values for feature channels representing key features by introducing an attention mechanism, so as to select corresponding key feature channels according to the channel weight values for feature extraction by the third Conv module; wherein the Conv module represents a convolution module.
[0017] Preferably, the effective channel attention mechanism module ECA outputs the channel weight value of the feature channel through the following calculation formula:
[0018]
[0019] Among them, ω represents the channel weight value, C1D represents one-dimensional convolution, k is the size of the convolution kernel, δ represents the Sigmoid activation function, W is the width of the feature map, H is the height of the feature map, and χ ij Represents the element value at position in the feature map.
[0020] Preferably, the adding positions of the convolutional attention mechanism module CBAM include: adding a first CBAM module between the third DC_C2f module and the fifth Conv module, adding a second CBAM module between the fourth DC_C2f module and the spatial pooling pyramid SPPF module, adding a third CBAM module between the second C2f module and the first output module, adding a fourth CBAM module between the third C2f module and the second output module, and adding a fifth CBAM module between the fourth C2f module and the third output module; each CBAM module is a dual attention mechanism combining channel attention and spatial attention, which is used to enhance the representation of useful features by learning the important weights of each channel, and to construct a spatial attention map by analyzing the spatial relationship between features.
[0021] Preferably, the CBAM module calculates the spatial relationship between the important weights of the channels and the features by the following calculation formula:
[0022]
[0023] Among them, M c (F) represents the important weight of the channel, M s (F) represents the spatial relationship of the features, σ represents the sigmoid activation function, F is the input feature, W0 and W1 correspond to the weight matrices of the two fully connected operations respectively, and F avg Represents the result obtained by global average pooling of input feature F, F max Represents the result obtained by global maximum pooling of input feature F, f r Represents a convolution operation with a kernel size of r.
[0024] Preferably, in the training of the improved YOLOv8 model using the sample training data, a total loss function determined by a bounding box loss function, a classification loss function and a confidence loss function is used for model training.
[0025] Preferably, the bounding box loss function is expressed as follows:
[0026]
[0027] Among them, L CIoU is the bounding box loss function, L IoU is the intersection-over-union ratio of the predicted box and the label box, α is the penalty term used to balance the parameters, v is the aspect ratio metric function used to measure the consistency of the aspect ratio, W g and H g are the length and width of the bounding box of the predicted box and the real box, respectively, (W i ,H i ) is the length and width of the overlapping part of the prediction box and the label box, (w,h) is the length and width of the prediction box, (w gt ,h gt ) is the length and width of the label box, (x, y) is the center position of the prediction box, (x gt ,y gt ) is the center position of the label frame;
[0028] The classification loss function is expressed as follows:
[0029]
[0030] Where C is the number of classification categories, y j is the real category, is the predicted class probability;
[0031] The confidence loss function is expressed as:
[0032]
[0033] Among them, p i is the true label of the target existence, is the confidence of the prediction, and N is the total number of prediction samples;
[0034] The total loss function is expressed as:
[0035] L=λ CIoU L CIoU +λ cls L cls +λ conf L conf
[0036] Among them, λ CIoU ,cls , conf are the weight hyperparameters corresponding to the bounding box loss function, classification loss function, and confidence loss function, respectively.
[0037] In a second aspect, the present invention provides a computing device, including a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, any method described in the first aspect is implemented.
[0038] It can be seen from the above technical scheme that in the target detection method of the image collected by the AR wearable device based on the improved YOLOv8 provided by the present invention, the historical image data of the AR wearable device is first obtained, and the historical image data is annotated to obtain sample training data; then an improved YOLOv8 model is constructed, and the improved YOLOv8 model is trained using the sample training data, so that the improved YOLOv8 model obtained by training is used to detect and identify the image data to be tested collected in real time, and the target recognition result of the power equipment can be obtained. In this scheme, compared with the traditional YOLOv8 model, the improved YOLOv8 model includes replacing the convolutional network module C2f in the backbone network of the traditional YOLOv8 model with a dynamic convolutional network module DC_C2f, adding an effective channel attention mechanism module ECA to the shallow layer of the backbone network, and adding a convolutional attention mechanism module CBAM to the deep layer of the backbone network and before the detection head of the head network. Replacing the convolutional network module C2f with the dynamic convolutional network module DC_C2f in the backbone network can dynamically generate convolution kernels according to the input features through the dynamic convolutional network module DC_C2f, thereby significantly improving the accuracy of feature extraction without increasing the depth or width of the network, and effectively improving the robustness of the model for image recognition under different conditions; and adding the effective channel attention mechanism module ECA can improve feature selectivity, reduce the impact of redundant information, and make the model more focused on the most recognizable features, thereby improving the accuracy of target detection; and adding the convolutional attention mechanism module CBAM in the deep layer of the backbone network can dynamically adjust the weight of the feature channel, thereby focusing on the most relevant features, which helps to improve the model's ability to recognize important targets; and adding the convolutional attention mechanism network CBAM in front of the detection head of the head network can make the model more focused on key features related to target detection, thereby improving the accuracy of classification and positioning, which is conducive to accurate detection in complex backgrounds or low-quality images. Therefore, this solution can improve the recognition accuracy of AR wearable devices for power equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A flowchart of a method for target detection in images collected by an AR wearable device based on improved YOLOv8 is provided in an embodiment of the present invention.
[0040] Figure 2 A framework diagram of an improved YOLOv8 algorithm provided in an embodiment of the present invention.
[0041] Figure 3 The following is a comparison of the parameters of the prediction box and the label box. DETAILED DESCRIPTION
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0043] like Figure 1 As shown, the present invention provides a target detection method for images collected by an AR wearable device based on improved YOLOv8, and the method may include the following steps:
[0044] Step 101: Acquire historical image data of the AR wearable device;
[0045] Step 102: annotate the historical image data to obtain sample training data;
[0046] Step 103: construct an improved YOLOv8 model; wherein the improved YOLOv8 model includes a backbone network, a neck network and a head network; the improvements of the improved YOLOv8 model compared to the traditional YOLOv8 model include: the convolutional network module C2f in the backbone network of the traditional YOLOv8 model is replaced by a dynamic convolutional network module DC_C2f, an effective channel attention mechanism module ECA is added to the shallow layer of the backbone network, and a convolutional attention mechanism module CBAM is added to the deep layer of the backbone network and before the detection head of the head network;
[0047] Step 104: Using sample training data to train the improved YOLOv8 model;
[0048] Step 105: Input the image data to be tested collected in real time by the AR wearable device into the trained improved YOLOv8 model to identify and detect the power equipment target.
[0049] For step 101 and step 102.
[0050] Step 101 and step 102 are mainly used to obtain historical image data of AR wearable devices, and then label the historical image data to form sample training data. Specific label information may include bounding box coordinates and category labels in YOLO format. Of course, before forming sample training data, the data should also be preprocessed, such as adjusting the input image to the input size required by the model, and enhancing the data through random cropping, flipping, rotation, brightness adjustment and other operations to improve the robustness of the trained model.
[0051] For step 103, an improved YOLOv8 model is constructed;
[0052] YOLOv8 (You Only Look Once version 8) is a deep learning algorithm in the field of target detection. It has fast reasoning ability and high-precision recognition effect, which can meet the dual requirements of real-time and accuracy of target detection in AR devices. YOLOv8 is mainly composed of three parts: backbone, neck and head. Among them, the backbone network is mainly responsible for feature extraction. Through multiple layers of convolutional layers and activation functions, it converts the input image into a high-dimensional feature representation; the neck is to fuse and process the features extracted from the backbone to generate a multi-scale feature map; the head is composed of multiple detection heads, which is responsible for converting the feature map generated by the neck into the final detection result. The improved YOLOv8 algorithm framework is as follows: Figure 2The backbone layer from top to bottom is the first Conv module, the second Conv module, the first DC_C2f module, the ECA module, the third Conv module, the second DC_C2f module, the fourth Conv module, the third DC_C2f module, the first CBAM module, the fifth Conv module, the fourth DC_C2f module, the second CBAM module and the SPPF module; among which, another output of the second DC_C2f module is to the second Concat module, another output of the third DC_C2f module is to the first Concat module, one output of the SPPF is to the first Upsample module, and another output is to the fourth Concat module; the neck network along the data flow direction is the first Upsample module, the fifth Conv module, the fifth DC_C2f module, the fifth CBAM module, the fifth Conv module, the fifth DC_C2f module, the fifth CBAM module and the SPPF module; among which, another output of the second DC_C2f module is to the second Concat module, another output of the third DC_C2f module is to the first Concat module, one output of the SPPF is to the first Upsample module, and another output is to the fourth Concat module. A Concat module, a first C2f module, a second Upsample module, a second Concat module, a second C2f module, and a third CBAM module. One output of the third CBAM module is sent to the first output module, and the other output is sent to the sixth Conv module. The sixth Conv module is followed by the third Concat module, the third C2f module, the seventh Conv module, the fourth Concat module, and the fourth C2f module. Another output of the first C2f module is sent to the third Concat module, another output of the third C2f module is sent to the fourth CBAM module, and the fourth CBAM module outputs to the second output module. The fourth C2f module outputs to the fifth CBAM module, and the fifth CBAM module outputs to the third output module.
[0053] Compared with the traditional YOLOv8, this embodiment mainly makes the following improvements:
[0054] (1) Replace the convolutional network module C2f in the traditional YOLOv8 model backbone network with the dynamic convolutional network module DC_C2f;
[0055] When using wearable AR devices for object recognition, there are often significant differences in quality factors such as image brightness and pixels. Effective measures need to be taken to improve the robustness of the model to ensure reliable image recognition under various environmental conditions. In addition, considering the computing load and energy consumption constraints faced by the model when running on AR devices, it is particularly important to choose the right convolution technology. In this context, dynamic convolution becomes an ideal solution.
[0056] Dynamic convolution can dynamically generate convolution kernels based on input features instead of applying fixed convolution kernels at each position. This method significantly improves the accuracy of feature extraction without increasing the depth or width of the network. This makes dynamic convolution particularly suitable for use on resource-limited end devices, and can effectively improve the robustness of the model for image recognition under different conditions.
[0057] In the traditional YOLOv8 model, C2f is mainly used for efficient extraction and feature fusion. In this embodiment, replacing the convolution in the C2f module of the backbone with dynamic convolution can bring significant benefits. First, dynamic convolution can improve the adaptability of the model to images of different qualities and conditions, and improve the accuracy of feature extraction, especially when dealing with low brightness, low resolution or complex backgrounds; secondly, dynamic convolution can reduce the dependence on network width or depth while improving feature extraction accuracy, and is suitable for resource-constrained environments such as AR wearable devices; in addition, dynamic convolution has the ability to flexibly process multi-scale features, which helps to enhance the model's recognition ability for small targets and complex structures, thereby improving the performance and robustness of YOLOv8 as a whole.
[0058] Dynamic convolution (DC) is an advanced operation in convolutional neural networks that aims to introduce stronger contextual information and spatiotemporal relationships to better understand the dynamic changes in the data. Traditional convolution operations are based on fixed-size convolution kernels to extract local features, which may lead to incomplete capture of information in some cases. Dynamic convolution uses convolution kernels of variable sizes to adapt to features of different scales and can be adjusted according to the dynamic changes of the data, which enables the model to better capture the temporal and spatial changes in the data, thereby improving the expressiveness of the model. Therefore, this solution uses a variable-size convolution kernel to extract features from the input data through the dynamic convolution network module DC_C2f to capture the temporal and spatial change characteristics in the data.
[0059] Specifically, the dynamic convolutional network module DC_C2f can extract features from the input data using the following calculation formula:
[0060]
[0061] Among them, x and y are the input feature map and the feature map enhanced by the dynamic convolutional network module, respectively. represents the weighted sum vector of K convolution kernels, and its transposed vector is represents the weight parameters of K convolution kernels, represents the bias weighted sum vector of K convolution kernels, g represents the ReLU activation function, π k (x) represents the weight coefficient obtained by the attention module, which is adaptively generated by nonlinearly aggregating k parallel convolutions through the attention mechanism.
[0062] It should be pointed out that the dynamic convolutional network first performs a depth-separable convolution operation, and takes different extraction features in each channel. It uses k convolution kernels for each layer of a single convolution, and nonlinearly aggregates k parallel convolutions through the attention mechanism to adaptively generate its weight coefficient πk (0<k≤K), which improves the computational efficiency and thus improves the feature expression capability of the model.
[0063] (2) Add an effective channel attention mechanism module ECA to the shallow layer of the backbone network;
[0064] In the YOLOv8 target detection process, shallow features usually contain rich edge and texture information, and in complex backgrounds or low-quality images, the expression of these features may be affected by noise. By introducing the attention mechanism, the model can dynamically identify and focus on the most important feature channels, thereby enhancing feature selectivity, increasing attention to key features, and enhancing the adaptability of the model, making the model more robust when facing diverse inputs and improving the accuracy of model recognition. In addition, due to the energy consumption and computing resource limitations of wearable AR devices, it is very necessary to choose a lightweight attention mechanism.
[0065] The benefits of introducing an effective channel attention mechanism are obvious. First, it can improve feature selectivity, reduce the impact of redundant information, and make the model more focused on the most discernible features, thereby improving the accuracy of target detection. Secondly, the channel attention mechanism is relatively lightweight and has low computational overhead. Although it will increase the amount of calculation, the overall computational efficiency is optimized by strengthening feature expression and improving the efficiency of subsequent layers. In addition, it helps the model perform more stably under various conditions and provides more flexible processing capabilities for diverse image inputs, thereby improving the performance of YOLOv8 as a whole.
[0066] Therefore, this embodiment considers adding an effective channel attention mechanism module ECA between the first DC_C2f module and the third Conv module, which is used to output channel weight values for feature channels representing key features by introducing an attention mechanism, so as to select corresponding key feature channels according to the channel weight values for feature extraction by the third Conv module; wherein the Conv module represents a convolution module.
[0067] Specifically, the effective channel attention mechanism module ECA can output the channel weight value of the feature channel through the following calculation formula:
[0068]
[0069] Among them, ω represents the channel weight value, C1D represents one-dimensional convolution, k is the size of the convolution kernel, δ represents the Sigmoid activation function, W is the width of the feature map, H is the height of the feature map, and χ ij Represents the element value of each position in the feature map, and g(χ) represents the channel global average pooling.
[0070] In this embodiment, the effective channel attention mechanism module ECA is improved on the basis of compression and excitation network. While preventing overfitting and improving robustness, it also improves the correlation between channels and enhances the aggregation of features. Moreover, the model has a simple structure and a small number of parameters. This module uses global average pooling and nonlinear adaptively determined one-dimensional convolution to replace the maximum pooling layer and the fully connected layer respectively, and finally outputs a corresponding weight value through the normalization function δ, so that the dimension reduction is avoided when learning the correlation between feature channels, and the number of parameters is greatly reduced while improving network performance.
[0071] (3) Add a convolutional attention mechanism module CBAM in the deep layer of the backbone network and before the detection head of the head network;
[0072] The deep part of the YOLOv8 backbone is mainly responsible for extracting high-level semantic features in the image and fusing features at different levels at multiple scales, helping the model to enhance its detection capabilities for objects of different sizes and shapes when processing multi-scale targets. The YOLOv8 model detection head is mainly responsible for classifying targets in the image based on the features extracted by the backbone and outputting information after feature extraction and fusion.
[0073] When dealing with complex scenes and diverse targets, deep features may be disturbed by background noise and redundant information. The convolutional attention mechanism module (CBAM) is an attention mechanism that combines the dual advantages of channel attention and spatial attention. By introducing the attention mechanism in the deep layer of the backbone, the channel attention mechanism adjusts the importance of each channel to focus on the most relevant features, that is, enhances the selectivity of features, so that the model pays more attention to key features, thereby improving classification and positioning accuracy; and adding an attention mechanism before the detection head can pay attention to the spatial dimension of the feature map, so that the model can focus more on key features related to target detection, ensuring that the model has stronger perception capabilities in key areas of the image and improving classification and positioning accuracy.
[0074] Specifically, the locations where the convolutional attention mechanism module CBAM is added include: adding a first CBAM module between the third DC_C2f module and the fifth Conv module, adding a second CBAM module between the fourth DC_C2f module and the spatial pooling pyramid SPPF module, adding a third CBAM module between the second C2f module and the first output module, adding a fourth CBAM module between the third C2f module and the second output module, and adding a fifth CBAM module between the fourth C2f module and the third output module; each CBAM module is a dual attention mechanism combining channel attention and spatial attention, which is used to enhance the representation of useful features by learning the important weights of each channel, and to construct a spatial attention map by analyzing the spatial relationship between features.
[0075] The CBAM module can calculate the important weights of channels and the spatial relationship of features through the following calculation formula:
[0076]
[0077] Among them, M c (F) represents the important weight of the channel, M s (F) represents the spatial relationship of the features, σ represents the sigmoid activation function, F is the input feature, W0 and W1 correspond to the weight matrices of the two fully connected operations respectively, and F avg Represents the result obtained by global average pooling of input feature F, F max Represents the result obtained by global maximum pooling of input feature F, f r Represents the convolution operation with a convolution kernel size of r, and MLP is the fully connected layer calculation.
[0078] In this embodiment, channel attention enhances the representation of useful features by learning the important weights of each channel, performs maximum pooling and average pooling on the input feature map, and aggregates the two pooling results to extract diverse information. Spatial attention constructs a spatial attention map by analyzing the spatial relationship between features to highlight important spatial areas.
[0079] For step 104, the improved YOLOv8 model is trained using the sample training data;
[0080] The training of the improved YOLOv8 model mainly includes the steps of input, parameter update, output and iterative training, which can be specifically included as follows:
[0081] (1) Sample training data preparation;
[0082] (2) Network initialization;
[0083] In this step, the improved YOLOv8 model architecture is loaded and the weight parameters are initialized as needed.
[0084] (3) Define the loss function
[0085] 1) Define the boundary loss function and use CIoU as the model bounding box loss function, which is expressed as follows:
[0086]
[0087] Among them, L CIoU is the bounding box loss function, L IoU is the intersection-over-union ratio of the predicted box and the label box, α is the penalty term used to balance the parameters, v is the aspect ratio metric function used to measure the consistency of the aspect ratio, W g and Hg are the length and width of the bounding box of the predicted box and the real box, respectively, (W i ,H i ) is the length and width of the overlapping part of the prediction box and the label box, (w,h) is the length and width of the prediction box, (w gt ,h gt ) is the length and width of the label box, (x, y) is the center position of the prediction box, (x gt ,y gt ) is the center position of the label box; the parameters of the prediction box and the label box are shown in the following figure Figure 3 shown.
[0088] 2) Define the classification loss function: Use cross entropy loss as the classification loss function, expressed as follows:
[0089]
[0090] Where C is the number of classification categories, y j is the real category, is the predicted class probability;
[0091] 3) Define the confidence loss function: Use cross entropy loss as the confidence function, expressed as follows:
[0092]
[0093] Among them, p i is the true label of the target existence, is the confidence of the prediction, and N is the total number of prediction samples;
[0094] 4) Define the total loss function: The total loss function is determined by the weighted sum of the three losses, expressed as follows:
[0095] L=λ CIoU L CIoU +λ cls L cls +λ conf L conf
[0096] Among them, λ CIoU , cls , conf They are the weight hyperparameters corresponding to the bounding box loss function, classification loss function, and confidence loss function, which are used to balance the impact of different losses on the model.
[0097] (4) Training loop
[0098] Perform N training iterations, and the training process for each batch is as follows:
[0099] 1) Input data: extract a batch of images and their corresponding labels from the training dataset;
[0100] 2) Forward propagation: pass the input image into the model to obtain the predicted bounding box, category probability and confidence;
[0101] 3) Calculate the loss: Calculate the loss value based on the predicted value and the true value label;
[0102] 4) Back propagation: Calculate the gradient and update the model parameters through back propagation, using the optimizer to adjust the weights to minimize the loss;
[0103] 5) Update parameters: Use hyperparameters such as learning rate to update the weights in the network.
[0104] (5) Validation and evaluation: After each iteration, the validation set is used to evaluate the model performance to ensure that the model is not overfitted.
[0105] (6) Save the model: Save the model weights periodically, usually when the validation loss is minimal, to facilitate subsequent use and reasoning.
[0106] (7) Output results: Output the final model weights and evaluation results, ready for inference and deployment.
[0107] Step 105: input the image data to be tested collected in real time by the AR wearable device into the trained improved YOLOv8 model to identify and detect the power equipment target.
[0108] In this step, after the model training is completed, the operation and maintenance personnel can collect image data through the wearable AR device, and transmit the real-time collected image data to the model for recognition and detection, and then identify the power equipment target to assist the operation and maintenance personnel in their work.
[0109] The present specification also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute a method in any one of the embodiments in the specification.
[0110] The present specification also provides a computing device, including a memory and a processor, wherein executable codes are stored in the memory, and when the processor executes the executable codes, a method in any embodiment of the present specification is implemented.
[0111] Since the device embodiments provided by the present invention are based on the same inventive concept as the method embodiments of this specification, the specific contents can be found in the description of the method embodiments of this specification and will not be repeated here.
[0112] The modules or units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. The above disclosure is only the preferred embodiment of the present invention, and of course it cannot be used to limit the scope of the rights of the present invention. Those skilled in the art can understand that all or part of the processes of the above embodiment are implemented, and the equivalent changes made according to the claims of the present invention still fall within the scope of the invention.
Claims
1. A target detection method for images collected by AR wearable devices based on improved YOLOv8, characterized in that: include: Obtain historical image data of AR wearable devices; Annotating the historical image data to obtain sample training data; An improved YOLOv8 model is constructed; wherein the improved YOLOv8 model includes a backbone network, a neck network and a head network; the improvements of the improved YOLOv8 model compared to the traditional YOLOv8 model include: the convolutional network module C2f in the backbone network of the traditional YOLOv8 model is replaced by a dynamic convolutional network module DC_C2f, an effective channel attention mechanism module ECA is added to the shallow layer of the backbone network, and a convolutional attention mechanism module CBAM is added to the deep layer of the backbone network and before the detection head of the head network; Using the sample training data to train the improved YOLOv8 model; The image data to be tested collected in real time by the AR wearable device is input into the trained improved YOLOv8 model to identify and detect power equipment targets.
2. The target detection method for images collected by an AR wearable device based on improved YOLOv8 according to claim 1, characterized in that: The dynamic convolutional network module DC_C2f is used to extract features from the input data using convolution kernels of variable sizes to capture the temporal and spatial variation characteristics in the data.
3. The target detection method for images collected by an AR wearable device based on improved YOLOv8 according to claim 1, characterized in that: The dynamic convolutional network module DC_C2f uses the following calculation formula to extract features from the input data: Among them, x and y are the input feature map and the feature map enhanced by the dynamic convolutional network module, respectively. represents the weighted sum vector of K convolution kernels, and its transposed vector is represents the weight parameters of K convolution kernels, represents the bias weighted sum vector of K convolution kernels, g represents the ReLU activation function, π k (x) represents the weight coefficient obtained by the attention module, which is adaptively generated by nonlinearly aggregating k parallel convolutions through the attention mechanism.
4. The target detection method for images collected by an AR wearable device based on improved YOLOv8 according to claim 1, characterized in that: The effective channel attention mechanism module ECA is added between the first DC_C2f module and the third Conv module, and is used to output channel weight values for feature channels representing key features by introducing an attention mechanism, so as to select corresponding key feature channels according to the channel weight values for feature extraction by the third Conv module; wherein the Conv module represents a convolution module.
5. The target detection method for images collected by an AR wearable device based on improved YOLOv8 according to claim 4, characterized in that: The effective channel attention mechanism module ECA outputs the channel weight value of the feature channel through the following calculation formula: Among them, ω represents the channel weight value, C1D represents one-dimensional convolution, k is the size of the convolution kernel, δ represents the Sigmoid activation function, W is the width of the feature map, H is the height of the feature map, and χ ij Represents the element value at position in the feature map.
6. The target detection method for images collected by an AR wearable device based on improved YOLOv8 according to claim 1, characterized in that: The locations where the convolutional attention mechanism module CBAM is added include: adding a first CBAM module between the third DC_C2f module and the fifth Conv module, adding a second CBAM module between the fourth DC_C2f module and the spatial pooling pyramid SPPF module, adding a third CBAM module between the second C2f module and the first output module, adding a fourth CBAM module between the third C2f module and the second output module, and adding a fifth CBAM module between the fourth C2f module and the third output module; each CBAM module is a dual attention mechanism combining channel attention and spatial attention, which is used to enhance the representation of useful features by learning the important weights of each channel, and to construct a spatial attention map by analyzing the spatial relationship between features.
7. The target detection method for images collected by an AR wearable device based on improved YOLOv8 according to claim 6, characterized in that: The CBAM module calculates the spatial relationship between the important weights of the channels and the features through the following calculation formula: Among them, M c (F) represents the important weight of the channel, M s (F) represents the spatial relationship of the features, σ represents the sigmoid activation function, F is the input feature, W0 and W1 correspond to the weight matrices of the two fully connected operations respectively, and F avg Represents the result obtained by global average pooling of input feature F, F max Represents the result obtained by global maximum pooling of input feature F, f r Represents a convolution operation with a kernel size of r.
8. The target detection method for images collected by an AR wearable device based on improved YOLOv8 according to claim 1, characterized in that: In the training of the improved YOLOv8 model using the sample training data, a total loss function determined by a bounding box loss function, a classification loss function, and a confidence loss function is used for model training.
9. The target detection method for images collected by an AR wearable device based on improved YOLOv8 according to claim 8, characterized in that: The bounding box loss function is expressed as follows: Among them, L CIoU is the bounding box loss function, L IoU is the intersection-over-union ratio of the predicted box and the label box, α is the penalty term used to balance the parameters, v is the aspect ratio metric function used to measure the consistency of the aspect ratio, W g and H g are the length and width of the bounding box of the predicted box and the real box, respectively, (W i ,H i ) is the length and width of the overlapping part of the prediction box and the label box, (w,h) is the length and width of the prediction box, (w gt ,h gt ) is the length and width of the label box, (x, y) is the center position of the prediction box, (x gt ,y gt ) is the center position of the label frame; The classification loss function is expressed as follows: Where C is the number of classification categories, y j is the real category, is the predicted class probability; The confidence loss function is expressed as: Among them, p i is the true label of the target existence, is the confidence of the prediction, and N is the total number of prediction samples; The total loss function is expressed as: L=λ CIoU L CIoU +λ cls L cls +λ conf L conf Among them, λ CIoU , cls , conf are the weight hyperparameters corresponding to the bounding box loss function, classification loss function, and confidence loss function, respectively.
10. A computing device comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Visual identification-based intelligent decision-making control method and system for body
CN120279536A
Infrared image adaptive detection method and system facing complex background interference
CN120298679A
Driver dangerous driving behavior visual identification method and device, terminal and medium
CN120340003A
Steel mill material taking behavior identification method and system based on target detection
CN120599531A
Disaster-causing marine organism and environment detection method and system
CN120778696A