Target detection method based on improved YOLOv11

By embedding the dual attention mechanism module in the Bottleneck module of the YOLOv11 model, combining frequency domain-space coordinated attention and channel attention, the problem of low detection accuracy of YOLOv11 in complex scenarios is solved, significantly improving the accuracy and robustness of target detection.

CN119942406APending Publication Date: 2025-05-06JIAXING UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510018109.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

YOLOv11 has low detection accuracy and long detection period in complex backgrounds, occlusions and small object detection scenarios. This is mainly due to the lack of attention to important information during feature extraction by traditional convolutional networks, which leads to the reduction of robustness of the model when dealing with complex scenarios.

Method used

The dual attention mechanism module is embedded in the C3K2-layer Bottleneck module of the YOLOv11 convolutional neural network detection model, combining frequency domain-space coordinated attention and channel attention to enhance the model's attention ability to pay attention to important features.

Benefits of technology

By enhancing the model's ability to pay attention to important features, the accuracy and robustness of object detection are significantly improved, especially in complex scenarios, occlusions and small object detections, identifying and positioning targets more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942406A_ABST
    Figure CN119942406A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method based on improved YOLOv11, and relates to the technical field of target detection. According to the target detection method provided by the invention and the target detection method based on the improved YOLOv11, a double attention mechanism module is embedded into a Bottleneck module of a C3K2 layer of a YOLOv11 convolutional neural network detection model, and training is carried out to obtain a target detection neural network model; the dual attention mechanism module is composed of a channel attention module and a frequency domain-space collaborative attention module which are parallel, so that multi-dimensional feature extraction of a target image is realized through a mode of combining channel attention and frequency domain-space collaborative attention. Therefore, the YOLOv11 convolutional neural network detection model can identify and position a target more accurately when processing a complex scene, a shielding object and a small target, and the overall precision of target detection is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a target detection method based on improved YOLOv11. Background Art

[0002] Object detection is an important task in the field of computer vision, which aims to identify specific objects in images or videos and determine their locations. In recent years, with the rapid development of deep learning technology, object detection algorithms based on convolutional neural networks (CNNs) have made significant progress. Among them, the YOLO (You Only Look Once) series of algorithms are widely used in various scenarios due to their high efficiency and real-time performance.

[0003] YOLOv11 is a version of YOLO released in September 2024. It improves the accuracy and speed of the algorithm's target detection compared to previous YOLO versions. However, although YOLOv11 performs well on multiple benchmark datasets in the field of image detection, it still has technical problems such as long detection cycles and low detection accuracy in target detection scenarios such as complex backgrounds, occlusions, and small target detection. These problems are mainly due to the traditional convolutional network's lack of attention to important information when extracting features, resulting in reduced robustness of the model when processing complex scenes. Summary of the invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a target detection method based on improved YOLOv11, by designing a dual attention mechanism module embedded in the Bottleneck module of the C3K2 layer of the YOLOv11 convolutional neural network detection model, and by combining frequency-space collaborative attention and channel attention, the model's ability to focus on important features is enhanced, thereby improving the accuracy and robustness of target detection. The technical solutions provided by the present invention are as follows:

[0005] According to one aspect of an embodiment of the present invention, a target detection method based on improved YOLOv11 is provided, characterized in that the method includes:

[0006] S10: Obtain a VOC image dataset corresponding to the target detection object;

[0007] S20: After marking the target features of each target detection object in the VOC image dataset, the dataset is divided into a training set, a validation set, and a test set according to a preset ratio;

[0008] S30: construct and initialize an improved YOLOv11 convolutional neural network detection model, train and optimize the improved YOLOv11 convolutional neural network detection model according to the training set, the validation set and the test set to obtain a target detection neural network model; the Bottleneck module of the C3K2 layer of the improved YOLOv11 convolutional neural network detection model is embedded with a dual attention mechanism module, and the dual attention mechanism module is composed of a parallel channel attention module and a frequency domain-space collaborative attention module;

[0009] The channel attention module is used to obtain channel feature data corresponding to the feature image data. The channel attention graph expression of the channel attention module is as shown in formula (1):

[0010]

[0011] In formula (1), F1 represents the output channel feature data, X represents the input feature image data, Represents element-wise multiplication;

[0012] The frequency domain-spatial collaborative attention module is used to obtain the frequency domain and spatial domain fusion feature data corresponding to the feature image data. The frequency domain-spatial collaborative attention map is expressed as formula (2):

[0013]

[0014] In formula (2), F2 represents the output frequency domain and spatial domain fusion feature data, X represents the input feature image data, Represents element-wise multiplication;

[0015] The output of the dual attention mechanism module is obtained by fusing the channel feature data output by the channel attention module and the frequency domain-spatial feature data output by the frequency domain-spatial collaborative attention module according to a preset weight. The expression of the dual attention mechanism module is as shown in formula (3):

[0016] F DAM =αF1+(1-α)F2 Formula (3)

[0017] In formula (3), F DAM represents the output fusion feature data, and α represents the learnable parameter;

[0018] S40: Using the trained target detection neural network model to detect and identify the image data corresponding to the target detection object.

[0019] Preferably, the Bottleneck module in S30 is composed of two convolutional layers and a residual connection. The image data x input to the Bottleneck module is firstly executed on two convolutional layers to obtain an output F(x). Then, the output F(x) is used as the feature image data X input to the dual attention mechanism module. After being processed by the dual attention mechanism module, the fused feature data F is obtained. DAM , and finally the fusion feature data F DAM Perform a residual connection with the image data x to obtain the output F of the joint feature acquisition module DA-Bottleneck composed of the Bottleneck module and the dual attention mechanism module res (x), the expression of the joint feature acquisition module DA-Bottleneck is as shown in formula (4):

[0020] F res (x) = x + F DAM (X) Formula (4).

[0021] Preferably, the implementation method of the channel attention module in S30 includes:

[0022] S30a: Perform global average pooling on the input feature image data X to obtain a global statistical vector z c , the global average pooling expression is as follows:

[0023]

[0024] In formula (5), X c represents the characteristic image data output by the cth channel, and H×W represents the characteristic image data X c The image size, i and j represent the feature image data X c The horizontal and vertical coordinates of each pixel point on the global statistical vector z c The size of is c×1×1;

[0025] S30b: After the number of input channels is expanded to twice the original number of input channels in the first convolutional layer, the global statistical vector z c After the first convolutional layer is processed, the first feature data y1 is obtained. The expression of the first feature data y1 is as shown in formula (6):

[0026] y1=ReLU6(W1z c ) Formula (6)

[0027] In formula (6), W1∈R C×2C , W1 is a one-dimensional vector storing the connection weights of the convolutional layer, which is used for global feature extraction and connecting neurons between different layers;

[0028] S30c: Divide the first feature data y1 into two parts, y2 and y3, wherein y2 is nonlinearly transformed by the Relu6 activation function, and y3 maintains the original linear mapping, and then multiply y2 and y3 by element-by-element multiplication to obtain the second feature data z1. The division process of y2 and y3 is expressed as formula (7), and the expression of the second feature data z1 is expressed as formula (8):

[0029] y2=y1[:,:C],y3=y1[:,C:] Formula (7)

[0030]

[0031] In formula (8), y2∈R B×C ,y3∈R B×C ;

[0032] S30d: The second feature data z1 is processed by a fully connected layer to obtain a channel attention weight z2, and the feature image data X is weighted by the channel attention weight z2 to obtain the channel feature data F1 corresponding to the feature image data X. The expression of the channel attention weight z2 is as shown in formula (9), and the expression of the channel feature data F1 is as shown in formula (10):

[0033] z2=W2z1 Formula (9)

[0034] F1=σ(z2)*X Formula (10)

[0035] In formula (9), W2∈RC ×C , W2 is a one-dimensional vector storing the connection weights of the fully connected layer, which is used for global feature combination and connects all neurons in the previous and next layers;

[0036] In formula (10), σ is the Sigmoid nonlinear activation operation.

[0037] Preferably, the implementation method of the frequency-domain-spatial collaborative attention module in S30 includes:

[0038] S30e: Perform average pooling and maximum pooling on the input feature image data X to obtain average pooling feature data And the maximum pooling feature data The average pooled feature data And the maximum pooling feature data Through the standard convolutional layer, concatenation and convolution are performed to obtain 2D spatial feature data A, and the expression of the 2D spatial feature data A is as shown in formula (11):

[0039]

[0040] In formula (11),

[0041] S30f: Perform Sigmoid nonlinear activation on the 2D spatial feature data A to obtain the spatial attention weight A s , the spatial attention weight A s The expression of is as follows:

[0042] A s =σ(A) Formula (12)

[0043] In formula (12), σ is the Sigmoid nonlinear activation operation.

[0044] S30g: Perform a two-dimensional fast Fourier transform on the feature image data X to obtain frequency domain feature data F(X)(u,v). The expression of the two-dimensional fast Fourier transform is as shown in formula (13):

[0045]

[0046] In formula (13), the frequency domain feature data F(X)(u,v) contains the information of the feature image data X at different frequencies, u and v are frequency coordinates, Y(x,y) is the spatial domain feature map, x and y are spatial coordinates, j is an imaginary unit, represents a complex exponential function;

[0047] S30h: Obtain the real part feature F of the frequency domain feature data F(X)(u,v) R and the imaginary characteristic F I , the real feature F is converted into R and the imaginary characteristic F I Concatenate to get the joint feature F C , and then the convolutional layer is used to combine the features F C Perform position encoding and combine it with the joint feature F in the form of residual C Merge to get fusion features The joint feature F C The expression of is as shown in formula (14), the fusion feature The expression of is as follows:

[0048] F C =cat(F R , F I ) Formula (14)

[0049]

[0050] In formulas (14) and (15), F R ∈R H×W×C , F I∈R H×W×C , F C ∈R H×W×2C ,

[0051] S30i: Fusion features corresponding to different frequency components Perform normalization to obtain the normalized frequency domain attention weight A f , the frequency domain attention weight A f The expression of is as follows:

[0052]

[0053] In formula (16), the denominator represents the sum of the spectral energy of each channel in the spatial dimension;

[0054] S30j: The frequency domain attention weight A f With the spatial attention weight A s Multiply them together to perform multi-dimensional feature fusion, and perform weighted processing with the feature image data X to obtain frequency domain and spatial domain fusion feature data F2. The expression of the frequency domain and spatial domain fusion feature data F2 is as shown in formula (17):

[0055] F2=X*Sigmoid(A f *A s ) formula (17).

[0056] Preferably, the initial hyperparameters used in the improved YOLOv11 convolutional neural network detection model include: image size imgsz=640×640, number of training rounds epochs=300, batch sample size batchsize=16, initial learning rate Init_lr=0.01, minimum learning rate Min_lr=0.0001, cosine learning rate scheduling, SGD as the optimizer, and Mosaic as the data enhancement method.

[0057] Compared with the prior art, the target detection method based on improved YOLOv11 provided by the present invention has the following advantages:

[0058] The present invention provides a target detection method based on improved YOLOv11, which embeds a dual attention mechanism module in the Bottleneck module of the C3K2 layer of the YOLOv11 convolutional neural network detection model and trains it to obtain a target detection neural network model. The dual attention mechanism module consists of a parallel channel attention module and a frequency domain-space collaborative attention module, so that multi-dimensional feature extraction of the target image is jointly realized by combining channel attention with frequency domain-space collaborative attention, so that the YOLOv11 convolutional neural network detection model can more accurately identify and locate targets when processing complex scenes, occluders and small targets, thereby significantly improving the overall accuracy of target detection.

[0059] Furthermore, the target detection neural network model provided by the present invention has a small amount of model parameters and calculations added, so that the target detection neural network model can maintain a high operating efficiency while improving the detection performance, and is suitable for real-time target detection applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0061] Figure 1 It is a method flow chart of a target detection method based on improved YOLOv11 according to an exemplary embodiment of the present invention.

[0062] Figure 2 It is a structural diagram of a dual attention mechanism module according to an exemplary embodiment of the present invention.

[0063] Figure 3 It is a structural schematic diagram of a Bottleneck module of a C3K2 layer of an improved YOLOv11 convolutional neural network detection model according to an exemplary embodiment of the present invention.

[0064] Figure 4 is a structural diagram of a channel attention module according to an exemplary embodiment of the present invention.

[0065] Figure 5 is a data processing diagram of a channel attention module according to an exemplary embodiment of the present invention.

[0066] Figure 6 It is a structural diagram of a frequency-domain-spatial collaborative attention module according to an exemplary embodiment of the present invention.

[0067] Figure 7 It is a schematic diagram of the structure of a target detection neural network model according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0068] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is described in detail below in combination with specific embodiments (but not limited to the embodiments) and the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0069] Figure 1 is a method flow chart of a target detection method based on improved YOLOv11 according to an exemplary embodiment of the present invention. Figure 1 As shown, the target detection method based on improved YOLOv11 includes:

[0070] A target detection method based on improved YOLOv11, characterized in that the method includes:

[0071] S10: Obtain a VOC image dataset corresponding to the target detection object.

[0072] S20: After marking the target features of each target detection object in the VOC image dataset, the dataset is divided into a training set, a validation set, and a test set according to a preset ratio.

[0073] S30: Construct and initialize an improved YOLOv11 convolutional neural network detection model, train and optimize the improved YOLOv11 convolutional neural network detection model according to the training set, the validation set and the test set to obtain a target detection neural network model; the Bottleneck module of the C3K2 layer of the improved YOLOv11 convolutional neural network detection model is embedded with a dual attention mechanism module, and the dual attention mechanism module is composed of a parallel channel attention module and a frequency domain-space collaborative attention module.

[0074] The channel attention module is used to obtain channel feature data corresponding to the feature image data. The channel attention graph expression of the channel attention module is as shown in formula (1):

[0075]

[0076] In formula (1), F1 represents the output channel feature data, X represents the input feature image data, Represents element-wise multiplication;

[0077] The frequency domain-spatial collaborative attention module is used to obtain the frequency domain and spatial domain fusion feature data corresponding to the feature image data. The frequency domain-spatial collaborative attention map is expressed as formula (2):

[0078]

[0079] In formula (2), F2 represents the output frequency domain and spatial domain fusion feature data, X represents the input feature image data, Represents element-wise multiplication;

[0080] The output of the dual attention mechanism module is obtained by fusing the channel feature data output by the channel attention module and the frequency domain-spatial feature data output by the frequency domain-spatial collaborative attention module according to a preset weight. The expression of the dual attention mechanism module is as shown in formula (3):

[0081] F DAM =αF1+(1-α)F2 Formula (3)

[0082] In formula (3), F DAM represents the output fusion feature data, and α represents the learnable parameter.

[0083] It should be noted that, in an embodiment of the present invention, the channel attention module adaptively adjusts the channel weights of the feature map to highlight important channel information. By paying attention to the importance of different channels, it can improve the sensitivity of the model to key features; while the frequency-space collaborative attention module focuses on the spatial position in the feature map and optimizes the identification of important areas. It helps the model better capture spatial information by emphasizing important areas in the feature map. During the above-mentioned element-by-element multiplication operation, the attention values ​​will be broadcast accordingly, wherein the channel attention values ​​are broadcast in the spatial dimension, and the frequency-space collaborative attention values ​​are broadcast in the channel dimension. In this way, the feature map is optimized in both the channel and spatial dimensions. Finally, a learnable parameter α is introduced to dynamically adjust their weights, thereby achieving effective fusion of input features, thereby improving the performance and accuracy of the model, especially in complex tasks.

[0084] Further, a schematic diagram of the structure of a dual attention mechanism module according to an embodiment of the present invention is shown as follows: Figure 2 As shown, in Figure 2 In the figure, A is the channel attention module branch in the dual attention mechanism module, and B is the frequency domain-spatial collaborative attention module branch in the dual attention mechanism module.

[0085] S40: Using the trained target detection neural network model to detect and identify the image data corresponding to the target detection object.

[0086] Preferably, the Bottleneck module in S30 is composed of two convolutional layers and a residual connection. The image data x input to the Bottleneck module is firstly executed on two convolutional layers to obtain an output F(x). Then, the output F(x) is used as the feature image data X input to the dual attention mechanism module. After being processed by the dual attention mechanism module, the fused feature data F is obtained. DAM , and finally the fusion feature data F DAM Perform a residual connection with the image data x to obtain the output F of the joint feature acquisition module DA-Bottleneck composed of the Bottleneck module and the dual attention mechanism module res (x), the expression of the joint feature acquisition module DA-Bottleneck is as shown in formula (4):

[0087] F res (x) = x + F DAM (X) Formula (4).

[0088] Further, a schematic diagram of the structure of the Bottleneck module of the C3K2 layer of an improved YOLOv11 convolutional neural network detection model provided by an embodiment of the present invention is shown as follows: Figure 3 As shown, in Figure 3 In (a), Conv is a convolutional layer, x is the image data input to the Bottleneck module, and F(x) is the feature image data output after being processed by the first two convolutional layers of the Bottleneck module; Figure 3 In (b), A is the channel attention module in the dual attention mechanism module, B is the frequency domain-spatial collaborative attention module in the dual attention mechanism module, X is the feature image data input to the dual attention mechanism module, and F res (x) is the output of the joint feature acquisition module DA-Bottleneck.

[0089] Preferably, the implementation method of the channel attention module in S30 includes:

[0090] S30a: Perform global average pooling on the input feature image data X to obtain a global statistical vector z c , the global average pooling expression is as follows:

[0091]

[0092] In formula (5), X c represents the characteristic image data output by the cth channel, and H×W represents the characteristic image data X c The image size, i and j represent the feature image data X cThe horizontal and vertical coordinates of each pixel point on the global statistical vector z c The size of is c×1×1.

[0093] S30b: After the number of input channels is expanded to twice the original number of input channels in the first convolutional layer, the global statistical vector z c After the first convolutional layer is processed, the first feature data y1 is obtained. The expression of the first feature data y1 is as shown in formula (6):

[0094] y1=ReLU6(W1z c ) Formula (6)

[0095] In formula (6), W1∈R C×2C , W1 is a one-dimensional vector storing the connection weights of the convolutional layer, which is used for global feature extraction and connecting neurons between different layers.

[0096] Expanding the number of input channels to twice the original number of channels allows the model to learn richer feature representations and thus capture more complex feature relationships.

[0097] S30c: Divide the first feature data y1 into two parts, y2 and y3, wherein y2 is nonlinearly transformed by the Relu6 activation function, and y3 maintains the original linear mapping, and then multiply y2 and y3 by element-by-element multiplication to obtain the second feature data z1. The division process of y2 and y3 is expressed as formula (7), and the expression of the second feature data z1 is expressed as formula (8):

[0098] y2=y1[:,:C],y3=y1[:,C:] Formula (7)

[0099]

[0100] In formula (8), y2∈R B×C ,y3∈R B×C .

[0101] The step design of S30c enables the model to obtain a high-dimensional nonlinear feature space from a low-dimensional input.

[0102] S30d: The second feature data z1 is processed by a fully connected layer to obtain a channel attention weight z2, and the feature image data X is weighted by the channel attention weight z2 to obtain the channel feature data F1 corresponding to the feature image data X. The expression of the channel attention weight z2 is as shown in formula (9), and the expression of the channel feature data F1 is as shown in formula (10):

[0103] z2=W2z1 Formula (9)

[0104] F1=σ(z2)*X Formula (10)

[0105] In formula (9), W2∈R C×C , W2 is a one-dimensional vector storing the connection weights of the fully connected layer, which is used for global feature combination and connects all neurons in the previous and next layers;

[0106] In formula (10), σ is the Sigmoid nonlinear activation operation.

[0107] The step design of S30d can enhance the features of important channels and suppress unimportant channels. As the training and learning progress, the model can adaptively adjust to ignore less important channels and emphasize important channels.

[0108] Further, a schematic diagram of the structure of a channel attention module provided by an embodiment of the present invention is shown as follows: Figure 4 As shown, a data processing schematic diagram of a channel attention module provided by an embodiment of the present invention is shown as follows Figure 5 shown.

[0109] Preferably, the implementation method of the frequency-domain-spatial collaborative attention module in S30 includes:

[0110] S30e: Perform average pooling and maximum pooling on the input feature image data X to obtain average pooling feature data And the maximum pooling feature data The average pooled feature data And the maximum pooling feature data Through the standard convolutional layer, concatenation and convolution are performed to obtain 2D spatial feature data A, and the expression of the 2D spatial feature data A is as shown in formula (11):

[0111]

[0112] In formula (11),

[0113] S30f: Perform Sigmoid nonlinear activation on the 2D spatial feature data A to obtain the spatial attention weight A s , the spatial attention weight A s The expression of is as follows:

[0114] A s =σ(A) Formula (12)

[0115] In formula (12), σ is the Sigmoid nonlinear activation operation.

[0116] It should be noted that S30e to S30f are the implementation steps for spatial attention acquisition of the frequency domain-spatial collaborative attention module, and the following S30g to S30i are the implementation steps for frequency domain attention acquisition of the frequency domain-spatial collaborative attention module.

[0117] S30g: Perform a two-dimensional fast Fourier transform on the feature image data X to obtain frequency domain feature data F(X)(u,v). The expression of the two-dimensional fast Fourier transform is as shown in formula (13):

[0118]

[0119] In formula (13), the frequency domain feature data F(X)(u,v) contains the information of the feature image data X at different frequencies, u and v are frequency coordinates, Y(x,y) is the spatial domain feature map, x and y are spatial coordinates, j is an imaginary unit, represents a complex exponential function;

[0120] S30h: Obtain the real part feature F of the frequency domain feature data F(X)(u,v) R and the imaginary characteristic F I , the real feature F is converted into R and the imaginary characteristic F I Concatenate to get the joint feature F C , and then the convolutional layer is used to combine the features F C Perform position encoding and combine it with the joint feature F in the form of residual C Merge to get fusion features The joint feature F C The expression of is as shown in formula (14), the fusion feature The expression of is as follows:

[0121] F C =cat(F R , F I ) Formula (14)

[0122]

[0123] In formulas (14) and (15), F R ∈R H×W×C , F I ∈R H×W×C , F C ∈R H×W×2C ,

[0124] S30i: Fusion features corresponding to different frequency components Perform normalization to obtain the normalized frequency domain attention weight A f, frequency domain attention weight A f The expression of is as follows:

[0125]

[0126] In formula (16), the denominator represents the sum of the spectral energy of each channel in the spatial dimension.

[0127] The fusion features corresponding to different frequency components Normalization ensures that each frequency component has a relatively fair influence in the attention calculation, which avoids the bias caused by energy differences.

[0128] S30j: The frequency domain attention weight A f With the spatial attention weight A s Multiply them together to perform multi-dimensional feature fusion, and perform weighted processing with the feature image data X to obtain frequency domain and spatial domain fusion feature data F2. The expression of the frequency domain and spatial domain fusion feature data F2 is as shown in formula (17):

[0129] F2=X*Sigmoid(A f *A s ) formula (17).

[0130] The design of S30j enables the important frequency components and spatial region features in the feature image data to be highlighted. This combination enhances feature selectivity and ensures that only features that are considered important in both the frequency domain and space are emphasized, thereby improving the performance and denoising capabilities of the model. Through this weighting mechanism, the model can extract and utilize key features more effectively.

[0131] Further, a schematic diagram of the structure of a frequency domain-spatial collaborative attention module provided by an embodiment of the present invention is shown as follows: Figure 6 shown.

[0132] The initial hyperparameters used in the improved YOLOv11 convolutional neural network detection model include: image size imgsz=640×640, number of training rounds epochs=300, batch sample size batchsize=16, initial learning rate Init_lr=0.01, minimum learning rate Min_lr=0.0001, cosine learning rate scheduling, SGD as optimizer, and Mosaic as data enhancement method.

[0133] Further, a schematic diagram of the structure of a target detection neural network model provided by an embodiment of the present invention is shown as follows: Figure 7 As shown, in Figure 7In the embodiment of the present invention, DA-Bottleneck is the Bottleneck module embedded with the dual attention mechanism module, that is, the above-mentioned joint feature acquisition module.

[0134] Compared with the prior art, the target detection method based on improved YOLOv11 provided by the present invention has the following advantages:

[0135] The present invention provides a target detection method based on improved YOLOv11, which embeds a dual attention mechanism module in the Bottleneck module of the C3K2 layer of the YOLOv11 convolutional neural network detection model and trains it to obtain a target detection neural network model. The dual attention mechanism module consists of a parallel channel attention module and a frequency domain-space collaborative attention module, so that multi-dimensional feature extraction of the target image is jointly realized by combining channel attention with frequency domain-space collaborative attention, so that the YOLOv11 convolutional neural network detection model can more accurately identify and locate targets when processing complex scenes, occluders and small targets, thereby significantly improving the overall accuracy of target detection.

[0136] Furthermore, the target detection neural network model provided by the present invention has a small amount of model parameters and calculations added, so that the target detection neural network model can maintain a high operating efficiency while improving the detection performance, and is suitable for real-time target detection applications.

[0137] Although the present invention has been described in detail above by means of general description, specific implementation methods and tests, it is obvious to those skilled in the art that modifications or improvements can be made based on the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection claimed by the present invention.

[0138] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention herein. The present invention is intended to cover any variations, uses or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art that are not disclosed by the present invention. It should be understood that the present invention is not limited to the precise structure that has been described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from the scope thereof.

Claims

1. A target detection method based on improved YOLOv11, characterized in that: The method comprises: S10: Obtain a VOC image dataset corresponding to the target detection object; S20: After marking the target features of each target detection object in the VOC image dataset, the dataset is divided into a training set, a validation set, and a test set according to a preset ratio; S30: construct and initialize an improved YOLOv11 convolutional neural network detection model, train and optimize the improved YOLOv11 convolutional neural network detection model according to the training set, the validation set and the test set to obtain a target detection neural network model; the Bottleneck module of the C3K2 layer of the improved YOLOv11 convolutional neural network detection model is embedded with a dual attention mechanism module, and the dual attention mechanism module is composed of a parallel channel attention module and a frequency domain-space collaborative attention module; The channel attention module is used to obtain channel feature data corresponding to the feature image data. The channel attention graph expression of the channel attention module is as shown in formula (1): In formula (1), F1 represents the output channel feature data, X represents the input feature image data, Represents element-wise multiplication; The frequency domain-spatial collaborative attention module is used to obtain the frequency domain and spatial domain fusion feature data corresponding to the feature image data. The frequency domain-spatial collaborative attention map is expressed as formula (2): In formula (2), F2 represents the output frequency domain and spatial domain fusion feature data, X represents the input feature image data, Represents element-wise multiplication; The output of the dual attention mechanism module is obtained by fusing the channel feature data output by the channel attention module and the frequency domain-spatial feature data output by the frequency domain-spatial collaborative attention module according to a preset weight. The expression of the dual attention mechanism module is as shown in formula (3): F DAM =αF1+(1-α)F2 Formula (3) In formula (3), F DAM represents the output fusion feature data, and α represents the learnable parameter; S40: Using the trained target detection neural network model to detect and identify the image data corresponding to the target detection object.

2. The target detection method according to claim 1, characterized in that: The Bottleneck module in S30 is composed of two convolutional layers and a residual connection. The image data x input to the Bottleneck module is firstly executed on two convolutional layers to obtain an output F(x). Then, the output F(x) is used as the feature image data X input to the dual attention mechanism module. After being processed by the dual attention mechanism module, the fused feature data F is obtained. DAM , and finally the fusion feature data F DAM Perform a residual connection with the image data x to obtain the output F of the joint feature acquisition module DA-Bottleneck composed of the Bottleneck module and the dual attention mechanism module res (x), the expression of the joint feature acquisition module DA-Bottleneck is as shown in formula (4): F res (x) = x + F DAM (X) Formula (4).

3. The method according to claim 1, characterized in that The implementation method of the channel attention module in S30 includes: S30a: Perform global average pooling on the input feature image data X to obtain a global statistical vector z c , the global average pooling expression is as follows: In formula (5), X c represents the characteristic image data output by the cth channel, and H×W represents the characteristic image data X c The image size, i and j represent the feature image data X c The horizontal and vertical coordinates of each pixel point on the global statistical vector z c The size of is c×1×1; S30b: After the number of input channels is expanded to twice the original number of input channels in the first convolutional layer, the global statistical vector z c After the first convolutional layer is processed, the first feature data y1 is obtained. The expression of the first feature data y1 is as shown in formula (6): y1=ReLU6(W1z c ) Formula(6) In formula (6), W1∈R C×2C , W1 is a one-dimensional vector storing the connection weights of the convolutional layer, which is used for global feature extraction and connecting neurons between different layers; S30c: Divide the first feature data y1 into two parts, y2 and y3, wherein y2 is nonlinearly transformed by the Relu6 activation function, and y3 maintains the original linear mapping, and then multiply y2 and y3 by element-by-element multiplication to obtain the second feature data z1. The division process of y2 and y3 is expressed as formula (7), and the expression of the second feature data z1 is expressed as formula (8): y2=y1[:,:C],y3=y1[:,C:] Formula (7) In formula (8), y2∈R B×C ,y3∈R B×C ; S30d: The second feature data z1 is processed by a fully connected layer to obtain a channel attention weight z2, and the feature image data X is weighted by the channel attention weight z2 to obtain the channel feature data F1 corresponding to the feature image data X. The expression of the channel attention weight z2 is as shown in formula (9), and the expression of the channel feature data F1 is as shown in formula (10): z2=W2z1 Formula (9) F1=σ(z2)*X Formula (10) In formula (9), W2∈R C×C , W2 is a one-dimensional vector storing the connection weights of the fully connected layer, which is used for global feature combination and connects all neurons in the previous and next layers; In formula (10), σ is the Sigmoid nonlinear activation operation.

4. The method according to claim 1, characterized in that The implementation method of the frequency-space collaborative attention module in S30 includes: S30e: Perform average pooling and maximum pooling on the input feature image data X to obtain average pooling feature data And the maximum pooling feature data The average pooled feature data And the maximum pooling feature data Through the standard convolutional layer, concatenation and convolution are performed to obtain 2D spatial feature data A, and the expression of the 2D spatial feature data A is as shown in formula (11): Formula (11) S30f: Perform Sigmoid nonlinear activation on the 2D spatial feature data A to obtain the spatial attention weight A s , the spatial attention weight A s The expression of is as follows: A s = σ(A) Equation (12) In formula (12), σ is the Sigmoid nonlinear activation operation. S30g: Perform a two-dimensional fast Fourier transform on the feature image data X to obtain frequency domain feature data F(X)(u, v). The expression of the two-dimensional fast Fourier transform is as shown in formula (13): In formula (13), the frequency domain feature data F(X)(u, v) contains the information of the feature image data X at different frequencies, u and v are frequency coordinates, Y(x, y) is the spatial domain feature map, x and y are spatial coordinates, j is an imaginary unit, represents a complex exponential function; S30h: Obtain the real part feature F of the frequency domain feature data F(X)(u, v) R and the imaginary characteristic F I , the real feature F is converted into R and the imaginary characteristic F I Concatenate to get the joint feature F C , and then the convolutional layer is used to combine the features F C Perform position encoding and combine it with the joint feature F in the form of residual C Merge to get fusion features The joint feature F C The expression of is as shown in formula (14), the fusion feature The expression of is as follows: F C =cat(F R , F I ) Formula (14) In formulas (14) and (15), F R ∈R H×W×C , F I ∈R H×W×C , F C ∈R H×W×2C , S30i: Fusion features corresponding to different frequency components Perform normalization to obtain the normalized frequency domain attention weight A f , the frequency domain attention weight A f The expression of is as follows: In formula (16), the denominator represents the sum of the spectral energy of each channel in the spatial dimension; S30j: The frequency domain attention weight A f With the spatial attention weight A s Multiply them together to perform multi-dimensional feature fusion, and perform weighted processing with the feature image data X to obtain frequency domain and spatial domain fusion feature data F2. The expression of the frequency domain and spatial domain fusion feature data F2 is as shown in formula (17): F2=X*Sigmoid(A f *A s ) formula (17).

5. The method according to claim 1, characterized in that The initial hyperparameters used in the improved YOLOv11 convolutional neural network detection model include: image size imgsz=640×640, number of training rounds epochs=300, batch sample size batchsize=16, initial learning rate Init_lr=0.01, minimum learning rate Min_lr=0.0001, cosine learning rate scheduling, SGD as optimizer, and Mosaic as data enhancement method.

Citation Information

Cited By

  • Photovoltaic panel dust retention detection method and device, electronic equipment and storage medium

    CN120298402A

  • Photovoltaic panel dust detection method, device, electronic device and storage medium

    CN120298402B