Target detection method based on space-frequency attention and channel transposition attention mechanism
By introducing an improved method of space-frequency attention and channel transpose attention mechanism into the YOLO object detection framework, the problem of low object detection accuracy in complex industrial images is solved, and higher detection accuracy and small object detection accuracy are achieved.
Patent Information
- Application Number
- CN202510143659.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
Existing object detection technology is difficult to effectively improve the accuracy of object detection in complex industrial images, especially when processing large objects, small objects and fuzzy imaging.
The improvement method based on the space-frequency attention and channel transpose attention mechanism is adopted to improve the YOLO object detection framework. Through image cropping, target annotation and data expansion processing, combined with the space-frequency attention mechanism and channel transpose attention mechanism, a joint attention module is built, and the C2F module in YOLOv8 is replaced to form an improved object detection model.
The accuracy of object detection and small object detection accuracy are improved, the model's adaptability to complex industrial images and the target positioning accuracy are enhanced, and false detection and missed detection are reduced.
Smart Images

Figure CN120070981A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent target detection, and particularly relates to an object detection method based on spatial-frequency attention and channel transposed attention mechanism. Background Art
[0002] In recent years, object detection technology based on deep learning has developed rapidly, especially the currently most popular and continuously iterated YOLO series of object detection methods. The YOLO series of detection algorithms are widely used in the field of industrial recognition and have achieved good implementation results. Among them, the quality of industrial images also directly affects the recognition effect of object detection algorithms. However, due to the complex industrial scenarios, there will be large targets, small targets, and problems such as blurred imaging caused by other reasons. Therefore, it is very important to improve the accuracy of detected targets.
[0003] Existing object detection frameworks and their improved variants mainly focus on how to improve the feature extraction ability of the algorithm itself and how to avoid losing more information during network downsampling. They do not start from the data itself to solve the problem of how to better improve the performance of object detection algorithms.
[0004] In view of the problems in the prior art, there is an urgent need to propose an object detection method based on spatial-frequency attention and channel transposed attention mechanism. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes an object detection method based on spatial-frequency attention and channel transposed attention mechanism, which is used to improve the recognition accuracy of object detection and is applied to tasks including but not limited to industrial defect detection, object recognition, and object classification, etc., to solve the problems existing in the above prior art.
[0006] To achieve the above object, the present invention provides an object detection method based on spatial-frequency attention and channel transposed attention mechanism, including the following steps:
[0007] Obtain an industrial image, and perform image cropping, target annotation, and data augmentation processing on the industrial image;
[0008] Improve the YOLO object detection framework based on the spatial-frequency attention mechanism and the channel transposed attention mechanism to obtain an improved object detection model;
[0009] Train the improved object detection model based on the processed industrial image to obtain a trained object detection model;
[0010] Perform object detection on the industrial image obtained in real time based on the trained object detection model.
[0011] Optionally, the process of cropping the industrial image includes: presetting a repetition ratio and cropping the industrial image based on the repetition ratio.
[0012] Optionally, the open-source LabelImg annotation software is used for target annotation of the industrial image.
[0013] Optionally, the methods for data augmentation of the industrial image include, but are not limited to, HSV transformation, angular rotation transformation, translation transformation, shearing transformation, perspective transformation, flipping transformation, image cropping and splicing, Mixup, Copy-paste, and erasing.
[0014] Optionally, the process of improving the YOLO object detection framework based on the spatial-frequency attention mechanism and the channel transpose attention mechanism includes:
[0015] Combining spatial-frequency attention and channel transpose attention to form a joint attention module; replacing the C2F module in YOLOv8 with the joint attention module to obtain an improved object detection model.
[0016] Optionally, the process of performing object detection on a real-time acquired industrial image based on the trained object detection model includes:
[0017] Performing spatial-frequency attention and channel transpose attention calculations on the convolutional output of each layer based on the joint attention module to obtain corresponding attention features, then splicing the corresponding attention features in the feature dimension, and inputting the spliced features into the next network layer for object detection.
[0018] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the method.
[0019] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and characterized in that the computer program, when executed by a processor, implements the steps of the method.
[0020] The present invention also provides a computer program product, including a computer program, and characterized in that the computer program, when executed by a processor, implements the steps of the method.
[0021] Compared with the prior art, the present invention has the following advantages and technical effects:
[0022] The present invention processes the images used for model training using image cropping, target annotation, and data augmentation, enabling the model to have strong adaptability and be able to handle different industrial image scenarios;
[0023] The present invention improves the YOLO model based on the spatial-frequency attention mechanism and the channel transpose attention mechanism. Through the spatial-frequency attention mechanism, the model can better capture the spatial and frequency features in the image, which helps to improve the target recognition ability, especially in the context of complex industrial image backgrounds. The combined use of the spatial-frequency attention mechanism and the channel transpose attention mechanism helps to restore image details and improve the accuracy of blurred target detection. At the same time, while the present invention restores image details, it also better preserves the detail information of small targets, avoiding the loss of more information after convolution and downsampling, and improving the detection accuracy of small targets. This is particularly important for the detection of small targets in industrial images, which can improve the target positioning accuracy of the model and reduce false detections and missed detections.
[0024] The improved target detection model of the present invention can achieve fast and accurate target detection for real-time acquired industrial images, which is crucial for industrial application scenarios that require rapid response. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0026] Figure 1 is a schematic structural diagram of the spatial-frequency attention module according to an embodiment of the present invention;
[0027] Figure 2 is a schematic structural diagram of the channel transpose attention module according to an embodiment of the present invention;
[0028] Figure 3 is a schematic structural diagram of the improved target detection model according to an embodiment of the present invention;
[0029] Figure 4 is a schematic diagram of the training process data of YOLOv8 according to an embodiment of the present invention;
[0030] Figure 5 is a schematic diagram of the training process data of the improved target detection model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0032] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0033] Embodiment 1
[0034] In this embodiment, a target detection method based on spatial-frequency attention and channel transpose attention mechanism is provided, including the following steps:
[0035] Obtain an industrial image, and perform image cropping, target annotation, and data augmentation processing on the industrial image;
[0036] Improve the YOLO target detection framework based on the spatial-frequency attention mechanism and the channel transpose attention mechanism to obtain an improved target detection model;
[0037] Train the improved target detection model based on the processed industrial image to obtain a trained target detection model;
[0038] Perform target detection on the industrial image obtained in real time based on the trained target detection model.
[0039] As an implementable way, the process of image cropping the industrial image includes: Since the images in the actual industrial scenario are high-resolution images, they need to be cropped to improve the detection accuracy. When cropping, set a certain overlap ratio to avoid cropping the target defect onto two images.
[0040] As an implementable way, the process of target annotation for the industrial image includes: In target annotation, use the open-source annotation software LabelImg to annotate the detected targets, including categories and the positions of the detection frames.
[0041] As an implementable way, the process of data augmentation for the industrial image includes: Data augmentation refers to performing data transformation on the annotated data to generate defects of different sizes, different shapes. Such as HSV transformation, angle rotation transformation, translation transformation, shear transformation, perspective transformation, flip transformation, image cropping and splicing, Mixup, Copy-paste, erasing and other image transformations.
[0042] As an implementable way, the process of improving the YOLO target detection framework based on the spatial-frequency attention mechanism and the channel transpose attention mechanism to obtain an improved target detection model includes:
[0043] In the target detection network, two attention modules are introduced in this embodiment, namely the spatial-frequency attention module and the channel transposed attention module. The spatial-frequency attention mechanism (Spatial-Frequency Attention (SFA)), as shown in Figure 1 , has been effectively verified in image super-resolution restoration. It pays more attention to the high-frequency part of the image during network training, that is, the details such as the edges and textures of the image. The spatial-frequency attention mechanism integrates the feature high-frequency and channel information into the self-attention information to enhance the restoration of high-frequency information in the features and improve its resolution. The frequency-domain wavelet transform used in the spatial-frequency attention mechanism is DTCWT (Dual-Tree Complex Wavelet Transform). Based on this, this embodiment introduces the spatial-frequency attention mechanism to improve the recognition rate of fuzzy target detection.
[0044] The channel transposed attention module (Channel Transposed Attention (CTA)), as shown in Figure 2 , adopts a different strategy from the spatial-frequency attention mechanism. Its self-attention calculation is along the channel direction, and then the spatial and channel information is fused and enhanced. The channel transposed attention module is used as a supplement to the spatial-frequency attention mechanism, which can improve the ability of the spatial-frequency attention mechanism to a certain extent.
[0045] Based on the performance of the spatial-frequency attention mechanism and the channel transposed attention module in super-resolution image restoration, this embodiment embeds them into the YOLO target detection framework and proposes an improved target detection model based on the spatial-frequency attention mechanism and the channel transposed attention mechanism. Its network structure is as shown in Figure 3 :
[0046] In the improved network, this embodiment replaces the C2F module in YOLOv8 with a joint attention module of spatial-frequency attention and channel transposed attention. In the joint attention module, the spatial-frequency attention and channel transposed attention calculations are respectively performed on the output of each layer of convolution to obtain different attention features, and then they are concatenated in the feature dimension as the input of the next network layer.
[0047] As an implementable way, the process of training the improved target detection model includes: after the network model is constructed, the pictures after data augmentation are input into the improved target detection model to obtain the training model.
[0048] Furthermore, the cropped classification pictures are divided into datasets, including a training set, a test set, and a validation set, and their division ratios are 0.8, 0.1, and 0.1 respectively.
[0049] As an implementable approach, the process of performing object detection on real-time industrial images based on the trained object detection model includes: converting the model based on the ONNX platform, parsing the model using C#, deploying it to the Windows host computer platform, connecting it to the image acquisition interface, and performing real-time object recognition on the images.
[0050] As an implementable approach, in this embodiment, the improved object detection model was verified on the coco128 dataset. With the same parameters, it was trained for 1000 epochs respectively, and its performance is as follows:
[0051] Table 1 shows the average performance metrics for all classes. It can be seen from the table that the improved algorithm is superior to the YOLOv8 prototype in all four metrics. Figure 4 And Figure 5 They are the training process data of the YOLOv8 algorithm and the improved YOLOv8 algorithm of this embodiment respectively. It can be seen that the improved algorithm obtains the minimum loss value in each class of loss function. This verifies the effectiveness of the algorithm in this embodiment.
[0052] Table 1
[0053]
[0054]
[0055] Embodiment 2
[0056] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method.
[0057] Embodiment 3
[0058] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. It is characterized in that when the computer program is executed by a processor, it implements the steps of the method.
[0059] Embodiment 4
[0060] This embodiment also provides a computer program product, including a computer program. It is characterized in that when the computer program is executed by a processor, it implements the steps of the method.
[0061] The above are only the preferred specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A target detection method based on space-frequency attention and channel transposition attention mechanism, characterized in that: The following steps are involved: Acquire industrial images, and perform image cropping, object annotation, and data expansion on the industrial images; The YOLO target detection framework is improved based on the space-frequency attention mechanism and the channel transposition attention mechanism to obtain an improved target detection model; Training the improved target detection model based on the processed industrial image to obtain a trained target detection model; Perform target detection on industrial images acquired in real time based on the trained target detection model.
2. The method according to claim 1, characterized in that The process of cropping the industrial image includes: presetting a repetition ratio, and cropping the industrial image based on the repetition ratio.
3. The method according to claim 1, characterized in that The LabelImg open source labeling software is used to label the industrial images.
4. The method according to claim 1, characterized in that: Methods used to expand the data of the industrial image include but are not limited to HSV transformation, angle rotation transformation, translation transformation, shear transformation, perspective transformation, flip transformation, image cropping and splicing, Mixup, Copy-paste and erasure.
5. The method according to claim 1, characterized in that The process of improving the YOLO target detection framework based on the space-frequency attention mechanism and the channel transposition attention mechanism includes: The spatial-frequency attention and the channel transposed attention are combined to form a joint attention module; the C2F module in YOLOv8 is replaced with the joint attention module to obtain an improved target detection model.
6. The method according to claim 5, characterized in that The process of performing target detection on industrial images acquired in real time based on the trained target detection model includes: Based on the joint attention module, space-frequency attention and channel transposed attention are calculated for each layer of convolution output respectively to obtain corresponding attention features, and then the corresponding attention features are spliced in the feature dimension, and the spliced features are input into the next network layer for target detection.
7. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Thangka Buddha statue identification method based on key area destruction and reconstruction learning
CN122416004A