Lightweight traffic target detection method and device based on LM-YOLO and medium

By introducing the LM-YOLO network with linear deformable convolution and hybrid local channel attention mechanism, the problems of detection accuracy and computational efficiency in complex traffic scenarios are solved, and high-precision lightweight target detection is achieved, which is suitable for vehicle-mounted and edge computing devices.

CN120689676AActive Publication Date: 2025-09-23HEFEI KUANGHANG INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510835522.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-23
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing deep learning traffic target detection technology has problems such as poor deformation adaptability, computational redundancy of the attention mechanism, and the contradiction between lightweight and accuracy in complex traffic scenarios, making it difficult to achieve high-precision real-time detection.

Method used

The linear deformable convolution (LDConv) module and the hybrid local channel attention (MLCA) mechanism are used to construct the LM-YOLO network, which enhances the target detection capability in complex traffic scenes through dynamic sampling and local-global feature collaborative optimization.

Benefits of technology

While maintaining its lightweight, it improves detection accuracy and robustness in complex traffic scenarios and is suitable for vehicle-mounted terminals and edge computing devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689676A_ABST
    Figure CN120689676A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight traffic target detection method and device based on LM-YOLO, and a medium, and aims to solve the problems of poor geometric deformation adaptability, attention mechanism calculation redundancy, contradiction between lightweight and precision and the like of a conventional detection network in a complex traffic scene. According to the method, a linear Deformation Convolution (LDCON) module is introduced, through a coordinate generation algorithm and a linear parameter growth mechanism, dynamic sampling is supported while the parameter quantity is reduced, and target deformation features are effectively captured. Based on a mixed local channel attention (MLCA) mechanism, a C2fMLCA module is constructed, and through dual-channel attention weighting, the key area focusing capability of a multi-scale target is enhanced. A complementary enhancement mechanism is formed through hierarchical deployment of the LDConv and MLCA modules, the detection robustness of vehicles and pedestrians in a complex traffic scene can be effectively improved, and the method has the advantages of light weight and high precision and is suitable for deployment of a vehicle-mounted terminal and edge computing equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a lightweight traffic target detection method, device, and storage medium based on LM-YOLO. Background Art

[0002] Current deep learning-based traffic object detection technologies generally employ fixed-structure convolutional neural networks, which present significant limitations when dealing with complex traffic scenarios. Standard convolution operations, constrained by a fixed sampling grid and square convolution kernel design, struggle to effectively capture the deformation and occlusion characteristics of vehicles and pedestrians. Although deformable convolution improves deformation adaptability by introducing an offset learning mechanism, its number of parameters increases quadratically with the kernel size, resulting in a dramatic increase in computational resources. Furthermore, it only supports regular sampling shapes (e.g., 3×3, 5×5) and lacks the flexibility to adapt to the irregular contours of objects in traffic scenarios. Furthermore, existing attention mechanisms (e.g., SE and CBAM) often employ global channel compression or independent spatial weighting strategies, ignoring the correlation between local spatial and channel features. Consequently, they face the dual challenges of insufficient feature representation and computational redundancy in real-time traffic monitoring scenarios.

[0003] In terms of lightweight network architectures, mainstream methods use depthwise separable convolutions to reduce parameter counts. However, over-compressing feature dimensions can easily lead to loss of detailed information about small objects (such as distant pedestrians and traffic signs). For example, while networks like MobileNetV3 perform well in general object detection, their SE attention module only performs global channel weighting, making it difficult to capture key local areas in traffic scenes (such as edge features of occluded areas by vehicles). Furthermore, existing technologies inadequately optimize the multi-scale feature fusion process, making it prone to feature misalignment, leading to reduced positioning accuracy, particularly when dealing with high-speed moving objects. To address these issues, there is an urgent need to develop novel network architectures that balance deformation adaptability, local feature enhancement, and computational efficiency to meet the demands of intelligent transportation systems for high-precision, real-time object detection. Summary of the Invention

[0004] The present invention proposes a lightweight traffic target detection method, device, and storage medium based on LM-YOLO, which can solve at least one of the technical problems in the background technology, while maintaining low computational overhead of the model and improving detection accuracy in complex traffic scenarios.

[0005] To achieve the above object, the present invention adopts the following technical solutions: A lightweight traffic target detection method based on LM-YOLO, including: S1, obtain the public traffic scene dataset and preprocess it; S2: Build the LM-YOLO network model. The LM-YOLO network uses YOLOv8n as its basic architecture, adopts the LDConv module to replace the conventional convolution module to improve the model's feature extraction ability for complex targets, and builds the C2f_MLCA ​​module before the three-stage detection head to enhance the model's feature discrimination ability for multi-scale targets. S3, configure the training environment and use the data set obtained in step S1 to perform LM-YOLO network training and verification analysis; S4: Input the traffic scene image to be detected into the pre-trained LM-YOLO network model for target detection.

[0006] Furthermore, in step S1, the preprocessing of the public dataset includes random data extraction, training set and validation set division, and image and label format conversion to adapt to the YOLOv8 model input.

[0007] Furthermore, in step S2, the LM-YOLO network retains only the first layer of feature extraction and uses conventional convolution for preliminary feature mapping. Subsequent downsampling uses the LDConv module to replace conventional convolution operations. The LDConv module eliminates the network's reliance on a regular square grid, supports convolution kernel sampling shapes with any number of parameters, adjusts the parameter growth trend from quadratic to linear, and dynamically adjusts the offset to adapt to target changes. This reduces the number of parameters while improving feature extraction capabilities for complex targets, providing a lightweight option for scenarios with limited hardware resources.

[0008] Furthermore, the LDConv module specifically includes: According to the set convolution kernel parameter N, the initial sampling shape and relative coordinates are generated by the initial sampling coordinate generation algorithm; Generate the absolute coordinates of the initial sampling points based on the resolution and stride of the input feature map; The offset generation module performs convolution operation on the input feature map to generate coordinate dynamic offset; Combine the relative coordinates, absolute coordinates and dynamic offset of the initial sampling point to calculate the final sampling point coordinates; Based on the coordinates of the final sampling point, the bilinear interpolation method is used to extract the feature value of the corresponding position from the input feature map to obtain the offset resampled feature map; Reshape the resampled feature map so that it can be processed by the convolution kernel of the main convolution module; The main convolution module is used to perform a convolution operation on the reshaped feature map to obtain the final output feature map.

[0009] Furthermore, in step S2, the LM-YOLO network introduces the MLCA attention mechanism into the C2f module preceding the three-stage detection head, constructing the C2f_MLCA ​​module to enhance the feature representation capabilities of multi-scale traffic targets. The MLCA attention mechanism employs a dual-branch structure, combining local average pooling (LAP) with global average pooling (GAP) to simultaneously model channel, spatial, local, and global information. It also employs 1D convolution for acceleration, achieving an efficient balance between parameters and computational effort, making it suitable for lightweight object detection networks.

[0010] Furthermore, in step S2, the C2f_MLCA ​​module structure specifically includes: Pass the input feature map through a 1x1 convolutional layer to expand the number of channels; Divide the expanded feature map into two parts along the channel dimension; Pass one of the feature maps through Bottleneck modules, the output of each module is used as the input of the next module, and finally we get feature maps; All feature maps are spliced ​​along the channel dimension to obtain a spliced ​​feature map; The concatenated feature map is input into the 1x1 convolution layer so that the number of channels is mapped to the target output channel number.

[0011] Furthermore, in step S2, the MLCA module specifically includes: The input feature map is divided into local regions, calculate the average value of each region, and obtain the local feature map; In the global branch, GAP is performed on the local feature map to compress it into a global feature map. The global feature map is reshaped and input into a 1D convolution operation to obtain the global channel attention weight. Finally, the global feature map is expanded to size; In the local branch, the local feature map is reshaped, and the local channel attention weight is calculated by 1D convolution, and the result is reshaped back size; The local feature maps and global feature maps obtained by the two branches are normalized and added together, and then adjusted back to the input feature map size through UNAP to obtain the final attention weight; The obtained attention weights are multiplied element-wise with the original input feature map to obtain the final enhanced feature map output.

[0012] In another aspect, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.

[0013] On the other hand, the present invention further discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0014] As can be seen from the above technical solutions, the lightweight traffic target detection method based on LM-YOLO of the present invention addresses the problems faced by traditional detection networks in complex traffic scenarios, such as poor adaptability to geometric deformation, redundant attention mechanism calculations, and the contradiction between lightweightness and precision. This method introduces the Linear Deformable Convolution (LDConv) module. Through a coordinate generation algorithm and a linear parameter growth mechanism, it supports dynamic sampling while reducing the number of parameters, effectively capturing target deformation characteristics. Based on the Mixed Local Channel Attention (MLCA) mechanism, a C2f_MLCA ​​module is constructed, which enhances the ability to focus on key areas of multi-scale targets through dual-channel attention weighting. By layering the LDConv and MLCA modules to form a complementary enhancement mechanism, it can effectively improve the robustness of vehicle and pedestrian detection in complex traffic scenarios, combining the advantages of lightweightness and high precision, and is suitable for deployment in vehicle terminals and edge computing devices.

[0015] Specifically, the beneficial effects of the present invention are as follows: 1. The lightweight traffic target detection method based on LM-YOLO described in this paper achieves significant improvements in lightweightness, detection accuracy, and hardware adaptability compared to traditional target detection methods by introducing linear deformable convolution and hybrid local channel attention modules into the YOLOv8 algorithm.

[0016] 2. The linear deformable convolution described in the present invention uses a dynamic offset learning mechanism to enable the convolution kernel sampling points to adapt to the target deformation and spatial distribution characteristics, effectively improving the feature extraction capability of complex targets such as vehicle deformation and occlusion in traffic scenes. By breaking away from the fixed shape limitation of traditional convolution kernels and converting the parameter growth trend from a square to a linear growth characteristic, a flexible feature representation space is provided for traffic target detection at different scales while maintaining hardware computing efficiency.

[0017] 3. The hybrid local channel attention mechanism described in the present invention overcomes the spatial and local information loss problems of the traditional attention mechanism by processing local and global features in parallel. By adopting a single-dimensional convolution acceleration strategy, it achieves the coordinated optimization of local-global features while avoiding channel dimensionality reduction, enhances the ability to perceive details of traffic targets, suppresses the interference of complex background noise, and improves the ability to express multi-target features while keeping the model lightweight.

[0018] 4. The LDConv-MLCA joint architecture described in the present invention realizes complementary enhancement between modules through layered deployment, which is specifically manifested as follows: the LDConv module is adopted in the shallow feature extraction stage of the network, and the geometric deformation characteristics of traffic targets are dynamically captured through the linear parameter growth strategy and dynamic sampling mechanism, thereby retaining detailed information while reducing computational overhead; the C2f_MLCA ​​module is embedded in the deep semantic enhancement stage, and the key feature areas are focused on through the local-global attention collaborative mechanism to enhance the semantic expression ability of multi-scale targets; further, the dynamic sampling mechanism of LDConv and the cross-regional attention of MLCA form a spatial-channel dual-dimensional adaptive characteristic, jointly modeling the deformation characteristics and contextual associations of traffic targets, and constructing a complete optimization link from bottom-level deformation adaptation to high-level semantic enhancement, thereby improving the detection accuracy and generalization ability of complex traffic scenes while maintaining the low computational complexity of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a diagram of the lightweight traffic target detection network structure based on LM-YOLO described in the present invention; Figure 2 This is the structural diagram of the LDConv module of the present invention; Figure 3 This is a structural diagram of the C2f-MLCA module of the present invention; Figure 4 This is a structural diagram of the MLCA module described in the present invention. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0021] like Figure 1 As shown, the lightweight traffic target detection method based on LM-YOLO described in this embodiment includes the following steps: S1, obtaining a public traffic scene dataset and performing preprocessing; S2, build LM-YOLO network model, refer to Figure 1As shown in the figure, the LM-YOLO network uses YOLOv8n as the basic architecture, adopts the LDConv module to replace the conventional convolution to enhance the model's feature extraction ability for complex targets, and builds the C2f_MLCA ​​module before the three-stage detection head to enhance the model's feature discrimination ability for multi-scale targets; S3, configure the training environment and use the data set obtained in step S1 to perform LM-YOLO network training and verification analysis; S4: Input the traffic scene image to be detected into the pre-trained LM-YOLO network model for target detection.

[0022] Preferably, in step S1, the BDD100k dataset is preprocessed, 10,000 data are randomly extracted, and the training set and validation set are divided into a ratio of 9:1, and the image and label formats are converted to adapt to the YOLOv8 model input.

[0023] Preferably, in step S2, the LM-YOLO network retains only the first layer of feature extraction and uses conventional convolution for preliminary feature mapping. Subsequent downsampling uses the LDConv module to replace the conventional convolution operation. The LDConv module can break away from the network's dependence on a regular square grid, support convolution kernel sampling shapes with any number of parameters, adjust the parameter growth trend from square to linear, and combine it with dynamic adjustment of the offset to adapt to target changes. It can improve the feature extraction capability of complex targets while reducing the number of parameters, providing a lightweight option for scenarios with limited hardware resources.

[0024] Preferably, referring to FIG3 , the LDConv module specifically includes: According to the set convolution kernel parameter N, the initial sampling coordinate generation algorithm is used to generate the initial sampling shape and relative coordinates. , the initial coordinate generation algorithm first passes the parameter Generate a regular square sampling grid, then fill in the remaining points and splice them into a complete sampling coordinate matrix; Generate the absolute coordinates of the initial sampling points based on the resolution and stride of the input feature map ; The offset generation module performs convolution operation on the input feature map to generate dynamic coordinate offsets , the offset generation module is in the form of ,in Indicates that each sampling point has two coordinate offsets (horizontally and vertically); Combine the relative coordinates, absolute coordinates and dynamic offset of the initial sampling point to calculate the final sampling point coordinates ; Based on the coordinates of the final sampling point, the bilinear interpolation method is used to extract the eigenvalues ​​of the corresponding positions from the input feature map to obtain the resampled feature map after offset, with a dimension of ; The resampled feature map along Direction stacking, reshaped into , so that it can be processed by the convolution kernel of the main convolution module; The main convolution module is used to perform convolution operation on the reshaped feature map to obtain the final output feature map. The main convolution module includes a convolution layer, a batch normalization layer and a SiLU activation function. The convolution layer is in the form of , to adapt to the feature map after dynamic sampling.

[0025] Preferably, in step S2, the LM-YOLO network introduces the MLCA attention mechanism in the C2f module before the three-stage detection head to construct a C2f_MLCA ​​module, which enhances the feature expression capability of multi-scale traffic targets and achieves high-precision target detection in complex scenarios. The MLCA attention mechanism adopts a dual-branch structure, combines LAP with GAP, and simultaneously models channel, spatial, local and global information, and uses one-dimensional convolution acceleration to achieve an efficient balance between parameters and computational complexity, making it suitable for lightweight target detection networks.

[0026] Preferably, in step S2, the C2f_MLCA ​​module structure is as follows: Figure 3 As shown, specifically including: Pass the input feature map through a 1x1 convolutional layer to expand the number of channels; The expanded feature map is divided into two parts along the channel dimension to obtain two feature maps and ; Will Pass in sequence Bottleneck modules, the output of each module is used as the input of the next module, and finally we get feature maps, respectively The Bottleneck module contains two convolutional layers and an MLCA module, and the module output uses a residual link; All feature maps Splicing is performed along the channel dimension to obtain a spliced ​​feature map; The concatenated feature map is input into the 1x1 convolution layer so that the number of channels is mapped to the target output channel number.

[0027] Preferably, in step S2, the MLCA module structure is as follows Figure 4 As shown, specifically including: Input feature map , after LAP, it is divided into local regions, calculate the average value of each region, and obtain the local features , thereby preserving spatial information and avoiding the loss of details caused by GAP alone; In the global branch, the extracted local feature map After GAP compression, it becomes a global feature map , reshape the global features into To adapt to 1D convolution calculation, thus obtaining global channel attention weight Finally, the global feature map is expanded to size; In the local branch, the extracted local feature map Reshaped into , also calculate the local channel attention weights through 1D convolution, and reshape the results back to size; Will and After normalization, the sum is added and adjusted back to the input feature map size through UNAP to obtain the final attention weight ; The obtained attention weight With the original input feature map The final enhanced feature map output is obtained by element-by-element multiplication, which enables the model to adaptively enhance important features and suppress redundant information.

[0028] Preferably, in step S3, the training environment of the LM-YOLO network is Ubuntu 22.04, NVIDIA GeForce RTX 4060 Ti, CUDA 11.3, Pytorch 1.11, Python 3.8, and the experimental parameter configuration is shown in Table 1.

[0029] Table 1 Experimental parameter settings

[0030] Preferably, in step S3, after the training is completed, in order to verify the detection performance of the LDConv and MLCA modules added to the LM-YOLO algorithm, an ablation experiment is performed on the BDD100k dataset using YOLOv8n as the basic network to evaluate the impact of each module on the performance of the target detection algorithm. The results are shown in Table 2.

[0031] Table 2 Ablation experiment results

[0032] In Table 2, FLOPS represents the model's floating-point operations, Paras represents the number of model parameters, and Weight represents the model weight file size. Smaller values ​​for these three indicate better real-time performance. mAP@0.5 represents the mean average precision of all categories at an intersection-over-union (IoU) threshold of 0.5, and mAP@0.5:0.9 represents the mean average precision of all categories at multiple IoU thresholds (ranging from 0.5 to 0.9, with a step size of 0.05). Larger values ​​for these two indicate better model detection performance. The results show that adding each module individually improves model detection accuracy. Combining the two modules reduces parameters by 4.8% and improves mAP@0.5:0.9 by 3.3% compared to the basic YOLOv8n model. Therefore, the LM-YOLO model maintains its lightweight nature while improving the model's detection accuracy for traffic objects.

[0033] In another aspect, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.

[0034] On the other hand, the present invention further discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0035] In another embodiment provided by the present application, a computer program product comprising instructions is further provided, which, when executed on a computer, enables the computer to execute any of the lightweight traffic target detection methods based on LM-YOLO in the above embodiments.

[0036] It is understandable that the system, device and storage medium provided in the embodiments of the present invention correspond to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts of the above methods.

[0037] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk, tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0038] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0039] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.

[0040] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A lightweight traffic target detection method based on LM-YOLO, characterized in that: The following steps are involved: S1. Obtain a public traffic scene dataset and preprocess it; S2. Build the LM-YOLO network model, using YOLOv8n as the basic architecture. Use the LDConv module to improve the model's feature extraction capabilities for complex targets. Build the C2f_MLCA ​​module before the three-stage detection head to enhance the model's feature discrimination capabilities for multi-scale targets. S3. Configure the training environment and use the data set obtained in step S1 to perform LM-YOLO network training and verification analysis; S4. Input the traffic scene image to be detected into the pre-trained LM-YOLO network model for target detection.

2. A lightweight traffic target detection method based on LM-YOLO according to claim 1, characterized in that: In step S1, the preprocessing of the public dataset includes random data extraction, training set and validation set division, and image and label format conversion to adapt to the YOLOv8 model input.

3. A lightweight traffic target detection method based on LM-YOLO according to claim 1, characterized in that: In step S2, the LM-YOLO network only retains the first layer of feature extraction and uses conventional convolution for preliminary feature mapping. Subsequent downsampling uses the LDConv module to replace the conventional convolution operation. The LDConv module specifically includes: According to the set convolution kernel parameter N, the initial sampling shape and relative coordinates are generated by the initial sampling coordinate generation algorithm; Generate the absolute coordinates of the initial sampling points based on the resolution and stride of the input feature map; The offset generation module performs convolution operation on the input feature map to generate coordinate dynamic offset; Combine the relative coordinates, absolute coordinates and dynamic offset of the initial sampling point to calculate the final sampling point coordinates; Based on the coordinates of the final sampling point, the bilinear interpolation method is used to extract the feature value of the corresponding position from the input feature map to obtain the offset resampled feature map; Reshape the resampled feature map so that it can be processed by the convolution kernel of the main convolution module; The main convolution module is used to perform a convolution operation on the reshaped feature map to obtain the final output feature map.

4. A lightweight traffic target detection method based on LM-YOLO according to claim 3, characterized in that: In step S2, the LM-YOLO network introduces the MLCA attention mechanism into the C2f module before the three-stage detection head to construct a C2f_MLCA ​​module. The C2f_MLCA ​​module structure specifically includes: Pass the input feature map through a 1x1 convolutional layer to expand the number of channels; Divide the expanded feature map into two parts along the channel dimension; Pass one of the feature maps through Bottleneck modules, the output of each module is used as the input of the next module, and finally we get feature maps; All feature maps are spliced ​​along the channel dimension to obtain a spliced ​​feature map; The concatenated feature map is input into the 1x1 convolution layer so that the number of channels is mapped to the target output channel number.

5. A lightweight traffic target detection method based on LM-YOLO according to claim 4, characterized in that: In step S2, the MLCA attention mechanism specifically includes: The input feature map is divided into local regions, calculate the average value of each region, and obtain the local feature map; In the global branch, GAP is performed on the local feature map, compressed into a global feature map, and the global feature map is reshaped and input into a 1D convolution operation to obtain the global channel attention weight. Finally, the global feature map is expanded to size; In the local branch, the local feature map is reshaped, and the local channel attention weight is calculated by 1D convolution, and the result is reshaped back size; The local feature maps and global feature maps obtained by the two branches are normalized and added together, and then adjusted back to the input feature map size through UNAP to obtain the final attention weight; The obtained attention weights are multiplied element-wise with the original input feature map to obtain the final enhanced feature map output.

6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 5.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Road defect detection method based on DRR module and SDFM

    CN119418285A

  • Lightweight target detection method based on local space mixed attention mechanism

    CN119625687A

  • Aluminum profile surface defect detection method and system based on CDA-YOLOv8s

    CN119991614A

  • Road crack detection method, medium and product

    US20250174019A1

Cited By

  • Blood cell detection method based on deep learning

    CN121214087A

  • Image data acquisition method under dangerous condition

    CN121811318A