Small target vehicle detection method and related equipment

By integrating SP-Module and Triplet Attention in the YOLOv8n network, adding upsampling modules, and designing the YOLOv8-SPT model, the problem of insufficient detection accuracy of small-target vehicles under complex operating conditions of autonomous driving is solved, and higher detection accuracy is achieved.

CN120339997APending Publication Date: 2025-07-18SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510424641.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing YOLO series algorithms have problems of mis-checking and missed detection when detecting small-target vehicles under complex operating conditions of autonomous driving, especially in long-distance, low-light or bad weather conditions.

Method used

Based on the YOLOv8n network, the new convolution module SP-Module and attention mechanism Triplet Attention are integrated, the upsampling module is added, and the YOLOv8-SPT model is designed. Through the improvement of the backbone network, neck network and detection head, the feature extraction and detection accuracy is improved.

Benefits of technology

It effectively improves the model's feature extraction and detection accuracy of small-target vehicles, meeting the strict perceptual performance requirements in autonomous driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339997A_ABST
    Figure CN120339997A_ABST
Patent Text Reader

Abstract

The invention discloses a small target vehicle detection method and related equipment. The method comprises the following steps: acquiring image data; inputting the obtained image data into a trained small target vehicle detection model, and outputting a detection result; the small target vehicle detection model comprises a backbone network, a neck network and a detection head; the backbone network fuses a convolution module and an attention mechanism on the basis of a YOLOv8n network, and is used for carrying out feature extraction on an image and providing feature information for a subsequent neck network; the neck network is additionally provided with an up-sampling module on the basis of a YOLOv8n network, feature map size alignment is carried out through an up-sampling and down-sampling bidirectional path, and feature map information is fully fused; the detection head is used for dividing feature information transmitted by the neck network into two paths and processing the two paths respectively, and completing two tasks of bounding box prediction and category prediction at the same time. According to the method, the feature extraction and detection precision of the model on the small target vehicle is effectively improved, and the strict perception performance requirement in an automatic driving scene is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and in particular, to a small target vehicle detection method and related devices. Background Art

[0002] In actual autonomous driving scenarios, the accuracy of small target vehicle detection directly affects the safety and reliability of the system. Although the existing YOLO series algorithms have achieved excellent results in the field of target detection, when facing small target vehicles under complex working conditions of autonomous driving, their detection performance still has deficiencies. Especially under conditions such as long distance, weak light, or bad weather, false detection and missed detection problems are more likely to occur. Therefore, how to further optimize the network structure and enhance the model's perception ability for small target vehicles is a key issue to improve the accuracy of environmental perception in autonomous driving. Summary of the Invention

[0003] To at least partly solve one of the technical problems existing in the prior art, an object of the present invention is to provide a small target vehicle detection method and related devices based on the YOLOv8-SPT model, aiming to effectively improve the feature extraction and detection accuracy of the model for small target vehicles to meet the strict perception performance requirements in autonomous driving scenarios.

[0004] The first technical solution adopted by the present invention is:

[0005] A small target vehicle detection method includes the following steps:

[0006] Obtain image data;

[0007] Input the obtained image data into a trained small target vehicle detection model and output the detection result;

[0008] Wherein, the small target vehicle detection model includes a backbone network, a neck network, and a detection head;

[0009] The backbone network integrates a new convolutional module SP-Module and an attention mechanism Triplet Attention on the basis of the YOLOv8n network, and is used to extract features from the image and provide feature information for the subsequent neck network;

[0010] The neck network adds an upsampling module on the basis of the original YOLOv8n network, aligns the feature map sizes through bidirectional paths of upsampling and downsampling, and fully integrates the feature map information;

[0011] The detection head is used to divide the feature information transmitted by the neck network into two paths and process them separately, and simultaneously complete two tasks of bounding box prediction and class prediction.

[0012] Furthermore, the backbone network further includes a downsampling convolution Conv, a residual module C2f, and a max pooling layer SPPF;

[0013] After the image is input into the backbone network, it sequentially passes through the first downsampling convolution Conv, the second downsampling convolution Conv, the first convolution module SP-Module, the first residual module C2f, the first attention mechanism Triplet Attention, the third sampling convolution Conv, the second convolution module SP-Module, the second residual module C2f, the second attention mechanism Triplet Attention, the fourth downsampling convolution Conv, the third convolution module SP-Module, the third residual module C2f, the third attention mechanism Triplet Attention, the fifth downsampling convolution Conv, the fourth convolution module SP-Module, the fourth residual module C2f, the fourth attention mechanism Triplet Attention, and the max pooling layer SPPF.

[0014] Furthermore, the neck network further includes a concatenation module Concat, a residual module C2f, and a downsampling convolution Conv;

[0015] Among them, the residual module C2f is used to perform differential processing on the input feature map and fuse the original feature map and the feature map processed by the Bottleneck module;

[0016] The concatenation module Concat is used to fuse and splice the feature maps output by different modules, so as to obtain feature maps with different scale information;

[0017] The upsampling module Upsample is used to improve the resolution of the feature map, magnify the deep low-resolution feature map, so that it can be aligned with the shallow high-resolution feature map, and finally realize the cross-scale feature map fusion.

[0018] Furthermore, the detection head includes two channels. The first channel is used to predict the target bounding box, and the second channel is used to predict the target category; each channel contains two convolution modules, and a Conv2d layer follows the convolution module to calculate the bounding box loss or the category loss.

[0019] Furthermore, the convolution module SP-Module includes:

[0020] A spatial-depth conversion module, which is used to downsample the feature map and fuse the sampled feature map, so as to reduce the size of the feature map while ensuring that the detailed information is not lost;

[0021] A convolution with a stride of 1, which is used to retain more feature information while changing the dimension of the feature map.

[0022] Furthermore, the working process of the space-depth conversion module includes:

[0023] The first step: Transform the size of the input feature map according to a preset arrangement method, combine the transformed feature maps into a new feature map, and the size of each pixel point on the new feature map is the average value of the pixels in the area centered on this point in the original feature map;

[0024] The second step: Perform a convolution operation on the new feature map obtained in the first step to further extract the fused information and reduce the dimension of the feature map at the same time.

[0025] Furthermore, the attention mechanism Triplet Attention uses three branch structures to calculate the attention weights for cross-dimensional analysis; in each branch, different permutation transformations are performed on the input tensors respectively, and then the Z-Pool module and convolution module in the structure are used to generate the required attention weight information;

[0026] Assume that the size of the input tensor is H×W×C, where H and W represent the height and width of the tensor respectively, and C represents the number of channels; the tasks of the three branches are as follows: the first branch is responsible for analyzing the relationship between the height H and the channel dimension C; the second branch is responsible for analyzing the relationship between the width W and the channel dimension C; the third branch directly analyzes the input tensor and extracts features through residual transformation; then the results of the three branches are averaged to obtain the output tensor of the entire module.

[0027] Furthermore, the three branches achieve cross-dimensional analysis through rotation:

[0028] In the first branch, the attention weights are solved by using the internal relationship between the height H and width W of the feature map; the input feature map does not undergo rotation transformation in this branch and directly enters the Z-pooling module for dimension reduction operation, and then passes through the convolution, normalization, and Sigmoid layers in sequence, and the feature size is reduced from the original C×H×W to 1×H×W;

[0029] In the second branch, the attention weights are solved by using the internal relationship between the channel dimension C and width W of the feature map; this branch first rotates the input feature map 90 degrees around the width W axis, then passes through the Z-pooling, convolution layer, normalization layer, and Sigmoid layer in sequence to obtain the attention weights, and finally rotates 90 degrees in the reverse direction around the width W axis to restore to the original order;

[0030] In the third branch, the attention weight is solved by using the internal relationship between the channel dimension C and the height H of the feature map. First, the input feature map is rotated 90 degrees around the height H axis to obtain a feature map with dimensions W×H×C, and then it is successively processed by Z pooling, a convolutional layer, a normalization layer, and a Sigmoid layer to obtain the attention weight. Finally, it is rotated back 90 degrees around the height H axis to restore to the original order (C×H×W).

[0031] The attention weights of the three branches are averaged to obtain the final feature map.

[0032] The second technical solution adopted by the present invention is:

[0033] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a small target vehicle detection method as described above.

[0034] The third technical solution adopted by the present invention is:

[0035] A computer-readable storage medium, and at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a small target vehicle detection method as described above.

[0036] The fourth technical solution adopted by the present invention is:

[0037] A computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes a small target vehicle detection method as described above.

[0038] The beneficial effects of the present invention are as follows: Based on the YOLOv8 network, the SP-Module is integrated into the YOLOv8 network, which fully improves the detection ability of the original model for small target vehicles. At the same time, a new attention mechanism TripletAttention is introduced, and a small target vehicle detection model based on YOLOv8-SPT is designed, which effectively improves the feature extraction and detection accuracy of the model for small target vehicles, meeting the strict perception performance requirements in the autonomous driving scenario. Description of the Drawings

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions of the present invention. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0040] Figure 1 is a flowchart of a small target vehicle detection method in an embodiment of the present invention;

[0041] Figure 2 is a structural diagram of the YOLOv8-SPT model in an embodiment of the present invention;

[0042] Figure 3 is a schematic diagram of the downsampling convolution Conv in an embodiment of the present invention;

[0043] Figure 4 is a schematic diagram of the residual module C2f in an embodiment of the present invention;

[0044] Figure 5 is a schematic diagram of the maximum pooling layer SPPF in an embodiment of the present invention;

[0045] Figure 6 is a structural schematic diagram of the detection head in an embodiment of the present invention;

[0046] Figure 7 is a structural diagram of the convolution module SP-Module in an embodiment of the present invention;

[0047] Figure 8 is a schematic diagram of the space-depth conversion module in an embodiment of the present invention;

[0048] Figure 9 is a schematic diagram of the convolution with a stride of 1 in an embodiment of the present invention;

[0049] Figure 10 is a schematic diagram of the convolution operation in an embodiment of the present invention;

[0050] Figure 11 is a structural diagram of the attention mechanism Triplet Attention in an embodiment of the present invention;

[0051] Figure 12 is a structural schematic diagram of the three branches of Triplet Attention in an embodiment of the present invention;

[0052] Figure 13 is a schematic diagram of the rotation operation of Triplet Attention in an embodiment of the present invention;

[0053] Figure 14 It is a display diagram of the SODA10M dataset in an embodiment of the present invention;

[0054] Figure 15 It is a distribution diagram of objects in the SODA10M dataset in an embodiment of the present invention;

[0055] Figure 16 It is a display diagram of the BDDQ00K dataset in an embodiment of the present invention;

[0056] Figure 17 It is a distribution diagram of objects in the BDD100K dataset in an embodiment of the present invention. Detailed implementation manners

[0057] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application. For the step numbers in the following embodiments, they are only set for the convenience of description and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adjusted adaptively according to the understanding of those skilled in the art.

[0058] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of the present application. The singular forms "a", "the" and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. In addition, unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0059] In the description of the present application, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.

[0060] In the description of the present application, the meaning of "several" is one or more, the meaning of "multiple" is two or more, and understandings such as "greater than", "less than", "exceeding", etc. do not include the number itself, and understandings such as "above", "below", "within", etc. include the number itself. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0061] In the description of this application, "and / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0062] Embodiment 1

[0063] See Figure 1 , this embodiment provides a small target vehicle detection method, including the following steps:

[0064] S1. Obtain image data, preprocess the image data, and construct a training set for training the model.

[0065] In some embodiments, before model training, data augmentation processing is performed on the original image data, including increasing the number of small target scene images and enhancing the data quality in harsh weather such as rain and snow, so as to improve the model's ability to solve the long-tail problem and make the model more robust.

[0066] S2. Construct and train the small target vehicle detection model YOLOv8-SPT.

[0067] As an implementation, the small target vehicle detection model includes:

[0068] Backbone network: The backbone network of YOLOv8-SPT integrates the new convolutional module SP-Module and the attention mechanism Triplet Attention on the basis of YOLOv8n, so that the output of the backbone network can provide richer feature information for the subsequent neck network, and finally improve the detection accuracy of the model.

[0069] Neck network: The neck network of YOLOv8-SPT mainly adds an upsampling module on the basis of the original YOLOv8n, and at the same time ensures that the size of the output feature map remains unchanged through convolutional downsampling.

[0070] Detection head: In the detection head part, by using the upsampling and downsampling modules in the neck network, a detection head for small target objects is added, and the category and position information related to small target vehicles are obtained using the output of the previous layer to achieve the purpose of detection.

[0071] S3. Obtain the image data to be detected, and input it into the trained small target vehicle detection model to output the detection result.

[0072] The above-mentioned small target vehicle detection model will be explained in detail below in conjunction with the accompanying drawings and specific embodiments.

[0073] (1) YOLOv8-SPT Model Structure

[0074] The overall network structure of the small target vehicle detection model YOLOv8-SPT is constructed based on YOLOv8n, integrating the new convolutional module SP-Module and the attention mechanism Triplet Attention, and designing an additional detection head S-Detector for small target scenarios. The main components of the model include the Backbone backbone network, the Neck neck network, and the Head detection head. Taking the convolutional neural network as the core, through operations such as convolution, sampling, and pooling, the input image data is trained and inferred. The output of each layer is used as the input of the next layer, and finally, the category and location information of the object to be detected are obtained through the output of the detection head. The network structure of the small target vehicle detection model YOLOv8-SPT is as Figure 2 shown.

[0075] (2) The Backbone Network of the YOLOv8-SPT Model

[0076] The main components of the backbone network of the YOLOv8-SPT model include the downsampling convolution Conv, the convolutional module SP-Module, the residual module C2f, the cross-dimensional attention mechanism Triplet Attention, and the maximum pooling layer SPPF. Its role is to be responsible for feature extraction. Through a series of convolutional and deconvolutional layers, as well as residual connections, the number of model parameters is reduced while ensuring the model's feature extraction ability. Below, the downsampling convolution Conv, the residual module C2f, and the maximum pooling layer SPPF will be briefly introduced.

[0077] 2.1) Downsampling Convolution Conv

[0078] The downsampling convolution Conv is the most basic module in the whole model, including the Conv2d downsampling layer, the BatchNorm2d batch normalization layer, and the SiLU activation function. Its structure is as Figure 3 shown.

[0079] Figure 3 In the figure, K represents the number of convolutional kernels, which affects the depth of the output feature map. Each convolutional kernel is responsible for a certain feature in the input. The combined use of different convolutional kernels can comprehensively and accurately reflect the features of the input object; S in the figure represents the stride of the convolutional kernel moving on the input feature map; P represents the padding number. Padding helps to maintain the spatial features of the input and reduce information loss; C represents the number of channels of the input object. For common RGB images, the number of channels is 3, representing the R, G, and B channels respectively.

[0080] After the convolution operation of the downsampling convolutional layer, the feature map enters the batch normalization layer. The main function of this layer is to improve the stability and convergence speed of model training. Here, the batch normalization layer normalizes the output of the previous convolutional layer, making the mean of each feature tend to 0 and the variance close to 1 in the mini-batch samples, thus ensuring that the data size in the network is appropriate.

[0081] The normalized data is finally fed into the SiLU activation function. As a non-linear activation function, SiLU combines the characteristics of the Sigmoid function, allows smooth gradients, and can avoid the problem of gradient vanishing, performing well in deep network models. The definition of this function is shown in Equation (1).

[0082]

[0083] 2.2) Residual module C2f

[0084] The residual module C2f ensures that the model has sufficient gradient information and at the same time makes the model lightweight. It first extracts the features of the input image through a convolution operation to generate an intermediate feature map; then, through multiple Bottleneck modules, a series of convolution, normalization, and activation operations are performed on the feature map to further extract useful detailed information; then, through a splicing module, the original feature map is fused with the processed feature map, so that information at different levels and scales can be fully utilized; finally, the feature map output after the splicing module in the previous step is convolved to output the final feature map for the next module to use. The network structure of C2f is as Figure 4 shown.

[0085] 2.3) Max pooling layer SPPF

[0086] The core role of the SPPF module in the backbone network is to solve the problem of the network's processing of multi-scale information in images, which is particularly useful in the object detection process. The main components of the SPPF module include the convolution module Conv, the MaxPool2d max pooling module, and the Concat splicing module. Its network structure is as Figure 5 shown. SPPF first extracts preliminary features through a convolution module. Due to downsampling, this step can also reduce the computational amount of the feature map; secondly, through three max pooling modules, spatial downsampling is performed on the input feature map in both the height and width directions, retaining the maximum value of the input tensor and reducing its dimension; then, the splicing module splices the feature maps generated by the three max pooling modules in the channel dimension, so that multi-scale feature information is reflected on a single feature map; finally, the spliced feature map is convolved again to fully fuse different scale information.

[0087] (3) Neck Network of YOLOv8-SPT Model

[0088] The neck network mainly consists of an upsampling module Upsample, a concatenation module Concat, a residual module C2f, and a downsampling convolution Conv. Among them, the implementation principles of the concatenation module, the residual module, and the downsampling convolution module are similar to those in the backbone network. The downsampling convolution can effectively reduce the dimension of the feature map, reduce the computational amount, and can also fuse different information. The residual module processes the input feature map differently, fusing the original feature map and the feature map processed by the Bottleneck module. The concatenation operation fuses and concatenates the feature maps output by different modules to obtain feature maps with different scale information.

[0089] The neck network of the YOLOv8-SPT model has an additional upsampling module Upsample compared to the backbone network. This module is a core component in the feature pyramid. It can improve the resolution of the feature map, magnify the deep low-resolution feature map so that it can be aligned with the shallow high-resolution feature map, and finally achieve cross-scale feature map fusion.

[0090] To improve the recognition accuracy of the model for small target vehicles, the YOLOv8-SPT model adds an ultra-small target detection layer and aligns the feature map sizes through bidirectional paths of upsampling and downsampling, fully fusing the feature map information, enabling the model to have the ability to recognize tiny targets.

[0091] (4) Detection Head of YOLOv8-SPT Model

[0092] See Figure 6 , the detection head processes the feature map transmitted from the neck network in two paths respectively, and simultaneously completes two tasks of bounding box prediction and class prediction.

[0093] As an anchor-free model, YOLOv8-SPT does not need to predict from the offsets of known anchor boxes, but directly detects the center of the object to be detected. This anchor-free detection method can greatly reduce the model's dependence on the number of prediction boxes and significantly accelerate the subsequent processing steps. The network structure of the detection head of the YOLOv8-SPT model is as Figure 6 shown. The detection head contains two channels. The first is used to predict the target bounding box, and the second is used to predict the target class. Each channel contains two convolutional modules, and a Conv2d layer follows the convolutional module to calculate the bounding box loss and class loss.

[0094] (5) SP-Module Convolutional Module

[0095] Convolutional neural networks play an excellent role in object detection tasks. However, when the image resolution decreases or the target pixels are small, the performance of the neural network cannot be guaranteed. In existing convolutional neural networks, stride convolutions are used to learn the feature information of the target object. When the number of convolutional layers increases, it is easy to cause the problem of loss of detailed features of the target object. Therefore, in this embodiment, a convolutional neural network module SP-Module is newly built, which can well replace the convolutional operation in the traditional neural network and largely avoid the problem of loss of target details during convolution. The structure of the SP-Module convolutional module is as Figure 7 shown.

[0096] SP-Module applies graphic transformation technology to convolutional neural networks and draws on graphic scaling technology in feature map mapping. Its network structure mainly includes two parts. First, there is a spatial-depth conversion module, which downsamples the feature map and fuses the sampled feature maps, achieving the purpose of reducing the size of the feature map while ensuring that the detailed information is not lost. Secondly, a convolution with a stride of 1 is used to change the dimension of the feature map while retaining more feature information.

[0097] 5.1) Spatial-depth conversion module

[0098] The spatial-depth conversion module mimics graphic transformation technology, converting the spatial dimension of the feature map into depth, increasing the number of channels while reducing the size of the feature map. This change makes the information extracted by the model more distinguishable, thus showing better generalization performance when dealing with image size and rotation transformations. After the input feature map undergoes this size transformation, it will be recombined, ultimately achieving the purpose of converting spatial dimension information into depth information. The mechanism of the spatial-depth conversion module is shown by Figure 8 shown.

[0099] The process of the spatial-depth conversion module is mainly divided into two steps. In the first step, the input feature map is resized according to the agreed arrangement, and the transformed feature maps are combined into a new feature map. The size of each pixel point on the new feature map is the average value of the pixels in the area centered on this point in the original feature map. In the second step, a convolution operation is performed on the new feature map obtained in the first step to further extract the fused information and reduce the dimension of the feature map at the same time. The downsampling operation performed in the spatial-depth conversion module does not cause loss of feature information because the feature maps obtained by downsampling are fused and stitched together in accordance with the agreed method, so that all the detailed information within the channel dimension is retained. Assuming that the size of the intermediate feature map is L×L×C1 and the downsampling scale factor is set to size, after the downsampling process of the spatial-depth conversion module, all feature map subsequences are obtained. Combining with the network structure diagram of the SP-Module module, the mathematical expression of the feature map subsequence output by the spatial-depth conversion module is shown in the following formula (2).

[0100]

[0101] In the formula, f size-1,size-1 represents one of the feature map subsequences generated by downsampling. Assuming that the downsampling scale factor size is selected as 2, according to the formula name shown above, the intermediate feature Figure X will be downsampled in the form of size equal to 2, and finally f 0,0 , f 0,1 , f 1,0 , f 1,1 four feature map subsequences will be obtained. And the size of the subsequence becomes half of the intermediate feature Figure X , that is, L / 2×L / 2×C1. Then, the subsequences are fused and stitched in the dimension of the number of channels to generate a new feature map, and the size of the new feature map is L / 2×L / 2×4C1.

[0102] So far, it can be found that the relationship between the size of the new feature map and the intermediate feature map can be expressed as follows. Assuming that the size of the intermediate feature map is L×L×C1 and the scale factor is size, the size of the generated new feature map is L / size×L / size×size 2 C1.

[0103] 5.2) Convolution with a stride of 1

[0104] In the convolution operation of a convolutional neural network, information loss usually occurs, especially when entering deep convolutions. Using a convolution with a stride of 1 can well avoid this problem because information loss only occurs in convolution operations with a stride greater than 1. After performing the convolution with a stride of 1, the number of channels is set to C2, and the relationship C2 < size 2×C1, after the convolution operation here, the size of the feature map output by the final spatial-depth conversion module becomes L / size × L / size × C2.

[0105] See Figure 9 , Figure 9 is a schematic diagram of a three-channel convolution with a stride of 1. After multiplying each convolution kernel by the pixel values on its corresponding feature map and summing them, and then summing the pixel values at the corresponding positions of each channel, the output feature map of the multi-channel convolution is finally obtained. It should be noted that the stride of the convolution kernel after each multiplication and summation is 1, which can retain the detailed information to the greatest extent.

[0106] See Figure 10 , after performing a convolution operation with a stride of 1 on a three-channel feature map with a size of 5×5×3 using a 3×3 convolution kernel, an output feature map with a size of 3×3×1 will be obtained after a 3×3×1 operation.

[0107] (6) Attention mechanism Triplet Attention

[0108] Channel attention mechanism and spatial attention mechanism play an increasingly important role in computer vision tasks. In this embodiment, from the cross-dimensional perspective, the lightweight attention mechanism TripletAttention is integrated into the YOLOv8 network. The basic principle of this module is to use three branch structures to calculate the attention weights for cross-dimensional analysis. In each branch, different permutation transformations are performed on the input tensor respectively, and then the Z-Pool module and convolution module in the structure are used to generate the required attention weight information. Because the computational cost is very small, this module does not increase the model parameters much and performs excellently in saving computational resources. The structure diagram of the Triplet Attention module is as Figure 11 shown.

[0109] Assume that the size of the input tensor is H×W×C, where H and W represent the height and width of the tensor respectively, and C represents the number of channels. As Figure 11 shown, the Triplet Attention module consists of three branches, and the tasks of the three branches are:

[0110] The first branch is responsible for analyzing the relationship between the height H and the channel dimension C; the second branch is responsible for analyzing the relationship between the width W and the channel dimension C; the third branch directly analyzes the input tensor and extracts features through residual transformation. Then, the results of the three branches are averaged to obtain the output tensor of the entire module. Figure 12 is the detailed network structure diagram of the Triplet Attention module.

[0111] In Figure 12Among them, the size of the input feature map is C×H×W, where C is the number of channels, H is the height, and W is the width. The three branches perform different operations, including Channel Pooling, Permute, 7×7 convolution, Z Pooling, Batch Normalization, and Sigmoid activation function. The Permute operation is mainly used to adjust the dimensions. Figure 12 The left branch is used to obtain the interaction relationship between spatial dimensions. This branch does not change the arrangement of the input feature map. After performing Z Pooling and convolution operations, it generates attention weights through the Sigmoid activation function. The middle branch is used to obtain the interaction relationship between the channel dimension C and the spatial dimension H. This branch passes through Z Pooling and convolution and finally obtains attention weights through the Sigmoid activation function. The right branch is used to obtain the interaction relationship between the channel dimension C and the spatial dimension W. Similarly, it obtains attention weights through Z Pooling, convolution, and the Sigmoid activation function. After each branch obtains the attention weights, they are sorted through the Permute module weights, and then the attention weights obtained from the three branches are averaged. This structure can comprehensively integrate the feature information in different dimensions through rotation and sorting operations, and better capture the internal features of the data. Compared with the previous attention mechanisms, the Triplet Attention module has the following three main improvement points.

[0112] 1) Cross-dimensional attention weight calculation: Calculate the attention weights through the internal interaction relationship between the three dimensions of the feature map channels, height, and width, enabling the model to more comprehensively capture the key features in space and semantics.

[0113] 2) Emphasize cross-dimensional interaction: The module emphasizes the deep feature correlation between different dimensions when calculating weights, and further enhances the model's sensitivity to local details through information fusion of cross paths.

[0114] 3) Rotation operation and residual transformation: Rotate the input feature map tensor and perform residual transformation to analyze the relationship between different dimensions, which not only improves the feature expression ability but also enhances the robustness and generalization ability of the model.

[0115] Refer to Figure 13 and emphasizes how the three branches of the Triplet Attention module perform cross-dimensional analysis through rotation. Figure 13 Each row in Figure 12 represents a branch in

[0116] Figure 13 The (a) branch in Figure 12The left branch in [description], in which the attention weights are solved by utilizing the internal relationship between the height H and width W of the feature map. The input feature map does not undergo rotation transformation in this branch and directly enters the Z-pooling module for dimensionality reduction operation. The principle of Z-pooling is to reduce the number of channels and concatenate the average pooling and maximum pooling features, and its mathematical calculation formula is shown as the following formula (3). Then, after passing through the convolution, normalization, and Sigmoid layers in sequence, the feature size is reduced from the original C×H×W to 1×H×W.

[0117] Z-pool(x)=[MaxPool 0d (x),AvgPool 0d (x)] (3)

[0118] Figure 13 The (b) branch in [description] corresponds to Figure 12 the middle branch in [description]. In this branch, the attention weights are solved by utilizing the internal relationship between the channel dimension C and width W of the feature map. Different from the (a) branch, this branch first rotates the input feature map 90 degrees around the width W axis, and then passes through modules such as Z-pooling, convolution layer, normalization layer, and Sigmoid layer in sequence to obtain the attention weights, and finally rotates it 90 degrees in the reverse direction around the width W axis to restore to the original order.

[0119] Figure 13 The (c) branch in [description] corresponds to Figure 12 the right branch in [description]. In this branch, the attention weights are solved by utilizing the internal relationship between the channel dimension C and height H of the feature map. First, the input feature map is rotated 90 degrees around the height H axis to obtain a feature map with the size of W×H×C. Then, it passes through modules such as Z-pooling, convolution layer, normalization layer, and Sigmoid layer in sequence to obtain the attention weights, and finally rotates it 90 degrees in the reverse direction around the height H axis to restore to the original order (C×H×W). The attention weights of the three branches are averaged to obtain the final feature map.

[0120] (7) Data preprocessing

[0121] In this embodiment, in order to verify the effect of model improvement, the SODA10M dataset is mainly used for experiments. At the same time, in order to detect the robustness of the model, another mainstream dataset BDD100K is added. The following is the introduction of the two datasets.

[0122] 7.1) SODA10M dataset

[0123] The SODA10M (Scenes, Objects, Depth and Annotations) dataset mainly contains 10 million unlabeled road scene images and 20,000 labeled images collected in 32 cities in China, and is widely used to evaluate the performance of object detection algorithms related to the field of autonomous driving. It includes six common detection categories on the road, namely Pedestrian, Cyclist, Car, Truck, Tram, and Tricycle. These six categories are all common road recognition objects in autonomous driving scenarios and cover different weather conditions, driving scenarios, and times. Figure 14 Several photos in the dataset are shown.

[0124] In this embodiment, 20,000 labeled images in this dataset are mainly used for model training and algorithm performance evaluation. The 20,000 annotated data are divided into 8,000 training data, 2,000 validation data, and 10,000 test data according to the ratio of 4:1:5. Figure 15 The distribution of various objects in the training dataset is shown.

[0125] It can be seen from the figure that the vehicle objects in the dataset are much more than pedestrians, and the proportion of cars is much more than that of other vehicles, which conforms to the recognition scenario in the field of autonomous driving. By checking the dataset, it is found that there are many occlusions and long-distance scenarios for cars in the data, so it meets the detection requirements for small target vehicles in this paper.

[0126] 7.2) BDD100K dataset

[0127] The BDD100K dataset is one of the largest and most diverse datasets for autonomous driving. This dataset contains 100,000 images, covering image data collected under different weather and road scenarios. And its categories are more than those of the SODA10M dataset, posing higher requirements for the accuracy and robustness of the model. Figure 16 Several photos in the dataset are shown.

[0128] In this embodiment, 5,000 labeled data are selected for testing, which includes ten categories, including 10 representative categories, namely bicycle, bus, car, motorcycle, person, rider, lamp, traffic sign, train, truck. Figure 17 The distribution of various objects in the selected test data is shown.

[0129] From Figure 17 it can be seen that among all categories, the number of cars is far more than that of other categories, and considering the actual situation of the images, it meets the research requirements for small target vehicles in this paper.

[0130] Before the formal experiment, it is necessary to process the format of the dataset. Through preliminary experiments, it is found that there are 21 training and validation data with annotation out-of-bounds problems in the SODA10M dataset, of which 18 belong to the training set and 3 belong to the validation set. After removing the unqualified data, there are still 7,982 training data and 1,997 validation data, and the quantity meets the experimental requirements. For the BDD100K dataset, first convert its json-format labels into txt-format labels that meet the requirements of YOLO. At the same time, remove the labels with too small pixel values, because when the pixel values are too small, it will affect the model's learning of effective features during training. When the learned features are biased or even wrong, it will lead to a decrease in the final accuracy. After preprocessing, both datasets meet the standards for model training.

[0131] (8) Experiment and Result Analysis

[0132] 8.1) Experimental Platform

[0133] This embodiment introduces the experimental platform of this chapter from the software and hardware configurations. First, in terms of hardware configuration, an Nvidia RTX 4060 (8.0GB) graphics card, 32GB of running RAM, and a processor of 13th Gen Intel(R) Core(TM) i5-13500HX (2.50GHz) are used. At the software level, the Windows 11 platform is adopted, the integrated development environment is PyCharm 2024.1, the programming language is selected as Python 3.6, the CUDA version number is 12.6, and the deep learning framework is selected as the Pytorch 2.0.0 version. The configuration of the experimental platform in this embodiment is shown in Table 1 below.

[0134] Table 1 Hardware Configuration Table for Model Training

[0135]

[0136] Based on the above-built experimental platform, some hyperparameters related to deep learning selected in this embodiment are: the data batch processing size (batch size) is 8, the number of training epochs (epochs) is 300, the image input resolution (imgsz) is 640×640, the number of data loading threads (workers) is 8, the optimizer selects Stochastic Gradient Descent (SGD), the initial learning rate (lr0) is 0.01, the momentum is 0.937, and the Intersection over Union (IoU) threshold is set to 0.7.

[0137] 8.2) Algorithm Flow

[0138] The overall training process of the algorithm in this embodiment is shown in Table 2 below.

[0139] Table 2 Flow Chart of YOLOv8-SPT Model Training Algorithm

[0140]

[0141]

[0142] Example 2

[0143] The embodiment of the present invention further provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and is loaded and executed by the processor to implement the method for detecting small target vehicles as Figure 1 shown.

[0144] It can be understood that the memory may include a Random Access Memory (RAM), and may also include a Read-Only Memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function, instructions for implementing the above-mentioned method embodiments, etc.; the data storage area can store data created according to the use of the server, etc.

[0145] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts within the entire server, and by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory, it executes various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor may integrate one or several combinations of a Central Processing Unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor and may be implemented separately by a single chip.

[0146] Since this electronic device is the electronic device corresponding to a small target vehicle detection method according to an embodiment of the present invention, and the principle of the electronic device to solve problems is similar to that of the method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0147] Embodiment 3

[0148] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one segment of program, a code set or an instruction set is stored, and the at least one instruction, the at least one segment of program, the code set or the instruction set is loaded and executed by a processor to implement Figure 1 a small target vehicle detection method as shown.

[0149] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, a magnetic disk memory, a tape memory, or any other computer-readable medium capable of carrying or storing data.

[0150] Since this storage medium is the storage medium corresponding to a small target vehicle detection method according to an embodiment of the present invention, and the principle of the storage medium to solve problems is similar to that of the method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0151] Embodiment 4

[0152] In some possible implementations, aspects of the method of the embodiments of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a small target vehicle detection method according to various exemplary embodiments described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in high-level programming languages such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0153] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0154] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0155] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and their purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.

Claims

1. A small target vehicle detection method, characterized in that, It includes the following steps: Obtain image data; Input the obtained image data into a trained small target vehicle detection model to output the detection result; Among them, the small target vehicle detection model includes a backbone network, a neck network, and a detection head; The backbone network integrates the convolutional module SP-Module and the attention mechanism Triplet Attention on the basis of the YOLOv8n network, and is used to extract features from the image and provide feature information for the subsequent neck network; The neck network adds an upsampling module on the basis of the YOLOv8n network, and aligns the feature map sizes through the bidirectional paths of upsampling and downsampling to fully integrate the feature map information; The detection head is used to divide the feature information transmitted from the neck network into two paths and process them separately, and simultaneously complete the two tasks of bounding box prediction and class prediction.

2. The small target vehicle detection method according to claim 1, characterized in that The backbone network also includes a downsampling convolution Conv, a residual module C2f, and a spatial pyramid pooling layer SPPF; After the image is input into the backbone network, it sequentially passes through the first downsampling convolution Conv, the second downsampling convolution Conv, the first convolutional module SP-Module, the first residual module C2f, the first attention mechanism Triplet Attention, the third sampling convolution Conv, the second convolutional module SP-Module, the second residual module C2f, the second attention mechanism Triplet Attention, the fourth downsampling convolution Conv, the third convolutional module SP-Module, the third residual module C2f, the third attention mechanism Triplet Attention, the fifth downsampling convolution Conv, the fourth convolutional module SP-Module, the fourth residual module C2f, the fourth attention mechanism Triplet Attention, and the spatial pyramid pooling layer SPPF.

3. A small target vehicle detection method according to claim 1, characterized in that, The neck network also includes a concatenation module Concat, a residual module C2f, and a downsampling convolution Conv; Among them, the residual module C2f is used to differentially process the input feature map and fuse the original feature map and the feature map processed by the Bottleneck module; The concatenation module Concat is used to fuse and splice the feature maps output by different modules to obtain feature maps with different scale information; The upsampling module Upsample is used to increase the resolution of the feature map, magnify the deep low-resolution feature map, so that it can be aligned with the shallow high-resolution feature map, and finally realize the cross-scale feature map fusion.

4. A small target vehicle detection method according to claim 1, characterized in that, The detection head includes two channels. The first channel is used to predict the target bounding box, and the second channel is used to predict the target class; each channel contains two convolutional modules, and a Conv2d layer follows the convolutional module to calculate the bounding box loss or class loss.

5. A small target vehicle detection method according to claim 1, characterized in that, The convolutional module SP-Module includes: A spatial-depth conversion module, which is used to downsample the feature map and fuse the sampled feature map, and achieve the purpose of reducing the feature map size while ensuring that the detailed information is not lost; A convolution with a step size of 1 is used to retain more feature information while changing the dimension of the feature map.

6. The small target vehicle detection method according to claim 5, characterized in that, The working process of the spatial-depth conversion module includes: The first step: The input feature map is resized according to a preset arrangement method, and the transformed feature maps are combined into a new feature map. The size of each pixel point on the new feature map is the average value of the pixels in the area centered on this point in the original feature map. The second step: A convolution operation is performed on the new feature map obtained in the first step to further extract the fused information and reduce the dimension of the feature map at the same time.

7. A small target vehicle detection method according to claim 1, characterized in that, The attention mechanism TripletAttention uses three branch structures to calculate the attention weights for cross-dimensional analysis; in each branch, different permutation transformations are performed on the input tensor respectively, and then the Z-Pool module and convolution module in the structure are used to generate the required attention weight information. Assume that the size of the input tensor is H×W×C, where H and W represent the height and width of the tensor respectively, and C represents the number of channels; the tasks of the three branches are as follows: the first branch is responsible for analyzing the relationship between the height H and the channel dimension C; the second branch is responsible for analyzing the relationship between the width W and the channel dimension C; the third branch directly analyzes the input tensor and extracts features through residual transformation; then the results of the three branches are averaged to obtain the output tensor of the entire module.

8. A small target vehicle detection method according to claim 7, characterized in that, The three branches achieve cross-dimensional analysis through rotation: In the first branch, the attention weights are solved by using the internal relationship between the height H and width W of the feature map; the input feature map does not undergo rotation transformation in this branch and directly enters the Z-pooling module for dimension reduction operation, and then passes through the convolution, normalization, and Sigmoid layers in sequence. After that, the feature size is reduced from the original C×H×W to 1×H×W. In the second branch, the attention weights are solved by using the internal relationship between the channel dimension C and width W of the feature map. In this branch, the input feature map is first rotated 90 degrees around the width W axis, and then passes through the Z-pooling, convolution layer, normalization layer, and Sigmoid layer in sequence to obtain the attention weights, and finally rotates 90 degrees in the reverse direction around the width W axis to restore to the original order; in the third branch, the attention weights are solved by using the internal relationship between the channel dimension C and height H of the feature map; first, the input feature map is rotated 90 degrees around the height H axis to obtain a feature map with a size of W×H×C, and then passes through the Z-pooling, convolution layer, normalization layer, and Sigmoid layer in sequence to obtain the attention weights, and finally rotates 90 degrees in the reverse direction around the height H axis to restore to the original order. The attention weights of the three branches are averaged to obtain the final feature map.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Power equipment intelligent identification method and system based on deep learning

    CN120894640A

  • Facial expression recognition method for autistic children based on deep learning

    CN121121829A