A target detection method for open road scenes based on thermal infrared recognition

By improving the YOLOv8s network structure and enhancing feature representation and robustness, the complex environment problem of target detection in open road scenes with thermal infrared recognition is solved, and efficient and accurate small target detection is achieved, which is suitable for real-time target detection in resource-constrained devices.

CN119919771BActive Publication Date: 2025-09-30SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510001111.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-09-30
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

Existing open-scene target detection technology based on thermal infrared recognition on roads has difficulty achieving efficient and accurate target detection, especially the recognition of small targets, in complex environments, especially under conditions of changing lighting and occlusion. In addition, the lightweight attention mechanism has high computational complexity on resource-constrained devices, making it difficult to meet real-time processing requirements.

Method used

The YOLOv8s network structure is improved, and the feature representation and robustness are enhanced through the depth-separable convolution layer, the improved SimAM attention mechanism layer, the Upsample-Extract feature extraction module, the Extract-Fusion feature fusion module, the micro-scale detection head and the improved SE attention mechanism layer, thereby improving the small target detection capability.

Benefits of technology

It improves the accuracy and robustness of target detection in complex environments, reduces computational complexity and energy consumption, and is suitable for real-time target detection in resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919771B_ABST
    Figure CN119919771B_ABST
Patent Text Reader

Abstract

This paper discloses a method for detecting targets in open road scenes based on thermal infrared recognition. The method collects image samples of scenes from the forward view of a thermal infrared dashcam, annotates the images using semi-automatic annotation tools and RectLabel annotation tools, and uses data augmentation techniques to obtain diverse training samples. The method also improves the C2f module in the neck of YOLOv8s and introduces an improved SimAM attention mechanism layer. The neck network structure and backbone feature extraction network of YOLOv8s are designed, and the SPPF module in the backbone network is improved. The head network of YOLOv8s is improved, with a decoupled small target detection head added and an improved SE attention mechanism layer introduced in the head. This method significantly improves target detection performance in complex road environments, particularly under challenging conditions such as illumination changes, vehicle lighting effects, and the diversity of different types of vehicles, achieving high robustness, accuracy, and efficiency in target detection tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection in computer vision, and in particular to a method for detecting targets in open road scenes based on thermal infrared recognition. Background Art

[0002] Road object detection technology is a core component of intelligent transportation systems. With the rapid development of urbanization and the surge in the number of vehicles, traffic conditions are becoming increasingly complex. Against this backdrop, accurate and real-time detection and identification of various road objects, including vehicles, pedestrians, and traffic signs, is crucial for improving road safety and traffic flow efficiency. Therefore, to enhance road traffic safety and ensure the safety of the public, road object detection technology based on thermal infrared recognition is developing towards full automation and intelligence, making intelligent road object detection a new direction for the development of intelligent transportation systems.

[0003] Traditional manual detection methods are no longer able to meet the growing demands of traffic management, and the efficiency of road object detection directly impacts smooth traffic flow and road safety. In recent years, the application of thermal infrared technology in object detection has significantly improved the efficiency and accuracy of detection tasks. Leveraging these technologies, object detection methods can accurately detect and extract target information from thermal infrared images in complex urban traffic scenes. However, in practical applications, traditional thermal infrared-based object detection methods for open-road scenes still face many challenges due to the variability of urban environments, such as lighting changes, weather conditions, and object occlusions. These issues limit the performance of detection technology in highly dynamic and complex urban traffic environments, particularly in identifying small and dim targets and handling occlusions. Therefore, to further enhance road driving safety, road object detection technology requires further research and improvement to achieve more efficient and intelligent traffic operations and management, and to promote the modernization of road traffic management.

[0004] In the prior art, some studies have improved the detection performance by improving the network structure to solve the problem of target detection in complex environments. For example, ZHANG et al. (ZHANG Xiu-Wei, ZHANG Yan-Ning, YANG Tao, ZHANG Xin-Gong, SHAO Da-Pei. Automatic Visual-thermal Image Sequence Registration Based on Co-motion. ACTA AUTOMATICA SINICA, 2010, 36(9): 1220-1231. doi: 10.3724 / SP.J.1004.2010.01220) proposed a method for automatic fusion of visible light and thermal infrared image sequences based on Co-motion. The method introduced the statistical features of Co-motion motion to solve the problem of fusion of heterogeneous image sequences, thereby avoiding the difficulty of similar image feature extraction and accurate motion detection of heterogeneous images. However, this method still has limitations under complex lighting changes and occlusion conditions. For example, in an outdoor multi-camera system, due to changes in lighting conditions and different camera positions, the observed objects may be very different, making it difficult to extract similar foregrounds. Moreover, in unstructured environments, due to the unpredictability of motion, this method may perform poorly in accurately tracking continuous motion, especially in the presence of a large number of external point interferences and susceptibility to large-scale changes. In addition, for the problem of small target detection, some studies have attempted to improve detection accuracy through multi-scale feature fusion. For example, Chen et al. (ChenY, PangY, Wang S, et al. Small Object Detectionin Aerial Images via Attention Module and Feature Fusion[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 59(7): 5991-6004.) used attention mechanisms and feature fusion techniques to improve the detection performance of small targets, but in practical applications, these methods often have high computational complexity and are difficult to meet the needs of real-time processing.

[0005] In the area of ​​attention mechanisms, some studies have proposed lightweight attention mechanisms to capture long-term dependencies and reduce computational burden. For example, Qiu et al. (Qiu R, Li J, Huang Z, et al. Rethinking Attention with Performers [J]. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2021, 10373-10382) proposed an efficient attention model, Performers, that reduces computational complexity by approximating the softmax function. However, these lightweight attention mechanisms still face performance bottlenecks when processing dense object scenes. On resource-constrained devices, such as embedded systems or mobile devices, while lightweight attention mechanisms reduce computational workload, they still consume high memory and energy in dense object detection tasks. In an object detection experiment on a mobile device, a system using the Performers model experienced increased battery consumption and a significant increase in computing resource usage when processing dense object scenes. Furthermore, lightweight attention mechanisms may perform well on specific datasets, but their generalization capabilities across different scenarios are limited.

[0006] In summary, while thermal infrared recognition-based object detection technology for open-road scenes has made some progress, further research and improvement are needed to better meet the high efficiency and high accuracy requirements of assisted driving. This will bring more efficient and intelligent operation and management to intelligent transportation, and promote the modernization of assisted driving. Summary of the Invention

[0007] The present invention aims to address the limitations of the above-mentioned prior art and proposes a target detection method for open road scenes based on thermal infrared recognition to improve the accuracy and robustness of target detection under complex road conditions.

[0008] To achieve the above objectives, the present invention provides a method for detecting targets in open road scenes based on thermal infrared recognition, comprising the following steps:

[0009] (1) Collect image samples of the scene in the forward view of the thermal infrared driving recorder;

[0010] (2) Preprocess and label the sample images;

[0011] (3) Training set production based on data augmentation;

[0012] (4) Improve the C2f module in the neck of YOLOv8s, also known as CSP Bottleneck with 2convolutions: By replacing the C2f module in the neck of the original YOLOv8s network with a depthwise separable convolution layer, a residual block formed by stacking three Bottleneck layers, a depthwise separable convolution layer, and an improved SimAM, also known as the Simple Parameter-FreeAttentionModule attention mechanism layer, and adding a DSCBS, also known as the Depthwise SeparableConvolution-Batch Normalization-SiLU module, to connect the split layer and the concat layer of C2f, the representation ability of feature information and the ability to recognize high-level semantic features of small targets are enhanced;

[0013] (5) Design the neck network structure and backbone network feature extraction network of YOLOv8s: Based on the original neck network of YOLOv8s, an Upsample-Extract feature extraction module is added to the top of the top-down aggregation path and an Extract-Fusion feature fusion module is added to the bottom of the bottom-up aggregation path. In addition, the SPPF and Spatial Pyramid Pooling-Fast modules in the backbone network are replaced with STN, a multi-layer cascade of Spatial Transformer Networks layers and ghost convolution layers to improve the robustness of the model and the efficiency of feature extraction.

[0014] (6) Improve the head network of YOLOv8s: Based on the original YOLOv8s detection head, a micro-scale detection head DDH is added to the C2f module of the Extract-Fusion module in the neck of YOLOv8s, namely Decoupled Detection Head. An improved SE, namely Squeeze-and-Excitation Networks attention mechanism layer, is introduced into the micro-scale detection head DDH and the original multi-scale detection head of YOLOv8s;

[0015] (7) Train the improved YOLOv8s deep network and input the test video to perform network forward reasoning to detect targets in open road scenes;

[0016] The image samples of the forward-view scene collected by the thermal infrared driving recorder refer to the thermal infrared image data of vehicles driving or parked on the road at different angles and distances under different urban and suburban conditions; the preprocessing and annotation of the sample images refer to the use of the OpenCV image processing library to adjust all images to a uniform size of 480×640 pixels, semi-automatically annotating the collected images through a semi-automatic annotation video training data tool based on a multi-target detector and tracker, and verifying feedback through the existing YOLOv8s model, and then manually annotating and correcting the image annotation samples with poor detection results using the RectLabel annotation tool; the training set preparation based on data augmentation refers to grouping the image samples, and then performing random rotation on them, adjusting the image contrast, adding Gaussian noise to the images, and applying elastic transformation; the improved YOLOv8s deep network refers to the deep network obtained by the improvement through steps (4), (5), and (6), and the training of the improved YOLOv8s deep network refers to completing the training of the deep network using the Adam algorithm.

[0017] Furthermore, the improved C2f module in the neck of YOLOv8s refers to replacing the C2f module in the neck of the original YOLOv8s network with a depthwise separable convolutional layer, a residual block formed by stacking three Bottleneck layers, a depthwise separable convolutional layer, and an improved SimAM attention mechanism layer in series, and adding a DSCBS module to connect the split layer and the concat layer of C2f to integrate multi-level feature information, improve sensitivity to key features, enhance the ability to recognize high-level semantic features of small targets, alleviate the problems of deep gradient disappearance and gradient explosion, and replace the convolutional layer with a depthwise separable convolutional layer to reduce the number of model parameters; the DSCBS refers to the existing CBS, also known as Convolution-Batch The convolution layer of the Normalization-SiLU module is replaced with a depth-wise separable convolution layer; the improved SimAM attention mechanism layer refers to adding a weight decay regularization operation to the SimAM module to preprocess the data input to the SimAM module, promote the dispersion and distinguishability of features, and replace the activation function from Sigmoid to Tanh to solve the non-zero centering problem of Sigmoid, improve the stability of feature representation, and alleviate gradient disappearance.

[0018] Furthermore, the design of the neck network structure and backbone network feature extraction network of YOLOv8s refers to adding an Upsample-Extract feature extraction module at the top of the top-down aggregation path and an Extract-Fusion feature fusion module at the bottom of the bottom-up aggregation path on the basis of the original YOLOv8s neck network to enhance the features of small targets, and replacing the SPPF module in the original YOLOv8s backbone network with a cascade of STN layer, ghost convolution layer, STN layer, three parallel MaxPool2d layers, Concat layer, two parallel STN layers, and ghost convolution layer, and adding a residual connection after the first convolution layer; the Upsample-Extract feature extraction module is a cascade of Concat layer, C2f module, Upsample layer, Concat layer, and C2f module; the Extract-Fusion feature fusion module is a cascade of C2f module, convolution layer, and Concat layer, thereby minimizing gradient attenuation and improving network structure efficiency while enhancing the detection accuracy of surface feature maps.

[0019] Furthermore, the improved head network of YOLOv8s refers to adding a multi-scale prediction of the decoupled small target detection head to the C2f module of the Extract-Fusion module at the neck of YOLOv8s, and introducing an improved SE attention mechanism layer to improve the head network of YOLOv8s; the added multi-scale prediction of the decoupled small target detection head refers to adding a small scale detection head DDH to the C2f module of the Extract-Fusion module at the neck of YOLOv8s on the basis of the original YOLOv8s detection head, and adding a small scale detection head DDH to the C2f module of the Extract-Fusion module at the neck of YOLOv8s. An improved SE attention mechanism layer is introduced into the original multi-scale detection head of LOv8s. The improved SE attention mechanism layer refers to adding a parallel branch consisting of a cascade of convolutional layer, normalization layer, LeakyReLU layer, convolutional layer, normalization layer, and Swish layer in the Excitation stage of the SE module, so that it can adaptively assign weights to each channel, thereby improving the detection performance of small targets. By adding an additional cascade structure, more complex features and relationships between channels can be captured. The introduction of LeakyReLU solves the dead ReLU problem and improves the adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flow chart of an embodiment of the present invention;

[0021] Figure 2 This is a fusion diagram of the YOLOv8s improved network and SE module according to an embodiment of the present invention;

[0022] Figure 3 This is a structural diagram of an improved SE module according to an embodiment of the present invention;

[0023] Figure 4 This is a structural diagram of an improved C2f module according to an embodiment of the present invention;

[0024] Figure 5 This is a structural diagram of an improved SPPF module according to an embodiment of the present invention;

[0025] Figure 6 This is a diagram of the depthwise separable convolution structure of an embodiment of the present invention;

[0026] Figure 7 This is a diagram of the ghost convolution structure of an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The technical approach of the present invention is further described in detail below through specific embodiments and drawings.

[0028] like Figure 1 As shown in FIG6 , an embodiment of the method for detecting targets in an open road scene based on thermal infrared recognition of the present invention includes the following steps:

[0029] 1. Collect image samples of the forward-view scene from the thermal infrared driving recorder

[0030] Use a high-resolution thermal infrared dashcam to capture image data while the vehicle is driving. These image samples should cover a variety of driving scenarios, such as urban roads, highways, and rural roads, to ensure that different target objects and environmental changes are captured.

[0031] Image data of vehicles at different angles and distances, in different states of motion and stillness, and at different speeds, is collected under different weather conditions, time periods, and road types. This ensures that the image samples account for changes in lighting, the impact of vehicle lights, and the diversity of different vehicle types, ensuring comprehensiveness and representativeness. Thermal infrared image samples can also provide image data at night or in low-light conditions. Thermal infrared dashcams can capture infrared wavelengths from 8 to 14 μm, providing clear images in a variety of lighting conditions, which is crucial for analyzing and identifying objects in varying weather and lighting conditions.

[0032] 2. Preprocess and label sample images

[0033] Establish a unified file naming convention and format, and adjust the size of image samples to meet the requirements of model training. Then, use manual or automated tools to annotate the objects in the image, marking the identified objects (such as pedestrians, vehicles, traffic signs, etc.) with boxes.

[0034] Specifically, all images were resized to 480×640 pixels using the OpenCV image processing library. The collected images were semi-automatically annotated using a tool for semi-automatically annotating video training data based on a multi-object detector and tracker. Verification feedback was provided using the existing YOLOv8s model. The RectLabel annotation tool was then used to manually correct image annotations with poor detection results. This approach, using semi-automatic tools for preliminary annotation of image samples and manual correction of image annotations that have been validated by the existing model, significantly improves the speed and accuracy of annotation.

[0035] 3. Training set creation based on data augmentation

[0036] Based on the acquired images and annotated data, data augmentation is performed, including image rotation, translation, scaling, flipping, and color transformation, to generate diverse training samples and improve the model's robustness and generalization capabilities. Image samples are grouped and then randomly rotated, contrast adjusted, Gaussian noise added, and elastic transformations applied to increase the diversity of the training data by creating samples that are different but related to the original data.

[0037] With the help of the existing data augmentation library, Keras's ImageDataGenerator, we set the combination and parameters of the augmentation method, then perform augmentation operations on each image and its corresponding annotations, and then generate a sufficient number of augmented samples to ensure that the dataset contains a combination of original samples and augmented samples. Finally, the generated training set, validation set, and test set are divided and saved in a suitable format for model training.

[0038] 4. Improve the C2f module of the neck in YOLOv8s

[0039] In order to overcome the problem of insufficient feature extraction and gradient disappearance in the C2f module at the neck of the original YOLOv8s, and at the same time improve the accuracy of the model in detecting small targets and its generalization ability in complex scenarios, this patent proposes to replace the C2f module at the neck of the original YOLOv8s network with a module composed of a depthwise separable convolution layer, a residual block formed by stacking three Bottleneck layers, a depthwise separable convolution layer, and an improved SimAM attention mechanism layer, and add a DSCBS module to connect the split layer and the concat layer of C2f. Specifically, Figure 4As shown, in the improved YOLOv8s model, after the feature map enters the C2f module, it first passes through a DSCBS layer (depthwise separable convolution (DSConv)) followed by batch normalization (BN) and SiLU activation to enhance feature representation. The feature map is then split into two parts: one part is directly passed to subsequent layers, while the other part enters three Bottleneck layers for deeper feature extraction. These Bottleneck layers use residual connections to maintain the network's depth and allow for efficient gradient propagation within the deep network. The feature map processed by the Bottleneck layer is merged with the directly passed part along the channel dimension, which helps integrate feature information from different paths. The merged feature map then passes through a second DSCBS layer to further refine and enhance the features. Finally, the feature map passes through an improved SimAM attention mechanism layer, which uses an energy function inspired by neuroscience theory to calculate the importance of each neuron and implements this process quickly through a closed-form solution to generate 3D attention weights. These attention weights are used to adjust the feature map so that the model can focus more on the key areas in the image and improve the detection ability of small objects and complex scenes.

[0040] Among them, in YOLOv8s, the CBS of the C2f module is replaced with DSConv-BN-SiLU. The original standard convolution layer is replaced by depth-wise separable convolution. Each block of the depth-wise separable convolution consists of: first a 3x3 depth-wise convolution, followed by BN and Relu layers, followed by 1x1 point-by-point convolution, and finally BN and Relu layers. This convolution method effectively reduces the number of parameters and computational complexity by first performing depth-wise convolution on each input channel, and then combining the channels through 1x1 point-by-point convolution. Next, the batch normalization layer is retained to maintain the stability of network training, and the output after convolution is normalized. Finally, the SiLU activation function is retained to introduce nonlinearity to the network to help the model capture more complex features. The efficiency of feature extraction is improved, and network performance is maintained. The parameter and computational ratio formulas of the replaced depth-wise separable convolution are:

[0041]

[0042]

[0043] Among them, P is the number of convolution parameters, C is the amount of convolution calculation, and D K Represents the size of the convolution kernel, M and N represent the number of input and output channels respectively, and D F is the dimension of the feature map.

[0044] like Figure 4As shown in the figure, adding a cascade of improved SimAM attention mechanism layers after the final depthwise separable convolution block of the C2f module in the neck of YOLOv8s can significantly improve the model's feature representation capabilities. Specifically, by introducing a 3D attention mechanism that does not require additional parameters, the model's ability to detect small targets and complex scenes is enhanced while maintaining the model's lightweight and fast inference speed. SimAM achieves this by minimizing an energy function that takes into account the linear separability between the target neuron and other neurons. The energy function expression is:

[0045]

[0046] Among them, w t and b t Represent the weight and bias of the linear transformation, λ is the regularization term, and M is the number of neurons in the channel. Through this energy function, the closed-form solution of the weight and bias can be derived, and then the importance of each neuron can be calculated. Finally, the importance of each neuron is given by Given, where is the minimum value of the energy function, and are the mean and variance of all neurons respectively, and the expressions are:

[0047]

[0048]

[0049] Among them, E contains all ⊙ represents element-by-element multiplication, and the sigmoid function is used to limit the values ​​in E to ensure that the relative importance of neurons is not affected by excessive values. The improvement to the SimAM attention mechanism layer is to add a weight decay regularization operation to the SimAM module to preprocess the data input to the SimAM module, promoting the dispersion and distinguishability of features, and replacing the activation function from Sigmoid with Tanh to solve the non-zero centering problem of Sigmoid. Weight decay regularization is achieved by adding an additional term to the loss function. This additional term is a multiple of the L2 norm of the model weights, usually represented by λ (lambda), and the formula is as follows:

[0050] L regularization =λ∑ i ||w i || 2 (6)

[0051] Among them L regularization represents the regularization term, w iis the model's weight parameter, and λ is the regularization coefficient, which is used to control the influence of the regularization term. This operation helps promote feature dispersion by reducing the redundancy of neuron activations in the feature map, enhancing the model's sensitivity to different features, and thus improving feature distinguishability. At the same time, the activation function in the SimAM module is replaced from Sigmoid to Tanh (hyperbolic tangent function). Compared to Sigmoid, the Tanh activation function has the characteristic of non-zero centering, which means that Tanh's output values ​​are symmetrically distributed around zero, helping to address the problem of the Sigmoid activation function's output values ​​being biased towards positive numbers. This non-zero centering property can improve the stability of feature representation, making the model's gradient updates more stable during training, thereby effectively alleviating the vanishing gradient problem, especially in deep networks. Through these improvements, the SimAM module can more effectively improve the expressiveness of features while maintaining its lightweight, enabling YOLOv8s to demonstrate higher accuracy and robustness in object detection tasks, especially when dealing with small objects and complex backgrounds.

[0052] Adding a DSCBS module to connect the split and concat layers of C2f further enhances the model's feature extraction and learning, especially when dealing with complex scenes and small objects. This helps the model learn deeper features while maintaining network stability and efficiency. This structural improvement helps improve the accuracy and robustness of YOLOv8s in object detection tasks.

[0053] 5. Design the neck network structure and backbone network feature extraction network of YOLOv8s

[0054] In order to improve the detection capability of the YOLOv8s model for small targets in complex scenarios, the present invention improves the neck network structure and the backbone network feature extraction network. Specifically, based on the original YOLOv8s neck network, in order to enhance the characteristics of small targets, an Upsample-Extract feature extraction module is added at the top of the top-down aggregation path, and an Extract-Fusion feature fusion module is added at the bottom of the bottom-up aggregation path. The SPPF module in the original YOLOv8s backbone network is replaced with a cascade of STN layer, ghost convolution layer, STN layer, three parallel MaxPool2d layers, Concat layer, two parallel STN layers, and ghost convolution layer, and a residual connection is added after the first convolution layer to enhance the robustness of the model and the efficiency of feature extraction. Figure 2As shown in the gray box on the far left, the Upsample-Extract feature extraction module first concatenates feature maps of different scales using the Concat layer, followed by further feature extraction using the C2f module. Next, the feature maps are upsampled using the Upsample layer, concatenated again using the Concat layer, and finally refined using the improved C2f module. This design helps minimize gradient decay while improving network efficiency and enhancing detection accuracy for surface feature maps.

[0055] At the same time, if Figure 2 As shown in the gray box in the middle, the Extract-Fusion feature fusion module integrates multi-level feature information through the improved C2f module, then further refines the features through the convolution layer, and finally fuses the feature maps of different levels through the Concat layer to improve the sensitivity to key features.

[0056] Improved SPPF modules, such as Figure 5 As shown in the figure, by replacing the STN layer, ghost convolution layer, STN layer, three parallel MaxPool2d layers, Concat layer, two parallel STN layers, and ghost convolution layer cascade structure, efficient multi-scale feature extraction and adaptive spatial transformation are achieved. The residual connection added after the first convolution layer helps the deep network effectively learn the identity mapping, thereby enhancing the model's robustness in recognizing objects in complex scenes.

[0057] The STN consists of three main modules: Localisation Network, Grid Generator and Sampler.

[0058] The role of the Localisation Network is to generate the regression transformation parameter θ. θ is a 6-dimensional vector that can represent different types of spatial transformations, such as translation, rotation, scaling, etc., or more complex affine transformations or nonlinear transformations, defining the transformation from the input image to the output image. The role of the localization network can be expressed as:

[0059] θ=f loc (U) (7)

[0060] Where U is the input feature map, f loc It is a localized network.

[0061] The Grid Generator calculates the coordinates of the corresponding transformed image according to the transformation parameter θ, and creates a corresponding sampling position in the input feature map U for each pixel on the output feature map V. The formula is expressed as:

[0062] Grid=G(θ,H,W) (8)

[0063] where G is the grid generator, θ is the transformation parameter, and H and W are the height and width of the output image, respectively.

[0064] The Sampler samples pixel values ​​from the input image U according to the coordinate grid provided by the Grid Generator to generate the transformed output image V. Since the coordinates in the grid may be non-integer, bilinear interpolation is needed to estimate the pixel values ​​of these non-integer coordinates. The specific formula is as follows:

[0065]

[0066] Among them, V c,i,j is the grayscale value of the cth channel (i, j) point on the output feature map, U c,n,m is the grayscale value of the cth channel point (n,m) on the input feature map, x i,j and y i,j is the coordinate of the sampling point. The bilinear interpolation method calculates the value of the new pixel by taking the weighted average of the four surrounding pixels, where the weight is determined by the linear attenuation of the distance.

[0067] By replacing the convolutional layer with a ghost convolutional layer, the efficiency of feature extraction can be effectively improved. Ghost convolutional layers reduce computational complexity, allowing the network to maintain performance while reducing the consumption of computing resources.

[0068] Modules in the neck region, such as the C2f module, are similar to those in the backbone and are used for feature fusion and multi-scale information processing. These modules in the neck region further refine and integrate features from the backbone network to facilitate effective object detection at different scales. After processing the original input image in the backbone layer, a series of multi-scale feature maps are generated, including 19×19, 38×38, 76×76, and 152×152. The final feature map is obtained through the cascade of the C2f module, the feature pyramid structure, the Upsample, the SE module, and the CBS module, and the SPPF module with a pooling kernel of 5×5. The feature pyramid structure achieves bidirectional feature fusion by introducing bottom-up and top-down paths, generating multi-scale feature maps. Its structure includes a top-down path and a bottom-up enhancement path. The top-down part achieves the fusion of features at different levels by upsampling and fusing with coarser-grained feature maps, while the bottom-up part further enhances the feature representation by fusing feature maps from different levels through a convolutional layer. Specifically, the top-down part operates in the following three steps:

[0069] (1) Upsample the feature map of the last layer to obtain a finer feature map;

[0070] (2) Fusing the upsampled feature map with the finer-grained feature map from the previous layer, by adding or concatenating them, yields a feature representation containing richer information.

[0071] (3) Repeat the above two steps and gradually merge upwards until the highest level feature map is reached, forming a top-down feature pyramid;

[0072] The bottom-up part fuses feature maps from different levels through one or more convolutional layers and upsampling operations to enhance feature representation. This process is divided into the following three steps:

[0073] (1) Starting from the lowest level feature map, first perform feature extraction or adjust the dimension through a convolutional layer to obtain a deeper feature table;

[0074] (2) The processed feature map is adjusted to the same spatial resolution as the feature map of the previous layer through upsampling operation, and the two are fused;

[0075] (3) Repeat the above two steps, layer by layer, until the highest feature map is reached;

[0076] Finally, the feature maps of the top-down part and the bottom-up part are fused to obtain the final feature map for target detection and recognition.

[0077] 6. Improve the head network of YOLOv8s

[0078] like Figure 2 As shown in the gray box on the far right, the detection performance of small targets is improved by adding a decoupled multi-scale prediction of the small target detection head to the C2f module of the Extract-Fusion module at the neck of YOLOv8s and introducing an improved SE attention mechanism layer.

[0079] Taking into account the challenges of target detection in open road scenes, such as lighting changes, the influence of vehicle lights, and the diversity of different types of vehicles, this patent proposes to add a decoupled small target detection head for multi-scale prediction. Specifically, on the basis of the original YOLOv8s detection head, a small-scale detection head DDH is added, and an improved SE attention mechanism layer is introduced into the small-scale detection head DDH and the original multi-scale detection head of YOLOv8s. The improved SE attention mechanism layer refers to adding a cascade branch consisting of a convolution layer, a normalization layer, a LeakyReLU layer, a convolution layer, a normalization layer, and a Swish layer in parallel in the Excitation stage of the SE module, so that it can adaptively assign weights to each channel, thereby improving the detection performance of small targets. By adding an additional cascade structure, more complex features and relationships between channels can be captured. The introduction of LeakyReLU solves the dead ReLU problem and improves the adaptability of the model. Among them, the addition of the small-scale detection head DDH enables the use of shallower high-resolution feature maps to calculate the prediction results of small targets.

[0080] The SE attention mechanism layer enables the network to focus more on the key information in the image, improving the accuracy of target detection. Due to its high computational efficiency, it is applicable to convolutional neural networks of various sizes. The specific formula is as follows:

[0081]

[0082] X'=S×X (11)

[0083] where Z c represents the global average value of the c-th channel, u c is the c-th channel of the input feature map, H and W are the height and width of the feature map respectively, W1 and W2 are the weight matrices of the fully connected layer, δ represents the ReLU activation function, σ represents the sigmoid activation function, z is the channel descriptor obtained by operation (1), and X' is the feature map after adjustment by the attention mechanism.

[0084] This characteristic of the Swish function enables it to perform better in deep neural networks, especially when processing input data with complex features. The specific formula is as follows:

[0085] Swish(x)=x·σ(βx) (12)

[0086] Where σ represents the sigmoid function and β is a trainable parameter used to control the nonlinearity of the Swish function.

[0087] The Leaky ReLU described above helps maintain the network's deep learning capabilities and maintains effective training even when faced with input data with complex features. The specific formula is as follows:

[0088]

[0089] Here, x is the input and α is a small positive number.

[0090] Based on the original YOLOv8s detection head, a micro-scale detection head DDH is added. For the detection and recognition tasks of small targets in open road scenes, it can specifically enhance the network's perception ability and recognition accuracy of small targets, better understand the relationship between the target and the surrounding environment, increase the flexibility and scalability of the network, and thus significantly improve the ability to detect and recognize small targets in open road scenes.

[0091] An improved SE attention mechanism layer is introduced in the YOLOv8s head. Specifically, the improved SE attention mechanism layer is introduced in the small-scale detection head DDH and the original multi-scale detection head of YOLOv8s. The SE attention mechanism layer proposes a new channel attention method, which improves the performance of convolutional neural networks by adaptively recalibrating channel features. Specifically, the SE module first integrates the spatial information of each channel through a global average pooling (Squeeze) operation to generate a channel descriptor. Then, two fully connected layers (Excitation) are applied to the descriptor to capture the dependencies between channels, thereby generating a weight coefficient for each channel.

[0092] Because the SE attention mechanism layer adaptively recalibrates channel features, allowing the network to focus on important feature representations, it performs well in improving model performance while maintaining high computational efficiency. In particular, in image recognition tasks with complex backgrounds and diverse objects, the SE layer can intelligently adjust the weights of each channel to highlight key features and suppress unimportant information.

[0093] 7. Train the improved YOLOv8s deep network and input the test video for network forward reasoning to detect targets in open road scenes

[0094] Use the Adam algorithm to train the improved YOLOv8s deep network;

[0095] Compared with the prior art, the present invention has the following beneficial effects:

[0096] (1) The present invention replaces the C2f module at the neck of the original YOLOv8s network with a module obtained by connecting a depthwise separable convolutional layer, a residual block formed by stacking three Bottleneck layers, a depthwise separable convolutional layer, and an improved SimAM attention mechanism layer in series, and adds a DSCBS module to connect the split layer and the concat layer of C2f to integrate multi-level feature information, improve the sensitivity to key features, enhance the ability to recognize high-level semantic features of small targets, alleviate the problems of deep gradient disappearance and gradient explosion, and replace the convolutional layer with a depthwise separable convolutional layer to reduce the number of model parameters.

[0097] (2) Based on the original YOLOv8s neck network, the present invention adds an Upsample-Extract feature extraction module at the top of the top-down aggregation path and an Extract-Fusion feature fusion module at the bottom of the bottom-up aggregation path to enhance the features of small targets. These cascades enhance the detection accuracy of surface feature maps while minimizing gradient attenuation and improving the efficiency of the network structure.

[0098] (3) The present invention replaces the SPPF module in the original YOLOv8s backbone network with a cascade of an STN layer, a ghost convolution layer, an STN layer, three parallel MaxPool2d layers, a Concat layer, two parallel STN layers, and a ghost convolution layer, and adds a residual connection after the first convolution layer to efficiently extract multi-scale features, adaptively perform spatial transformation on the feature map, enhance the robustness of the model in identifying targets in complex scenes, enable the deep network to effectively learn the identity mapping, and replace the convolution layer with a ghost convolution layer to effectively improve the feature extraction efficiency.

[0099] (4) Based on the original YOLOv8s detection head, the present invention adds a micro-scale detection head DDH to the C2f module of the Extract-Fusion module at the neck of YOLOv8s, so that the model can use a shallower high-resolution feature map to calculate the prediction results of small targets, and introduces an improved SE attention mechanism layer in the micro-scale detection head DDH and the original multi-scale detection head of YOLOv8s. At the same time, in the Excitation stage of the SE module, a cascade branch consisting of a convolution layer, a normalization layer, a LeakyReLU layer, a convolution layer, a normalization layer, and a Swish layer is added in parallel, so that it can adaptively assign weights to each channel, thereby improving the detection performance of small targets. By adding an additional cascade structure, more complex features and relationships between channels can be captured, and LeakyReLU is introduced to solve the dead ReLU problem, thereby improving the adaptability of the model.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific implementation methods of the present invention can still be modified or replaced without departing from the spirit and scope of the present invention. Any modification or replacement should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for detecting targets in open road scenes based on thermal infrared recognition, characterized in that: Improve the C2f module of the neck network in YOLOv8s, that is, CSP Bottleneck with 2 convolutions; then add an Upsample-Extract feature extraction module and an Extract-Fusion feature fusion module to the neck network, and replace the SPPF, that is, Spatial Pyramid Pooling-Fast module with STN, that is, Spatial Multi-layer cascade of TransformerNetworks layer and ghost convolution layer; finally, by adding decoupled small target detection head, integrating improved C2f module and SE attention mechanism layer; specifically including the following steps: (1) collecting image samples of the forward view scene of thermal infrared driving recorder; (2) preprocessing and labeling the image samples; (3) making training set based on data augmentation; (4) improving the C2f module of the neck in YOLOv8s; (5) designing the neck network structure and backbone network feature extraction network of YOLOv8s; (6) improving the head network of YOLOv8s; (7) training the improved YOLOv8s deep network, inputting the test video for network forward reasoning to realize the detection of targets in open road scenes; the said collection of image samples of the forward view scene of thermal infrared driving recorder refers to collecting images of cars driving or parked on the road at different angles and distances under different urban and suburban conditions. The thermal infrared image data of the vehicle; the preprocessing and labeling of the sample images refers to adjusting all images to a uniform size of 480×640 pixels using the OpenCV image processing library, semi-automatically labeling the collected images using a semi-automatic labeling video training data tool based on a multi-target detector and tracker, and verifying feedback using the existing YOLOv8s model, and then manually labeling and correcting the image labeling samples with poor detection results using the RectLabel labeling tool; the preparation of the training set based on data augmentation refers to grouping the image samples, randomly rotating them, adjusting the image contrast, adding Gaussian noise to the images, and applying elastic transformation; the improved YOLOv8s deep network refers to the deep network obtained by improving steps (4), (5), and (6), and the training of the improved YOLOv8s deep network refers to completing the training of the deep network using the Adam algorithm.

2. The method for detecting targets in open road scenes based on thermal infrared recognition according to claim 1, characterized in that: The improved C2f module in the neck of YOLOv8s refers to replacing the C2f module in the neck of the original YOLOv8s network with a depthwise separable convolution layer, a residual block formed by stacking 3 Bottleneck layers, a depthwise separable convolution layer, and an improved SimAM, that is, a Simple Parameter-Free Attention Module attention mechanism layer connected in series, and adding a DSCBS, that is, a Depthwise Separable Convolution-Batch Normalization-SiLU module to connect the split layer and the concat layer of C2f to integrate multi-level feature information, improve sensitivity to key features, enhance the ability to recognize high-level semantic features of small targets, alleviate the problems of deep gradient disappearance and gradient explosion, and replace the convolution layer with a depthwise separable convolution layer to reduce the number of parameters of the model; the DSCBS refers to the existing CBS, that is, Convolution-Batch The convolution layer of the Normalization-SiLU module is replaced with a depth-wise separable convolution layer; the improved SimAM attention mechanism layer refers to adding a weight decay regularization operation to the SimAM module to preprocess the data input to the SimAM module, and replacing the activation function from Sigmoid to Tanh.

3. The method for detecting targets in open road scenes based on thermal infrared recognition according to claim 1, characterized in that: The design of the neck network structure and backbone network feature extraction network of YOLOv8s refers to adding an Upsample-Extract feature extraction module at the top of the top-down aggregation path and an Extract-Fusion feature fusion module at the bottom of the bottom-up aggregation path on the basis of the original YOLOv8s neck network to enhance the features of small targets, and replacing the SPPF module in the original YOLOv8s backbone network with a cascade of an STN layer, a ghost convolution layer, an STN layer, three parallel MaxPool2d layers, a Concat layer, two parallel STN layers, and a ghost convolution layer, and adding a residual connection after the first convolution layer; the Upsample-Extract feature extraction module is a cascade of a Concat layer, a C2f module, an Upsample layer, a Concat layer, and a C2f module; the Extract-Fusion feature fusion module is a cascade of a C2f module, a convolution layer, and a Concat layer.

4. The method for detecting targets in open road scenes based on thermal infrared recognition according to claim 1, characterized in that: The improved YOLOv8s head network refers to adding a multi-scale prediction of a decoupled small target detection head to the C2f module of the Extract-Fusion module at the neck of YOLOv8s, and introducing an improved SE, that is, a Squeeze-and-Excitation Networks attention mechanism layer to improve the head network of YOLOv8s; the added multi-scale prediction of the decoupled small target detection head refers to adding a small-scale detection head DDH, that is, a Decoupled DetectionHead, to the C2f module of the Extract-Fusion module at the neck of YOLOv8s on the basis of the original YOLOv8s detection head, and introducing an improved SE attention mechanism layer into the small-scale detection head DDH and the original multi-scale detection head of YOLOv8s; the improved SE attention mechanism layer refers to adding a parallel branch consisting of a cascade of a convolutional layer, a normalization layer, a LeakyReLU layer, a convolutional layer, a normalization layer, and a Swish layer in the Excitation stage of the SE module.