An Infrared Image Vehicle Detection Method Based on a Single-Stage Network

Through the infrared image vehicle detection method based on a single-stage network, combined with ShuffleNetV2 and the dual-branch adaptive fusion attention module, the detection head network is optimized, and the problem of insufficient detection capability of infrared image vehicle detection in complex environments is solved, and high-precision and real-time multi-scale vehicle object detection is achieved.

CN116543228BActive Publication Date: 2025-07-25TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310578973.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-22
Publication Date
2025-07-25
Estimated Expiration
2043-05-22

AI Technical Summary

Technical Problem

The existing infrared image vehicle detection technology lacks detection capabilities in complex environments, and has problems such as poor generalization and poor real-time performance, and traditional methods are difficult to adapt to multi-scale vehicle targets.

Method used

The infrared image vehicle detection method based on a single-stage network is adopted, and the ShuffleNetV2 lightweight network is combined with the central differential convolution and the dual-branch adaptive fusion attention module to improve detection accuracy through feature extraction and fusion networks, and optimize the detection head network through task alignment ideas to achieve real-time accurate detection.

Benefits of technology

It improves the accuracy and anti-interference ability of vehicle detection in infrared scenarios, and can detect multi-scale vehicle targets in complex environments, with low model complexity and real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543228B_ABST
    Figure CN116543228B_ABST
Patent Text Reader

Abstract

The present invention relates to an infrared image vehicle detection method based on a single-stage network, which includes the following steps: Image acquisition: An infrared camera performs field of view scanning to obtain a video stream and convert it into an image format; Image preprocessing; Network detection: The infrared image is sent into a trained vehicle detection network for vehicle target detection. Feature extraction is performed through the backbone network, and feature fusion is performed using the optimized fusion network. The vehicle target category and position information are predicted through the classification and regression branches of the vehicle detection network; The vehicle detection network is a single-stage network structure, and the RetinaNet network is used as the basic network for optimization and improvement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of vehicle detection in infrared image processing technology, and particularly relates to an infrared image vehicle detection method based on a single-stage network. Background Art

[0002] As a basic means of transportation, vehicles play an important role in various fields such as residents' travel, urban logistics, industrial production, and national defense and military. With the improvement of social production capacity and residents' consumption level, the usage of vehicles in various industries in China has also shown a rapid growth trend. The growth of the number of vehicles has promoted the demands in aspects such as urban traffic planning, intelligent logistics transportation, and military tracking tasks. Especially in recent years, with the emergence of concepts such as "intelligent transportation", the detection technology for vehicle targets has gradually come into people's view and become a research hotspot in the scientific research field.

[0003] Early vehicle detection work was usually based on sensor and radar detection technologies. The detection system detected and judged through signals such as sound and electromagnetic waves triggered and fed back by vehicle targets. The above technical principles are simple, but often have many problems such as poor concealment, weak anti-interference ability, and high installation and maintenance costs, and it is difficult to be widely applied to actual vehicle detection tasks. In recent years, with the rapid development of infrared imaging technology, it has gradually attracted the attention of researchers in various fields. This technology uses the thermal radiation of the object itself to form an image, is not easily affected by harsh environments such as wind, frost, rain, and snow, and has advantages such as strong anti-interference ability, good concealment, and wide coverage area compared with early technical means. In addition, compared with visible light imaging technology, infrared imaging technology does not need to rely on external ambient light, is not affected by day and night, and can work 24 hours a day. Based on the above advantages, infrared imaging vehicle detection technology has begun to be gradually applied to fields such as autonomous driving, urban management, tracking guidance, and military reconnaissance.

[0004] Traditional infrared image vehicle detection technology mainly judges targets based on vehicle features and machine learning ideas. This type of technology usually relies on manually extracted features, has poor generalization ability, and is difficult to adapt to complex detection scenarios. In recent years, with the improvement of the level of intelligence and informatization, the field of deep learning has developed rapidly. The object detection technology based on deep learning can automatically extract object features using a convolutional neural network, so as to obtain richer feature information, and has better detection effects and higher working efficiency compared with traditional methods, so it is highly favored by researchers. Although deep learning technology has demonstrated powerful advantages in visual image detection tasks, for vehicle targets with scarce feature information in infrared scenarios, its detection ability still needs to be further improved. In addition, most current excellent object detection models have problems such as high complexity and poor real-time performance. Summary of the Invention

[0005] In view of the limitations and deficiencies of the prior art, the present invention fully considers the characteristics of vehicle targets in infrared images and the actual detection requirements, and provides a new vehicle detection method based on a single-stage network structure, which can realize real-time and accurate vehicle detection, and fully meet the needs of vehicle detection tasks in actual infrared scenarios. The technical solution of the present invention is: an infrared image vehicle detection method based on a single-stage network, including the following steps:

[0006] An infrared image vehicle detection method based on a single-stage network includes the following steps:

[0007] Step 1, Image acquisition: The infrared camera performs a field scan to obtain a video stream and converts it into an image format;

[0008] Step 2, Image processing: Preprocess the obtained infrared image;

[0009] Step 3, Network detection: Send the infrared image into a pre-trained vehicle detection network for vehicle target detection. Feature extraction is performed through the backbone network, and feature fusion is performed using an optimized fusion network. The vehicle target category and position information are predicted through the classification and regression branches of the vehicle detection network; The vehicle detection network is a single-stage network structure, based on the RetinaNet network, and is optimized and improved, including:

[0010] 1) Select the ShuffleNetV2 lightweight network as the backbone network, use the local information correlation of the infrared image to assist target recognition, to improve the network's extraction effect on invariant fine-grained information. Design a convolution method that combines central difference convolution and convolution operation, and embed this convolution method into the core module ShuffleBlock of the backbone network, dynamically adjust the weight ratio of the two convolutions for feature extraction. The convolution kernel slides and scans the feature map for sampling, extracts the pixel points in the area corresponding to the convolution kernel before aggregation, and calculates the difference between the pixel value of its center point and the values of the other pixel points in turn to obtain the updated pixel value, and then perform a dot product aggregation of the pixel value and the convolution kernel weight to obtain the final output value;

[0011] 2) A dual-branch adaptive fusion channel attention module DBAM is constructed in the vehicle detection network. The global information and local information of the infrared image are utilized through two parallel branches of global average pooling and global maximum pooling. Dynamic one-dimensional convolution is used to generate channel weights to better complete information interaction between channels; The method of dual-branch adaptive fusion is used to dynamically adjust the weight ratio of the two pooling branches in the fusion; The specific implementation process of DBAM is: First, perform global average pooling and global maximum pooling operations on the input feature with height, width and number of channels of H×W×C in the channel dimension, and then generate the channel weight matrices M of the two branches through fast one-dimensional convolution with a convolution kernel of size kavg and M max ; The two-channel weight matrices are used to obtain a summarized channel weight matrix through an adaptive fusion structure, and through element-wise multiplication with the original input features, the weights are mapped into a feature map of H×W×C;

[0012] 3) Using the ShuffleBlock module with central difference convolution as the basic unit, and embedding a dual-branch channel attention module at the same time to construct the backbone network; The feature information with different depth levels extracted from each stage of the lightweight network ShuffleNetV2 is fed into the optimized fusion network. The optimized fusion network is based on the feature pyramid structure and adopts a bidirectional cross fusion network structure. A bottom-up aggregation path is additionally added in some feature layers to improve the fusion effect, and a lateral connection is added between the original input and the final output nodes of the feature layers at the same scale; A fast normalization fusion strategy is used to add additional weights to different input nodes to distinguish the contribution degrees of different nodes;

[0013] 4) Design a detection head network based on the task alignment idea, called the task alignment head TAHead: It consists of a feature extractor and two optimized task branches. The feature extractor is used to perform multi-level extraction on the output features of the feature fusion network. Each task branch contains a calibration branch, which is used to perform probability adjustment and spatial adjustment on the preliminary prediction results;

[0014] 5) Design a network training and learning strategy based on the task alignment idea; By setting a calibration factor to clarify the alignment degree between the classification and regression tasks; The calibration factor is introduced into the loss calculation, and by synchronously optimizing the classification and regression losses, the prediction results of the two tasks for the same sample point are adjusted to improve the spatial misalignment problem.

[0015] Furthermore, in 2), the convolution kernel size k represents the range of cross-channel interaction, and there is a non-linear mapping relationship between it and the channel dimension C. The calculation formula is as follows:

[0016]

[0017] In the formula, |t| odd represents the nearest odd number of t, and γ and b are two adjustable parameters.

[0018] Furthermore, in 2), the summarized channel weight matrix is:

[0019]

[0020] In the formula, δ is the sigmoid activation function, and μ and ν are two floating-point learnable parameters, which are dynamically learned with the network model and the initial values are set to 1; and They are element-wise addition of matrices and element-wise multiplication of matrices respectively.

[0021] Furthermore, 4) includes:

[0022] Using a feature extractor to perform N convolutional operations on the output of feature fusion to obtain a multi-layer task interaction feature stack, which serves as the common feature basis for the two task branches; the two improved task branches are the classification branch and the regression branch respectively. Among them, the classification task branch first performs concat splicing and convolutional operations on the task interaction feature stack, and obtains a dense classification score of H×W×1 through the sigmoid activation function as the preliminary classification prediction result; the regression task branch also obtains a regression bounding box score of H×W×4 through concat splicing and convolutional operations as the preliminary regression prediction result.

[0023] Furthermore, 4) also includes: In TAHead, a parallel calibration branch is constructed for each of the two tasks to clearly adjust the preliminary prediction results of the two tasks; the predictions of the two task branches are adjusted simultaneously through the obtained task interaction features and subsequent task alignment learning strategies; the two calibration branches use the task interaction features to generate a spatial probability map and a spatial offset map. The spatial probability map learns the prediction consistency between the two tasks at each spatial position through the backpropagation process, and then adjusts the preliminary prediction result of the classification task to obtain the final classification result; the spatial offset map learns the spatial offset between the current anchor box and the surrounding best anchor box through the backpropagation process, and then adjusts the preliminary prediction result of the regression task to obtain the final regression result.

[0024] Furthermore, in 5), the calibration factor is obtained through the classification and regression tasks, and its calculation process is as follows:

[0025] t = s α × u β

[0026] In the formula, t is the calibration factor, s and u respectively represent the classification score and IOU value obtained by the classification and regression tasks for each anchor box, and α and β are set to control the influence of the two tasks on the calibration factor; select the m anchor boxes with the largest t value as positive samples, and the rest as negative samples.

[0027] Furthermore, in 5), the Focal Loss is used as the classification loss function to alleviate the imbalance between positive and negative samples. On this basis, the original positive sample anchor box label value is replaced with the calibration factor to improve the classification score of the anchor boxes with a higher alignment degree.

[0028] For the backbone network part of the present invention, the lightweight network ShuffleNetV2 that takes into account both accuracy and speed is used. Its core module is optimized based on central difference convolution to enhance the network's anti-interference ability against environmental changes. And at each stage of it, a dual-branch adaptive fusion attention module is embedded, which significantly improves the backbone network's ability to extract local detail features and overall feature information in infrared images while maintaining real-time performance; for the fusion network part, a bidirectional cross network structure is constructed on the basis of the original FPN, and the multi-scale feature information is enriched through a fast normalization fusion strategy, comprehensively improving the network's detection ability for multi-scale vehicle targets; the detection head network constructed based on task interaction features can effectively enhance the information interaction between the classification and regression tasks during training and detection, and reduce the spatial misalignment phenomenon of the prediction results of the two tasks; during the training process of the network, the sample allocation and loss calculation are optimized based on the alignment degree of the two tasks, further aligning the classification and regression prediction results, and comprehensively improving the detection effect of the network. The optimized and improved vehicle detection network has good detection performance and is more suitable for vehicle detection tasks in infrared scenarios. The present invention not only has high detection accuracy and good detection effects for vehicle targets of various scales and forms under complex environmental backgrounds, but also has strong anti-interference ability. At the same time, the model complexity is low, and it can achieve real-time detection processes, having good comprehensive performance and application value. Description of the Drawings

[0029] Figure 1 It is a schematic diagram of the overall details of the present invention;

[0030] Figure 2 It is a schematic diagram of the detection process of the present invention;

[0031] Figure 3 It is a schematic diagram of the overall structure of the vehicle detection network in the present invention;

[0032] Figure 4 It is a structural diagram of the ShuffleBlock module in the present invention;

[0033] Figure 5 It is a structural diagram of the dual-branch adaptive fusion attention module in the present invention;

[0034] Figure 6 It is a structural diagram of the optimized feature fusion network in the present invention;

[0035] Figure 7 It is a structural diagram of the detection head network based on task interaction features in the present invention;

[0036] Figure 8 It is a schematic diagram of the task alignment learning strategy in the present invention; Detailed Implementation Modes

[0037] The technical solution of the present invention is as follows: An infrared image vehicle detection method based on a single-stage network, comprising the following steps:

[0038] Step 1, image acquisition: Automatically or manually control an infrared camera through an interaction terminal to perform a field scan, obtain a real-time video stream, and convert it into an image format in a computer.

[0039] Step 2, image processing: Preprocess the obtained real-time infrared image using median filtering technology to weaken the influence of irrelevant noise.

[0040] Step 3, network detection: Send the real-time image into a pre-trained vehicle detection network for vehicle target detection. First, perform feature extraction through the backbone network, then use the optimized fusion network for feature fusion, and finally predict the vehicle target category and position information through the classification and regression branches of the detection network. The overall vehicle detection network is a single-stage network structure, based on the RetinaNet network, and adaptively optimized and improved according to the characteristics of vehicle targets in infrared images and actual vehicle detection requirements. Its main innovations include:

[0041] 6) Select the ShuffleNetV2 lightweight network as the backbone network, and use the local information correlation of the image to assist target recognition, better improving the network's extraction effect on invariant fine-grained information. Specifically, a convolution method combining central difference convolution and traditional convolution operation is designed. Taking a 3×3 convolution kernel as an example, the convolution kernel slides and scans the feature map for sampling. Before aggregation, 9 pixel points in the corresponding area of the convolution kernel are extracted, and the pixel value of the center point is sequentially differentiated from the values of the other pixel points to obtain the updated pixel value. Then, the pixel value is dot-product aggregated with the convolution kernel weight to obtain the final output value. The calculation process is as follows:

[0042]

[0043] Among them, x is the input feature map, w is the weight value, p0 is the position of the center of the convolution kernel on the current input feature map, and p n is used to enumerate the positions in R. For example, the local area R of a 3×3 convolution kernel is {(-1, -1), (-1, 0),... (1, 1)}. The hyperparameter θ ∈ (0, 1) is used to balance the contributions of standard convolution and difference convolution to the network. Embed the new convolution method into the core module ShuffleBlock of the backbone network, and dynamically adjust the weight ratio of the two convolutions for feature extraction, so as to better adapt to the lack of detailed information in infrared images and the influence of environmental changes, enhancing the backbone network's ability to extract detailed information without bringing additional computational complexity and improving the network detection effect;

[0044] 7) To further improve the feature extraction ability of the backbone network, a channel attention module with dual-branch adaptive fusion, namely DBAM (Dual-Branch Attention Module), is constructed in the network. By means of the strategy of two parallel branches of global average pooling and global maximum pooling, the global information and local information of infrared images are comprehensively utilized. To avoid a large number of parameters brought by the fully connected layer, dynamic one-dimensional convolution is adopted to generate channel weights to better complete the information interaction between channels. Finally, the method of dual-branch adaptive fusion is used to dynamically adjust the weight ratio of the two pooling branches in the fusion. The specific implementation process of DBAM is as follows: First, global average pooling and global maximum pooling operations are respectively performed on the input features with height, width, and number of channels of H×W×C in the channel dimension, and then the channel weight matrices M avg and M max of the two branches are generated through fast one-dimensional convolution with a convolution kernel of size k. The two channel weight matrices complete the summary of channel weights through the adaptive fusion structure, and are element-wise multiplied with the original input features to map the weights into the feature map of H×W×C.

[0045] Among them, the convolution kernel size k represents the range of cross-channel interaction, and there is a non-linear mapping relationship between it and the number of channels C. It does not need to be manually adjusted through cross-validation. The calculation process formula is as follows:

[0046]

[0047] |t| odd represents the nearest odd number of t, and γ and b are two adjustable parameters used to ensure the best matching relationship between the convolution kernel size k and the channel C, and the two parameter values are set to 2 and 1 in accordance with the original idea of ECA to ensure the best cross-channel information interaction effect of the network.

[0048] The two feature branches generate the channel attention weight matrices M avg and M max through fast one-dimensional convolution, and the adaptive fusion method is adopted to obtain the finally summarized channel weight matrix:

[0049]

[0050] In the formula, M F is the final channel weight matrix, δ is the sigmoid activation function, and μ and ν are two floating-point learnable parameters that are dynamically learned with the network model and the initial values are set to 1. To prevent the two parameters from decaying to 0 and thus causing the loss of feature information, the weights of 1 / 2 of each of the two branches are added in the fusion. and They are element-wise addition of matrices and element-wise multiplication of matrices respectively. The aggregated channel weights are multiplied element-wise with the original input feature F, and finally the new feature F' corrected by the channel attention module is obtained:

[0051]

[0052] Compared with other classical attention mechanisms, the dual pooling branches better aggregate the global and local key features in the image, enriching the image information. This attention module is embedded into the outputs of each stage of the backbone network to make up for the deficiency of the backbone network in extracting the overall feature information of infrared images.

[0053] 8) The backbone network is constructed with the ShuffleBlock module after introducing central difference convolution as the basic unit, and at the same time, the dual-branch channel attention module is embedded to construct the backbone network. The specific parameters of the network structure are shown in the following table. Among them, stride 1 and 2 represent the standard module and the downsampling module of ShuffleBlock respectively, and the repeated stacking times of the modules at different stages are also reflected in the table.

[0054] Backbone network structure parameters

[0055]

[0056] 9) The feature information with different depth levels extracted from each stage of the lightweight network ShuffleNetV2 is fed into the optimized fusion network. Based on the traditional feature pyramid structure, the optimized fusion network designs a two-way cross fusion network structure, adds a bottom-up aggregation path in some feature layers to improve the fusion effect, and adds a lateral connection between the original input and the final output nodes of the same-scale feature layers. In addition, the fast normalization fusion strategy is adopted to add additional weights to different input nodes, more delicately distinguishing the contribution degrees of different nodes and enhancing the network's expression ability for multi-scale vehicle target features; taking the fourth-layer feature output as an example, the fast normalization fusion formula is:

[0057]

[0058]

[0059] Among them, w1, w2, w'1, w'2, w'3 are the additional weights of each feature input node, M4, M5 are the intermediate features of the 4th and 5th layers of the top-down path, P3, P4 are the output features of the 3rd and 4th layers of the bottom-up aggregation path. Among them, Resize in the top-down and bottom-up paths is realized through upsampling and downsampling operations respectively, and at the same time, the parameter ε = 0.0001 is set to ensure numerical stability.

[0060] 10) A detection head network based on the task alignment idea is designed. The improved network is called Task Aligned Head (TAHead), and its overall structure is as follows Figure 7 shown. It consists of a feature extractor and two optimized task branches. The feature extractor is used to perform multi-level extraction on the output features of the feature fusion network, and each task branch contains a calibration branch to perform probability adjustment and spatial adjustment on the preliminary prediction results.

[0061] First, the feature extractor performs N convolutional operations on the output of the feature fusion to obtain a multi-layer task interaction feature stack, which serves as the common feature basis for the two task branches. This design provides multi-scale effective receptive fields and multi-level features for the two tasks, promoting information interaction between the two tasks. The calculation process of the task interaction features is shown in the following formula:

[0062]

[0063] In the formula, k ∈ {1, 2,..., N}, X fpn is the output feature of the feature fusion network, is the k-th layer of the task interaction feature, conv k and δ represent the k-th convolutional layer and a ReLU function respectively. The features of each layer are obtained by performing one convolution and ReLU activation on the features of the previous layer, and finally, more abundant multi-level feature information is extracted from the single-level features. Subsequently, the network performs preliminary classification and localization on each sample point based on the calculated task interaction feature stack.

[0064] The two improved task branches are the classification and regression branches respectively, as shown in Figure 7 . Among them, the classification task branch performs concat splicing and convolutional operations on the task interaction feature stack in sequence, and obtains a dense classification score of H×W×1 through the sigmoid activation function as the preliminary classification prediction result. The regression task branch also obtains a regression bounding box score of H×W×4 through concat splicing and convolutional operations as the preliminary regression prediction result.

[0065] To further improve the spatial misalignment problem of the two task prediction results, a parallel calibration branch is constructed for each of the two tasks in TAHead to explicitly adjust the preliminary prediction results of the two tasks. The prediction of the two task branches is adjusted simultaneously through the obtained task interaction features and subsequent task alignment learning strategies. The specific operation process is that the two calibration branches use the task interaction features to generate a spatial probability map and a spatial offset map, and their generation processes are expressed by the following formulas:

[0066]

[0067]

[0068] in It is the interactive feature stack of the entire task. Spm and Som are the spatial probability map and spatial offset map on the calibration branch, respectively. conv1 and conv3 are both 1×1 dimensionality reduction convolution operations. δ is the ReLU activation function, σ is the sigmoid activation function, and conv2 and conv4 are 3×3 convolution operations, which are used to further differentiate the spatial probability map and spatial offset map of the two branches. The spatial probability map learns the prediction consistency between the two tasks at each spatial position through the back propagation process, and then adjusts the preliminary prediction results of the classification task to obtain the final classification result. At the same time, the spatial offset map learns the spatial offset between the current anchor box and the surrounding best anchor box through the back propagation process, and then adjusts the preliminary prediction results of the regression task to obtain the final regression result. The adjustment process of the two calibration branches on the preliminary prediction results is shown in the following formula:

[0069]

[0070] R align (i,j,c)=R ori (i+Som(i,j,2×c),j+Som(i,j,2×c+1),c)

[0071] In the formula, C align and R align is the final prediction result of the classification and regression tasks after adjustment by the calibration branch, C ori and R ori The index (i, j, c) represents the first prediction result of the classification branch and the regression branch. c The (i,j)th spatial position of the channel, due to R align The characteristic size of is very small and its computational overhead is negligible. align Each channel in the proposed method is learned independently, and each boundary of the current anchor box has its own learned offset, allowing the network to make more accurate predictions of the four spatial values, so that TAHead can not only promote the prediction alignment of the two tasks, but also accurately learn each boundary value of the anchor box, thereby further improving the regression accuracy.

[0072] 11) A network training and learning strategy based on the idea of task alignment is designed. The alignment degree of the classification and regression tasks is clearly measured by setting the calibration factor. The calibration factor is integrated into the sample allocation strategy and loss function of the vehicle detection algorithm to dynamically refine the prediction on each anchor box, thereby guiding the network to adjust the prediction results. Among them, the calibration factor is obtained through classification and regression tasks, and its calculation process is as follows:

[0073] t = s α × u β

[0074] In the formula, t is the calibration factor, s and u respectively represent the classification score and IOU value obtained for each anchor box in the classification and regression tasks. The influence of the two tasks on the calibration factor is controlled by setting α and β. Select the m anchor boxes with the largest t value as positive samples, and the rest as negative samples. The calibration factor t, as a high-order combination of the classification and regression results, represents the alignment degree of the prediction results of the two tasks. Therefore, using it as the sample evaluation criterion can make the network pay more attention to the sample boxes with a high alignment degree, and finally obtain detection results with excellent classification and regression quality.

[0075] Introduce the calibration factor into the loss calculation, and adjust the prediction results of the two tasks for the same sample point by synchronously optimizing the classification and regression losses to improve the spatial misalignment problem. Specifically, use the Focal Loss as the classification loss function to alleviate the imbalance between positive and negative samples, and on this basis, replace the original positive sample anchor box label value with the calibration factor, so as to increase the classification score of the anchor box with a higher alignment degree. The final classification loss and regression loss calculation formulas are:

[0076]

[0077]

[0078] The above classification loss is composed of the loss calculations of positive and negative samples. Among them, BCE is the binary cross-entropy loss function, i represents the i-th anchor box among the N positive samples corresponding to each instance, t pos represents the calibration factor t value corresponding to this anchor box, j represents the j-th anchor box among the N i negative samples, s neg ,s i ,s j is the predicted value, and γ is the focusing parameter in the Focal Loss. Similar to the classification loss, use the GIOU Loss as the regression loss function, and re-weight the regression loss for each anchor box based on t to obtain the final regression loss L reg ,in the formula t i also represents the calibration factor corresponding to the current anchor box, b i and b' i respectively represent the i-th predicted bounding box and the ground truth box.

[0079] Step 4: Detection result feedback. If the predicted result of the network is that there is a vehicle target within the field of view, then identify this vehicle target in the display device, and feedback the vehicle target position information to the host computer, calculate the corresponding turntable parameters, and the control end controls the camera turntable angle automatically or manually by the user to further adjust the camera field of view. The publicly programmable interface parameter information in the system is shown in the following table:

[0080] Partial interface parameter information in the system

[0081]

[0082] For further illustration, the specific implementation details of the present invention will be described in detail in the form of the accompanying drawings below, but they cannot be understood as limiting the protection scope of the present invention.

[0083] As Figure 1 shown, an infrared image vehicle detection method based on a single-stage network, the overall implementation details are as follows:

[0084] 1) Construction of the dataset and training of the network model: Since there is a lack of high-quality publicly available datasets in the field of infrared scene vehicle detection, a vehicle dataset containing 10,807 infrared images is self-built by integrating the existing resources of the laboratory. The data covers different time periods and weather conditions to ensure the objective diversity of the external environment, and at the same time contains vehicle targets with different angles and sizes. Based on this dataset, the optimized single-stage vehicle detection network is trained. To facilitate the distinction of target scales, the targets with the detection box area accounting for less than 1% of the image area are defined as small targets, those with a ratio greater than 1% and less than 4% are medium targets, and those greater than 4% are large targets. After statistics, the numbers of small, medium, and large targets in the entire dataset are 20,187, 18,663, and 13,884 respectively. During the training process, through a task-aligned learning strategy, sample allocation and loss calculation are carried out from the perspective of joint optimization, guiding the network to pay more attention to the results with higher classification and regression alignment, and dynamically guiding the training process of the network. At the same time, to prevent overfitting, data enhancement methods such as random flipping and brightness transformation are used to process the input pictures.

[0085] After successful training, a dedicated vehicle detection network model will be obtained, which can be applied to vehicle detection work in the infrared scene, and its detection process is as Figure 2 shown.

[0086] 2) Experimental verification: After the network training, comparative tests and ablation experiments are carried out on some modules in the network to verify the effectiveness of the network for vehicle detection tasks in the infrared scene. Specifically, it includes the comparison of lightweight backbone networks, the comparison of attention modules, and the self-comparison experiment of calibration factors. The experimental results are shown in the following table:

[0087] Comparison experiment results of each backbone network

[0088]

[0089] Comparison experiment results of the attention mechanism

[0090]

[0091]

[0092] Comparison experiment results of the calibration factor parameters

[0093]

[0094] The ablation experiment results for the improvements of each module are shown in the following table, including the impacts of each improvement on the overall detection accuracy, the number of parameters, and the computational complexity of the network, as well as the impacts on the detection accuracy of vehicle targets at each scale.

[0095] Overall ablation experiment results

[0096]

[0097] Ablation experiment results of multi-scale detection accuracy

[0098]

[0099] In addition, comparative experiments are designed for different classical detection network algorithms, and the experimental results are as follows:

[0100] Comparison experiment results with classical networks

[0101]

[0102] The above experimental results verify the effectiveness of the various optimizations and improvements to the network, meet the requirements of high-precision real-time performance, and the network can be applied to the actual detection system.

[0103] 3) Image acquisition: First, control the infrared camera turntable through the interaction terminal. The real-time video stream under the current field of view can be obtained by means of automatic scanning or manual control. Read the video frame and perform normalization processing, and uniformly process the infrared image into a resolution of 640*512 to ensure that the clarity of the image can meet the detection requirements.

[0104] 4) Network detection: Preprocess the obtained real-time image using median filtering technology to reduce the influence of irrelevant noise. Subsequently, send the frame image into the trained vehicle detection network for prediction to determine whether there are vehicle targets in the image. The overall structure of the vehicle detection network in the present invention is as Figure 3As shown. After the network receives the input image, it first needs to go through each stage of the lightweight network ShuffleNetV2 optimized by adding a dual-branch adaptive fusion attention module, extract feature information with different depth levels, and send it to the fusion network. Among them, the optimized ShuffleBlock module and the attention module are as Figure 4 , 5 shown. The optimized feature fusion network is as Figure 6 shown, where C3, C4, and C5 are the output features from the backbone network. All three feature maps undergo channel transformation through 1×1 convolution, and then generate P3, M4, and M5 through upsampling and inter-layer addition operations. Starting from P3, a bottom-up aggregation path is added and laterally connected to the input of the same-scale layer, finally generating P4 and P5. Based on P5 supplemented with detailed information, downsampling operations are performed to generate P6 and P7. The input feature layer is deeply fused through a bidirectional cross-fusion network structure, and the five different-scale features are output to the detection head network for classification and localization. The detection head network is different from the traditional detection network structure. Based on task interaction features, it classifies and locates samples more uniformly. Its network structure is as Figure 7 shown. First, the feature extractor performs N convolution operations on the output of the feature fusion to obtain a task interaction feature stack, which is used as the common feature basis for the two task branches. The interaction features pass through the two prediction branches to obtain preliminary classification and regression results, and at the same time, the prediction results are further adjusted through the learning in the reverse process of the spatial probability map and the spatial offset map. The overall detection logic based on task alignment is as Figure 8 shown.

[0105] 5) Detection feedback: If the vehicle detection network detects vehicle target information, the target will be displayed on the display device, and at the same time, the vehicle target position information within the current field of view will be fed back to the control platform. After calculation, the corresponding turntable parameters are obtained to further adjust the camera field of view.

Claims

1. An infrared image vehicle detection method based on a single-stage network, comprising the following steps: Step 1, image acquisition: An infrared camera performs a field scan to obtain an infrared image; Step 2, image processing: Preprocess the obtained infrared image; Step 3, network detection: Send the preprocessed infrared image into a trained vehicle detection network for vehicle target detection. Feature extraction is performed through the backbone network, feature fusion is carried out using an optimized fusion network, and the vehicle target category and position information are predicted through the classification and regression branches of the vehicle detection network; The vehicle detection network is a single-stage network structure, based on the RetinaNet network and optimized and improved, including: 1) Select the ShuffleNetV2 lightweight network as the backbone network, use the local information correlation of infrared images to assist target recognition to improve the network's extraction effect on invariant fine-grained information, design a convolution method that combines central difference convolution and convolution operation, embed this convolution method into the core module ShuffleBlock of the backbone network, dynamically adjust the weight ratio of the two convolutions for feature extraction, the convolution kernel slides and scans for sampling on the feature map, extracts the pixel points in the area corresponding to the convolution kernel before aggregation, and calculates the difference between the pixel value of its center point and the values of the other pixel points in turn to obtain the updated pixel value, and then perform a dot product aggregation of the pixel value and the convolution kernel weight to obtain the final output value; 2) A Dual-branch Adaptive Fusion Channel Attention Module (DBAM) is constructed in the vehicle detection network. By means of two parallel branches of global average pooling and global maximum pooling, the global information and local information of the infrared image are utilized, and dynamic one-dimensional convolution is adopted to generate channel weights to better complete the information interaction between channels. The method of dual-branch adaptive fusion is used to dynamically adjust the weight ratio of the two pooling branches in the fusion. The specific implementation process of DBAM is as follows: First, global average pooling and global maximum pooling operations are respectively performed on the input features with height, width and number of channels of H×W×C in the channel dimension, and then the channel weight matrices M avg and M max of the two branches are generated through fast one-dimensional convolution with a convolution kernel of size k. The two channel weight matrices obtain the aggregated channel weight matrix through the adaptive fusion structure, and after element-wise multiplication with the original input features, the weights are mapped into the feature map of H×W×C; 3) Use the ShuffleBlock module with central difference convolution introduced as the basic unit, and at the same time embed a dual-branch channel attention module to construct the backbone network; The feature information with different depth levels extracted from each stage of the lightweight network ShuffleNetV2 is sent to the optimized fusion network. The optimized fusion network is based on the feature pyramid structure, adopts a two-way cross fusion network structure, adds an additional bottom-up aggregation path in some feature layers to improve the fusion effect, and adds a lateral connection between the original input and the final output nodes of the same-scale feature layer; Adopt a fast normalization fusion strategy to add additional weights to different input nodes to distinguish the contribution degrees of different nodes; 4) Design a detection head network based on the task alignment idea, called the task alignment head TAHead: It consists of a feature extractor and two optimized task branches. The feature extractor is used to perform multi-level extraction on the output features of the feature fusion network, and each task branch contains a calibration branch to perform probability adjustment and spatial adjustment on the preliminary prediction results; 5) Design a network training and learning strategy based on the task alignment idea; Set a calibration factor to clarify the alignment degree of the classification and regression tasks; Introduce the calibration factor into the loss calculation, and adjust the prediction results of the two tasks for the same sample point by synchronously optimizing the classification and regression losses to improve the spatial misalignment problem.

2. The infrared image vehicle detection method according to claim 1, wherein In 2), the convolution kernel size k represents the range of cross-channel interaction, and there is a non-linear mapping relationship between it and the channel dimension C. The calculation formula is as follows: where |t| odd denotes the nearest odd number of t, and γ and b are two adjustable parameters.

3. The infrared image vehicle detection method according to any one of claims 1 to 2, characterized in that In 2), the aggregated channel weight matrix is: where δ is the sigmoid activation function, μ and ν are two learnable floating-point parameters that are dynamically learned with the network model and are initially set to 1; and represent element-wise addition and element-wise multiplication of matrices, respectively.

4. The infrared image vehicle detection method according to any one of claims 1 to 3, characterized in that, 4) It includes: Performing N convolutional operations on the output of feature fusion using a feature extractor to obtain a multi-layer task interaction feature stack, which serves as the common feature basis for the two task branches; the two improved task branches are the classification branch and the regression branch respectively. Among them, the classification task branch first performs concat splicing and convolutional operations on the task interaction feature stack, and obtains a dense classification score of H×W×1 through the sigmoid activation function as the preliminary classification prediction result; the regression task branch also obtains a regression bounding box score of H×W×4 through concat splicing and convolutional operations as the preliminary regression prediction result.

5. The infrared image vehicle detection method according to any one of claims 1 to 4, characterized in that, 4) It also includes: constructing a parallel calibration branch for each of the two tasks in TAHead to clearly adjust the preliminary prediction results of the two tasks; simultaneously adjusting the predictions of the two task branches through the obtained task interaction features and subsequent task alignment learning strategies; the two calibration branches use the task interaction features to generate a spatial probability map and a spatial offset map. The spatial probability map learns the prediction consistency between the two tasks at each spatial position through the backpropagation process, and then adjusts the preliminary prediction result of the classification task to obtain the final classification result; the spatial offset map learns the spatial offset between the current anchor box and the surrounding best anchor box through the backpropagation process, and then adjusts the preliminary prediction result of the regression task to obtain the final regression result.

6. The infrared image vehicle detection method according to claim 1, characterized in that 5) In, the calibration factor is obtained through the classification and regression tasks, and its calculation process is as follows: t = s α × u β In the formula, t is the calibration factor, s and u respectively represent the classification score and IOU value obtained by the classification and regression tasks for each anchor box, and α and β are set to control the influence of the two tasks on the calibration factor; select the m anchor boxes with the largest t value as positive samples, and the rest as negative samples.

7. The infrared image vehicle detection method according to claim 6, characterized in that, 5) In, the Focal Loss is used as the classification loss function to alleviate the imbalance between positive and negative samples. On this basis, the calibration factor is used to replace the original positive sample anchor box label value to improve the classification score of the anchor boxes with a higher alignment degree.