A multi-spectral pedestrian detection method and system based on intelligent vehicles

The method improves pedestrian detection in smart vehicles by preprocessing and feature enhancement techniques, addressing inadequate fusion mechanisms and underutilized semantic information to enhance detection accuracy and reduce false negatives.

CN115457456BActive Publication Date: 2025-07-15WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211008039.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-22
Publication Date
2025-07-15
Estimated Expiration
2042-08-22

AI Technical Summary

Technical Problem

The existing multispectral pedestrian detection method is relatively rough in the design of the fusion mechanism, resulting in insufficient pedestrian position information, weak background distinction, and insufficient use of the high-level semantic information in deep fusion features, limiting the expressiveness of multi-scale features.

Method used

The multi-spectral pedestrian detection method based on intelligent vehicles is adopted. By pre-processing visible light and infrared images, the edge features of the infrared images are acquired, and feature extraction is performed using the improved ResNet50 network, and feature fusion and advanced semantic information extraction are performed. Combined with feature aggregation and pedestrian position information enhancement processing, the detection effect is improved.

Benefits of technology

The location information and semantic characteristics of pedestrian detection are enhanced, the detection accuracy is improved, the missed detection rate is reduced, and the efficiency of all-weather pedestrian detection is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457456B_ABST
    Figure CN115457456B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-spectral pedestrian detection method and system based on intelligent vehicles, including: 1) preprocessing paired visible light and infrared images; 2) calculating gradients of the preprocessed infrared images to obtain edge feature maps and obtain infrared images with supplementary edge features; 3) respectively inputting the paired visible light and infrared images with supplementary edge features into a ResNet50 network for feature extraction; 4) adding and fusing the output features of the same stage of the two modalities to obtain fused features; 5) performing high-level semantic feature extraction on the fused features of the highest layer to obtain high-level semantic information; 6) performing pedestrian position information enhancement processing on the fused feature maps with supplementary high-level semantic information; 7) performing pedestrian detection on the fused feature maps of different scales after feature supplementation and enhancement to obtain pedestrian detection results. The present invention supplements and enhances the information of the fused features, achieving better detection effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing technology, and in particular to a multispectral pedestrian detection method and system based on intelligent vehicles. Background Art

[0002] Traditional pedestrian detection systems usually use visible light images for pedestrian detection. In good lighting conditions, visible light images can better reflect information such as the color and texture of pedestrian targets and achieve good detection effects. However, in cases where the target is occluded, lighting conditions are poor, and the contrast is weak, the performance of pedestrian detection systems based on visible light images is poor. Compared with visible light sensors, near-infrared (0.75 - 1.3 μm) cameras or long-wave infrared (7.5 - 13 μm, i.e., thermal) cameras provide supplementary information outside the visual spectrum. Compared with near-infrared (NIR) cameras, long-wave thermal infrared (LWIR) cameras are less affected by environmental illumination through the data obtained by detecting the thermal radiation emitted by the human body itself, and can collect images with significant pedestrian contour information in the night environment with insufficient lighting. However, due to the imaging principle, long-wave infrared images have the disadvantages of low resolution and lack of information such as color and texture, and the detection effect is poor when the thermal radiation difference between the target and the background is small. Therefore, long-wave thermal infrared cameras cannot completely replace visible light images for pedestrian detection. Visible light and long-wave infrared cameras have their own advantages. Combining the two in the all-day pedestrian detection task can give play to the complementarity of bimodal data and theoretically achieve better detection effects, realizing the detection of pedestrian targets under different lighting conditions all day long.

[0003] However, the existing multispectral pedestrian detection methods are generally rough in the design of the fusion mechanism, resulting in insufficient position information of pedestrians in the fused features and weak distinguishability between pedestrians and the background. In addition, in different fusion architectures, the high-level semantic information contained in the deep fused features has not been fully utilized, restricting the enhancement of the multi-scale feature expressiveness.

[0004] Therefore, it is very necessary to provide a solution for supplementing and enhancing the fused features to make full use of the information beneficial to pedestrian detection and improve the pedestrian detection ability of intelligent vehicles. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a multispectral pedestrian detection method and system based on intelligent vehicles aiming at the defects in the prior art.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows: A multispectral pedestrian detection method based on intelligent vehicles includes the following steps:

[0007] 1) Preprocess the paired visible light and infrared images;

[0008] 2) Calculate the gradient of the preprocessed infrared image to obtain an edge feature map and an infrared image with supplementary edge features;

[0009] 3) Input the paired visible light and the infrared image with supplementary edge features into the ResNet50 network respectively for feature extraction, and output the features at different stages in the feature extraction network;

[0010] 4) Add and fuse the output features at the same stage of the two modalities to obtain a fused feature with the same structure as ResNet50;

[0011] 5) Perform high-level semantic feature extraction on the fused feature of the highest layer to obtain high-level semantic information;

[0012] One stage contains several layers. Here, the highest layer refers to the last layer of the last stage;

[0013] 6) Transmit the high-level semantic information in step 5) to the shallow fused features of different scales in step 4), and fuse them through a feature aggregation module to obtain a fused feature with supplementary high-level semantic information;

[0014] 7) Perform pedestrian position information enhancement processing on the fused feature to enhance the pedestrian position information in the fused feature;

[0015] 8) Perform pedestrian detection on the fused feature maps of different scales after feature supplementation and enhancement to obtain pedestrian detection results.

[0016] According to the above scheme, in step 2), calculate the gradient of the infrared image and output an edge feature map, specifically as follows:

[0017] 2.1) Arbitrarily select one of the three channels of the infrared image after data preprocessing as the original image P for gradient calculation;

[0018] 2.2) Use the Sobel operator to calculate the horizontal and vertical gradients of the original image P;

[0019] Let P represent the original image, G x and G y respectively represent the horizontal and vertical edge features obtained by horizontal and vertical gradient calculations. The formulas are as follows:

[0020]

[0021] Adopt the following formula to combine the horizontal and vertical edge features to obtain the edge feature map G:

[0022]

[0023] 2.3) After obtaining the edge feature map G, splice the edge feature map with the preprocessed 3-channel infrared image to obtain a 4-channel infrared image with supplementary edge features.

[0024] According to the above solution, in step 3), the ResNet50 network used is an improved ResNet50 network;

[0025] In the {conv1, conv2_x, conv3_x, conv4_x, conv5_x} structure included in the original ResNet50, replace the standard convolutions included in the conv4_x and conv5_x stages with deformable convolutions; the deformable convolution adds an offset to the original standard convolution calculation position, so that the shape of the convolution is no longer square, and the attention of the convolution will change with the target of interest;

[0026] The calculation method of the deformable convolution layer is as follows:

[0027]

[0028]

[0029] Among them, is the sampling point on the convolutional regular grid, w is the weight of the convolution, x is the input feature map, p0 is the center of the grid, p n Traverse the positions of each sampling point in, Δp n is the offset learned for each sampling point, so the offset sampling position is p n + Δp n By adding the weight coefficient Δm n to distinguish whether the introduced area is the area of the target of interest. If the convolution is not interested in the area of this sampling point, the weight is set to 0.

[0030] According to the above solution, in step 5), obtain the high-level semantic information, specifically as follows:

[0031] 5.1) Perform pyramid pooling on the highest-level fused feature conv5_x_f to extract the high-level semantic feature f1;

[0032] 5.1.1) Input the fused feature conv5_x_f into the pooling layer sub-branches of 4 scales in sequence. The pooling layer branches divide the feature map into different sub-regions and form different pooled feature representations. The spatial sizes of the feature maps output by the 4 pooling layer sub-branches are 1×1, 2×2, 4×4, and 6×6 respectively;

[0033] 5.1.2) When the number of input feature channels is N, use a 1×1 convolutional layer after each pooled feature to reduce the number of channels of the feature to N / 4;

[0034] 5.1.3) Upsample the feature after dimensionality reduction through bilinear interpolation to obtain a feature with the same size as the fused feature conv5_x_f, and concatenate the features of different branches to obtain the final output feature f1 after pyramid pooling. The scale and number of channels of f1 are consistent with the input feature;

[0035] 5.2) Perform cross-attention processing on the highest-level fused feature conv5_x_f to extract the high-level semantic feature f2;

[0036] 5.2.1) For the input feature map Apply two 1×1 convolutional layers on conv5_x_f to generate two feature maps Q and K respectively, where C′ is the number of channels, and after dimensionality reduction calculation, the number of channels C′ is less than C. Among them, represents the input feature map, W represents the width of the input feature map, H represents the height of the input feature map, C and C' are the number of channels, and after dimensionality reduction calculation, the number of channels C' is less than the number of channels C;

[0037] 5.2.2) At each position u in the spatial dimension of Q, obtain a vector Extract the feature vectors at the positions in the same row and the same column as position u in K to obtain the set Calculate the correlation degree of the two elements at the cross positions;

[0038] is the i-th element of Ω u The formula for calculating the correlation degree is defined as follows:

[0039]

[0040] Among them, d i,u ∈ D is the correlation degree between the feature Q u and Ω i,u The value range of i is [1,..., H + W - 1],

[0041] 5.2.3) Apply the softmax layer to the correlation degree set D to calculate the attention map A;

[0042] 5.2.4) Apply another 1×1 convolutional layer on conv5_x_f to generate used to adjust the feature; at each position u in the spatial dimension of V, obtain a vector and a set of The set Φ uis the set of feature vectors in V that are in the same row or column as the position u. The context information of the highest-level fused features can be collected through the following formula:

[0043]

[0044] where, [conv5_x_f]′ u is the feature vector at the position u, and A i,u is the attention scalar value at the channel i and the position u in A; adding the context information [conv5_x_f]′ u to the feature [conv5_x_f] u can enhance the pixel-level representation of the feature; the high-level semantic feature finally output by the cross-attention module is [conv5_x_f]′, abbreviated as f2.

[0045] 5.3) Add and fuse the information obtained by the two semantic information extraction sub-modules to obtain the fused high-level semantic feature f.

[0046] According to the above scheme, in step 6), the high-level semantic information is passed to the shallow fusion features of different scales and fused through a feature aggregation module, where the feature aggregation module mainly includes feature upsampling, feature fusion, and depth pooling processing operations. The specific steps are as follows:

[0047] S61) Upsample the high-level semantic feature to obtain the same scale as the shallow fusion feature;

[0048] S62) Upsample the fusion feature of the previous stage by 2 times to obtain the same scale as the adjacent shallow fusion feature;

[0049] S63) Fuse the upsampled high-level semantic feature, the fusion feature of the previous stage upsampled by 2 times, and the shallow fusion feature of this stage by element-wise addition;

[0050] S64) Perform depth pooling processing on the fusion feature obtained in S63 to reduce the interference caused by the upsampling of the high-level semantic feature and the fusion feature of the previous stage; for the depth pooling processing, first input the fused feature map into the average pooling layer at different downsampling rates to convert it into different ratio spaces, and then input it into a 3×3 convolutional layer for feature learning. Then, upsample the convolutional feature, and then merge the upsampled feature maps from different branches together. Finally, process the fusion feature through a 3×3 convolutional layer.

[0051] The depth pooling processing allows each spatial position to view the context information of the feature map in different scale spaces, further refining the feature after fusing the high-level semantic information.

[0052] According to the above solution, in step 7), pedestrian position information enhancement processing is performed as follows: using the existing pedestrian annotations, a binary image of 0 and 1 is generated as the ground truth of the pedestrian target in the pedestrian position information enhancement processing. The probability value of each position in the fused feature map being in the pedestrian area is predicted through supervised learning. The predicted probability map is multiplied element-wise with the fused feature after supplementing high-level semantic information, and the product result is used as the residual to enhance the position of the pedestrian in the fused feature. The process is expressed as:

[0053]

[0054]

[0055] where \(F\) is a set of fused feature maps of different scales for pedestrian detection, \(F_i\) i refers to the \(i\)-th feature map; represents the convolution operation, which is used to predict the possibility of each position in the feature map being in the pedestrian area; \(\delta\) represents the sigmoid activation function, which is used to convert the predicted possibility into a probability value within 0 - 1; \(w_i\) i is the probability feature map, and its size is the same as that of \(F_i\) i ; is the feature map with enhanced pedestrian area.

[0056] According to the above solution, when predicting the probability value of each position in the fused feature map being in the pedestrian area in step 7) through supervised learning, the loss function uses two loss functions, BCE and IoU, to improve the prediction probability of a certain position area where the pedestrian is located;

[0057]

[0058]

[0059]

[0060] where \(l_i\) (i) is the loss of the \(i\)-th feature map. In the formula, \(G(r, c) \in \{0, 1\}\) is the ground truth label of the pixel \((r, c)\), and \(S(r, c)\) is the predicted probability that the pixel is in the area where the pedestrian is located. Using the BCE loss can maintain a smooth gradient for all pixels during training, and at the same time, using the IoU loss can focus more on the pedestrian foreground.

[0061] A multi - spectral pedestrian detection system based on intelligent vehicles, which is used to deploy a multi - spectral pedestrian detection method based on intelligent vehicles. The system includes a data acquisition module, a calculation and storage module, and a display output module. Among them, the data acquisition module mainly includes a visible - light camera and an infrared camera. The calculation and storage module is implemented using a development board. By deploying the trained and optimized multi - spectral pedestrian detection model in the memory of the development board and running the model on the development board, the detection results can be obtained. The display output module displays the prediction boxes of pedestrians in the corresponding frames acquired by the camera through a display. The system can not only be installed on vehicles for real - time detection based on the driving data of the vehicles, but also be used in other pedestrian detection fields, such as video surveillance fields, robots, etc. The specific steps of the system are as follows:

[0062] Step 1: Train and optimize a multi - spectral pedestrian detection model based on intelligent vehicles using a multi - spectral pedestrian dataset.

[0063] Step 2: Deploy and store the trained multi - spectral pedestrian detection model in the memory of the development board.

[0064] Step 3: Acquire picture input data based on the visible - light and infrared cameras in the data acquisition module, and transmit the input data to the processor of the development board.

[0065] Step 4: The processor pre - processes the input original picture data, and at the same time loads the detection inference code in the memory into the inference calculation module.

[0066] Step 5: Input the data into the inference calculation module and run the detection model deployed in Step 2.

[0067] Step 6: The inference calculation module outputs the detection results, and the display module outputs the detection boxes of pedestrians in the input picture.

[0068] The beneficial effects of the present invention are as follows:

[0069] The present invention can obtain the edge features of pedestrian targets in infrared images by calculating the gradients of infrared images, which supplements the fused information. Through the extraction of high - level semantic information, the spatial and semantic information of the image fusion features is better retained. Through feature aggregation, the high - level semantic information in the high - level fusion features can be efficiently fused with the shallow features with detailed texture information, enriching the information of feature maps at different scales. Through the enhancement operation of pedestrian position information, the position area of pedestrians in the feature map becomes more prominent, which is conducive to the positioning of detection boxes.

[0070] Based on the operations of the designed multi-spectral pedestrian detection task, the detection model has achieved high detection accuracy in the KAIST public detection dataset, and the missed detection rate of pedestrian targets has been significantly reduced. The multi-spectral pedestrian detection system based on intelligent vehicles provided by the present invention is not limited to the field of intelligent vehicle environment perception, but can also be extended to other applications related to pedestrian detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0072] Figure 1 is the flowchart of the method according to the embodiment of the present invention;

[0073] Figure 2 is the network framework diagram according to the embodiment of the present invention;

[0074] Figure 3 is the schematic diagram of the principle of pyramid pooling for high-level semantic information extraction according to the embodiment of the present invention;

[0075] Figure 4 is the schematic diagram of the principle of cross-attention for high-level semantic information extraction according to the embodiment of the present invention;

[0076] Figure 5 is the schematic diagram of the principle of depth pooling operation according to the embodiment of the present invention;

[0077] Figure 6 is the schematic diagram of the principle of the feature aggregation module according to the embodiment of the present invention;

[0078] Figure 7 is the schematic diagram of the principle of enhancing pedestrian position information according to the embodiment of the present invention;

[0079] Figure 8 is the structural schematic diagram of a multi-spectral pedestrian detection system provided by the present invention;

[0080] Figure 9 is the layout schematic diagram of a multi-spectral pedestrian detection system provided by the present invention; wherein: 1 - data acquisition module, 2 - display output module, 3 - calculation and storage module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0082] First, the technical terms of the present invention will be explained:

[0083] ResNet50: ResNet is a deep residual network; ResNet50 represents a ResNet network with 50 convolutional layers.

[0084] Sobel operator: The Sobel operator is an edge operator that detects edges by weighted summing the gray values of the four neighboring pixels above, below, left, and right of each pixel in an image.

[0085] Deformable convolution: Since the regular grid sampling in standard convolution makes it difficult for the network to adapt to the geometric deformation of the target, deformable convolution learns offsets based on a parallel network, causing the convolution kernel to shift at the sampling points of the input feature map, focusing on the regions or targets of interest and improving the robustness of the feature map after convolution calculation.

[0086] As Figure 1 shown, a multi-spectral pedestrian detection method based on intelligent vehicles includes the following steps:

[0087] S1) Preprocess the paired visible light and infrared images;

[0088] In step S1), the paired visible light and infrared images are subjected to data augmentation and preprocessing. The operations used for data augmentation are color jittering, scale scaling, and random horizontal flipping; data preprocessing mainly normalizes the input 3-channel image at the channel level according to the mean of the pixels in each channel and performs normalization.

[0089] S2) Calculate the gradient of the preprocessed infrared image, mine the significant gradient information in the infrared image to obtain an edge feature map, and obtain an infrared image with supplementary edge features after splicing with the infrared image;

[0090] As Figure 2 shown by the upper left gradient calculation module, the specific steps are as follows:

[0091] S21: Arbitrarily select one of the three channels of the 3-channel thermal image after data preprocessing as the original image for gradient calculation;

[0092] S22: Use the Sobel operator to calculate the gradient and obtain the edge features of the thermal image;

[0093] Since the Sobel operator contains two 3×3 weight matrices in the horizontal and vertical directions, the Sobel operator can be assigned to two 3×3 convolutions, and the operator is convolved with the image in the plane to respectively obtain the approximate values of the horizontal and vertical brightness differences of the image and obtain the edge features. Let P represent the original image, G x and G y respectively represent the horizontal and vertical edge features obtained by horizontal and vertical gradient calculations, and the formula is as follows:

[0094]

[0095] The horizontal and vertical edge features are combined using the following formula to obtain the edge feature map G:

[0096]

[0097] The present invention uses the sum of absolute values to approximate the operation of the square root of the sum of squares and combines the gradient information of each pixel to calculate the gradient for each pixel of the image to obtain the edge feature, where the number of channels of the edge feature is 1:

[0098] G = |G x | + |G y |

[0099] S23: After obtaining the edge feature map G, the edge feature map is stitched with the pre-processed 3-channel thermal image to obtain a 4-channel thermal image that supplements the gradient edge information;

[0100] S24: The 4-channel thermal image that supplements the gradient edge information passes through a 3×3 convolution to change the number of channels to 3 to adapt to the input of the feature extraction network. The 3×3 convolution is used to change the number of channels, and its initialization uses kaiming initialization.

[0101] 3) The paired visible light and the infrared image supplemented with edge features are respectively input into the ResNet50 network for feature extraction, and the features at different stages in the feature extraction network are output;

[0102] The improved ResNet50 network adopted in this embodiment is specifically: in the {conv1, conv2_x, conv3_x, conv4_x, conv5_x} structure included in the original ResNet50, the standard convolutions included in the conv4_x and conv5_x stages are replaced with deformable convolutions. Since the standard convolution unit samples the input feature map at fixed positions, in the same layer of convolution calculation, the receptive fields of all activation units are the same. However, since different positions may correspond to objects of different scales or deformations, the standard convolution has the disadvantage of being difficult to adapt to the geometric deformations of objects, which limits the neural network from extracting high-quality feature maps.

[0103] Among them, the calculation method of the standard convolution is:

[0104]

[0105] Compared with standard convolution, deformable convolution adds an offset to the original convolution calculation position, making the shape of the convolution no longer square, and the attention of the convolution will change with the target of interest. This property enables the features after convolution to capture pedestrian information with shape changes and improve the quality of the features. The calculation method of deformable convolution is as follows:

[0106]

[0107]

[0108] where are the sampling points on the convolutional regular grid, w is the weight of the convolution, x is the input feature map, p0 is the center of the grid, and p n traverses the positions of each sampling point in, and Δp n is the offset of each sampling point learned after network training, and p n +Δp n is the sampling position after the offset of the deformable convolution. By adding the weight coefficient Δm n to distinguish whether the introduced area is the area of the target of interest. If the area of this sampling point is not the area of interest, the weight is 0.

[0109] In the design of the convolutional neural network, it is not the case that the more the number of deformable convolutions, the better. When the number of deformable convolutions reaches a certain value, the detection accuracy of the model will reach saturation or even decrease. In the present invention, the standard convolution is replaced with deformable convolution in the conv4_x and conv5_x stages of ResNet50.

[0110] S4) Add and fuse the output features of the same stage of the two modalities to obtain a fused feature with the same structure as ResNet50;

[0111] Input the paired visible light and infrared images with supplementary edge features into the improved ResNet50 network for feature extraction. The output features of different stages are {conv1_v, conv2_x_v, conv3_x_v, conv4_x_v, conv5_x_v} and {conv1_t, conv2_x_t, conv3_x_t, conv4_x_t, conv5_x_t} respectively.

[0112] The fused feature obtained by adding and fusing the output features of the same stage of the two modalities in step S4) is represented as {conv1_f, conv2_x_f, conv3_x_f, conv4_x_f, conv5_x_f}.

[0113] 5) Perform high-level semantic feature extraction on the fused features of the highest layer to obtain high-level semantic information;

[0114] One of the following strategies can be adopted for high-level semantic feature extraction:

[0115] Perform pyramid pooling on the fused features conv5_x_f of the highest layer to extract high-level semantic features;

[0116] Perform cross-attention processing on the fused features conv5_x_f of the highest layer to extract high-level semantic features;

[0117] Pyramid pooling focuses on semantic information within adjacent regions, while cross-attention focuses on semantic information between regions that are farther apart. The information obtained from the above two semantic feature extractions can also be added and fused to obtain high-level semantic features; specifically as follows:

[0118] 5.1) Perform pyramid pooling on the fused features conv5_x_f of the highest layer to extract high-level semantic feature f1;

[0119] 5.1.1) As Figure 3 shown, sequentially input the fused feature conv5_x_f into 4 pooling layer sub-branches of different scales. The pooling layer branches divide the feature map into different sub-regions and form different pooled feature representations. The spatial sizes of the feature maps output by the 4 pooling layer sub-branches are 1×1, 2×2, 4×4, and 6×6 respectively;

[0120] 5.1.2) When the input feature channels are N, use a 1×1 convolutional layer after each pooled feature to reduce the number of channels of the feature to N / 4;

[0121] Upsample the dimension-reduced features through bilinear interpolation to obtain features of the same size as the fused feature conv5_x_f, and splice the features of different branches to obtain the final output feature f1 after pyramid pooling. The scale and number of channels of f1 are the same as those of the input features;

[0122] 5.2) Perform cross-attention processing on the fused features conv5_x_f of the highest layer to extract high-level semantic feature f2;

[0123] 5.2.1) As Figure 4 shown, input the feature map Apply two 1×1 convolutional layers on conv5_x_f to generate two feature maps Q and K respectively, where C′ is the number of channels, and the number of channels C′ is less than C after dimension reduction calculation;

[0124] 5.2.2) At each position u in the spatial dimension of Q, obtain a vector Extract the feature vectors at the positions in the same row and column as the position u in K to obtain a set Calculate the correlation degree of the two elements at the cross position;

[0125] is the i-th element of Ω u The formula for calculating the correlation degree is as follows:

[0126]

[0127] where, d i,u ∈ D is the correlation degree between the feature Q u and Ω i,u i = [1,..., H + W - 1],

[0128] 5.2.3) Apply the softmax layer to the set D of correlation degrees to calculate the attention map A;

[0129] 5.2.4) Apply another 1×1 convolutional layer on conv5_x_f to produce used to adjust the features; obtain a vector at each position u in the spatial dimension of V and a set of sets The set Φ u is the set of feature vectors in V that are in the same row or column as the position u. The context information of the highest-level fusion features can be collected through the following formula:

[0130]

[0131] where, [conv5_x_f]′ u is the feature vector at the position u, and A i,u is the attention scalar value at the channel i and the position u in A; adding the context information [conv5_x_f]′ u to the feature [conv5_x_f] u can enhance the pixel-level representation of the feature; the final output high-level semantic feature of the cross-attention module is [conv5_x_f]′, abbreviated as f2.

[0132] 5.3) Add and fuse the information obtained by the two semantic information extraction sub-modules to obtain the fused high-level semantic feature f.

[0133] Step 6) Pass the high-level semantic feature obtained in Step 5) to the shallow fusion features of different scales and fuse them through the feature aggregation module, as Figure 6 shown. The feature aggregation module mainly includes feature upsampling, feature fusion, and depth pooling processing operations. The specific steps are as follows:

[0134] S61) Upsample the high - level semantic features to obtain the same scale as the shallow - layer fusion features;

[0135] S62) Upsample the fusion features of the previous stage by a factor of 2 to obtain the same scale as the adjacent shallow - layer fusion features;

[0136] S63) Fuse the upsampled high - level semantic features, the fusion features of the previous stage after 2 - fold upsampling, and the shallow - layer fusion features of this stage by element - wise addition;

[0137] S64) Perform depth pooling on the fusion features obtained in S63 to reduce the interference brought by the upsampling of high - level semantic features and the fusion features of the previous stage; The depth pooling process is as Figure 5 shown. First, input the fused feature map into the average pooling layer at different downsampling rates to convert it into different proportional spaces. The features in different proportional spaces are respectively input into a 3×3 convolutional layer for feature learning. Then, upsample the convolutional features and concatenate them in the channel dimension. Finally, process the fused features through a 3×3 convolutional layer.

[0138] The depth pooling process allows each spatial position to view the context information of the feature map in different scale spaces, further refining the features after fusing high - level semantic information.

[0139] S7) Perform pedestrian position information enhancement on the fused feature map obtained by fusing the high - level semantic information in step S4) in step 5) to enhance the pedestrian position information in the fused feature map, as Figure 7 shown;

[0140] Specifically, use the existing pedestrian annotations to generate a binary image of 0 and 1 as the ground truth of the pedestrian target in the pedestrian position information enhancement module. Through supervised learning, predict the probability that each position in the fused feature map is in the pedestrian area. Multiply the predicted probability map element - wise with the fused feature output in 6) and use the product result as the residual to enhance the positions of pedestrians in the fused feature; The process is expressed as:

[0141]

[0142]

[0143] where F is a set of fused feature maps of different scales for pedestrian detection, and F i refers to the i - th feature map that fuses high - level semantic information; represents the convolution operation, which is used to predict the possibility that each position in the feature map is in the pedestrian area; δ represents the sigmoid activation function, which is used to convert the predicted possibility into a probability value within 0 - 1; wi is a probability feature map, whose size is the same as that of F i consistent; is the fused feature map after enhancing the pedestrian position information.

[0144] When predicting the probability value of each position in the fused feature map being in the pedestrian area through supervised learning, the loss function uses two loss functions, BCE and IoU, to improve the prediction probability of a certain position area where the pedestrian is located;

[0145]

[0146]

[0147]

[0148] where l (i) is the loss of the i-th feature map. In the formula, G(r, c) ∈ {0, 1} is the ground truth label of the pixel (r, c), and S(r, c) is the predicted probability that the pixel is in the area where the pedestrian is located. Using the BCE loss can maintain a smooth gradient for all pixels during training, and at the same time, using the IoU loss can focus more on the pedestrian foreground.

[0149] 8) Perform pedestrian detection on the fused feature maps of different scales after feature supplementation and enhancement to obtain the pedestrian detection results.

[0150] After the above steps, the feature map has been supplemented with gradient information (edge features), supplemented with high-level semantic information, and enhanced with pedestrian position information. By inputting the feature maps of different scales after feature supplementation and enhancement into the region proposal network to generate candidate boxes, and then inputting the candidate boxes into the detection module, the predicted box and confidence of the pedestrian can be obtained.

[0151] To verify the effectiveness of the present invention, the present invention is trained on the KAIST dataset and compared with other algorithms under the KAIST test set. The results are shown in Table 1. The commonly used evaluation index for the multi-spectral pedestrian detection method is the Log-average Miss Rate (LAMR), and the lower its value, the better the detection effect.

[0152] The experimental results show that the method of the present invention obtains the best detection results in all-weather and night scenes, and the detection speed is the fastest. The faster the detection speed of the model, the more beneficial it is for intelligent vehicles to perform rapid detection in real application scenarios.

[0153] Table 1 Comparison results of the present invention and other methods on the KAIST multi-spectral pedestrian test set

[0154]

[0155] According to the above method, correspondingly, we can obtain a multi-spectral pedestrian detection device based on an intelligent vehicle, including:

[0156] An image preprocessing module, configured to preprocess paired visible light and infrared images;

[0157] A gradient calculation module, configured to calculate the gradient of the preprocessed infrared image, mine significant gradient information in the infrared image, obtain an edge feature map, and obtain an infrared image with supplementary edge features;

[0158] A feature extraction module, configured to respectively input paired visible light and infrared images with supplementary edge features into a ResNet50 network for feature extraction, and output features at different stages in the feature extraction network;

[0159] A feature fusion module, configured to add and fuse the output features of the same stage of the two modalities to obtain a fusion feature with the same structure as ResNet50; among them, the highest-layer feature of each stage is selected as the output feature;

[0160] A high-level semantic information extraction module, configured to perform high-level semantic feature extraction on the highest-layer fusion feature of the last stage in the feature fusion module to obtain high-level semantic information;

[0161] A pedestrian position information enhancement module, configured to perform pedestrian position information enhancement processing on the fusion feature map supplemented with high-level semantic information to enhance the pedestrian position information in the fusion feature map;

[0162] A pedestrian detection module, configured to perform pedestrian detection on the fusion feature maps of different scales after feature supplementation and enhancement to obtain pedestrian detection results.

[0163] Based on the devices used, we also provide a multi-spectral pedestrian detection system based on an intelligent vehicle. The multi-spectral pedestrian detection system based on an intelligent vehicle according to the present invention can be used to deploy a multi-spectral pedestrian detection model based on an intelligent vehicle. The system is integrated in hardware and can work on an intelligent vehicle, and performs pedestrian detection based on real-time collected data during vehicle driving. In addition, the system can also be used in other fields for implementing pedestrian detection tasks.

[0164] As Figure 8 shown, the system includes a data acquisition module, a calculation and storage module, and a display output module. Among them, the data acquisition module mainly includes a visible light camera and an infrared camera; the calculation and storage module is implemented using a development board, and deploys the trained and optimized multi-spectral pedestrian detection model in the memory of the development board; the display output module displays the prediction boxes of pedestrians in the corresponding frame acquired by the camera through a display.

[0165] The specific steps are as follows:

[0166] Step 1: Train and optimize a multi-spectral pedestrian detection model based on intelligent vehicles using a multi-spectral pedestrian dataset.

[0167] Step 2: Deploy and store the trained detection model in a development board.

[0168] Step 3: Obtain picture input data based on visible light and infrared cameras in the data acquisition module, and transmit the input data to the processor of the development board.

[0169] Step 4: The processor preprocesses the input original picture data, and at the same time loads the model detection inference code in the memory into the inference calculation module.

[0170] Step 5: After inputting the data into the inference calculation module, run the calculation program related to the detection method, and the running steps are as described in the foregoing steps S1 - S8.

[0171] Step 6: The inference calculation module outputs the detection result, and outputs the detection frame of the pedestrians in the picture through the display module.

[0172] Before the system is officially put into work, the model needs to be trained, tested and deployed. The model training process is usually completed on a device with high-performance data processing capabilities (such as a computer). The computer reads the existing multi-spectral pedestrian detection dataset and uses the above-mentioned multi-spectral pedestrian detection method based on intelligent vehicles to complete the training and testing of the model. Then, the debugged and optimized model parameters are saved to deploy the model in hardware devices such as the development board of the detection system for use. When the intelligent vehicle multi-spectral pedestrian detection system works, it first loads the detection model deployed in the memory, and loads the picture data collected by visible light and infrared cameras. Each frame of data read is preprocessed and then input into the detection model, and the detection result is output through the inference module.

[0173] As Figure 9 shown, when the system is applied to an intelligent vehicle, the data acquisition module can be arranged on the top of the vehicle to obtain picture data of the environment, the calculation and storage module can be arranged under the dashboard of the driver's seat, and the display module can be arranged at the vehicle console or use the console display screen to display the detected pedestrian targets.

[0174] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A multi-spectral pedestrian detection method based on intelligent vehicles, characterized in that, It includes the following steps: 1) Preprocess the paired visible light and infrared images; 2) Calculate the gradient of the preprocessed infrared image to obtain an edge feature map, and obtain an infrared image with supplementary edge features; 3) Input the paired visible light and the infrared image with supplementary edge features into the ResNet50 network respectively for feature extraction, and output the features at different stages in the feature extraction network; 4) Add and fuse the output features at the same stage of the two modalities to obtain a fused feature with the same structure as ResNet50; among them, the highest-layer feature of each stage is selected as the output feature; 5) Extract high-level semantic features from the highest-layer fused feature of the last stage to obtain high-level semantic information; Obtain high-level semantic information, specifically as follows: 5.1) Perform pyramid pooling on the highest-layer fused feature conv5_x_f to extract high-level semantic feature f1; 5.2) Perform cross-attention processing on the highest-layer fused feature conv5_x_f to extract high-level semantic feature f2; 5.3) Add and fuse the information obtained by the two semantic information extraction sub-modules to obtain the fused high-level semantic feature f; 6) Transmit the high-level semantic information in step 5) to the shallow fused features at different scales in step 4), and fuse them through a feature aggregation module to obtain a fused feature with supplementary high-level semantic information; 7) Perform pedestrian position information enhancement processing on the fused feature to enhance the pedestrian position information in the fused feature map; The pedestrian position information enhancement processing is specifically as follows: use the existing pedestrian annotations to generate a binary image of 0 and 1 as the ground truth of the pedestrian position; predict the probability value of each position in the fused feature map being in the pedestrian area through supervised learning; multiply the predicted probability map element-wise with the fused features at different scales; use the product result as a residual to enhance the pedestrian position in the fused feature; this process is expressed as: Among them, F is a set of fused feature maps of different scales for pedestrian detection, and F i refers to the i-th fused feature map; represents a convolution operation, which is used to predict the possibility that each position in the feature map is in the pedestrian area; δ represents the sigmoid activation function, which is used to convert the predicted possibility into a probability value within 0 to 1; w i is the probability feature map, and its size is the same as that of F i ; is the feature map with enhanced pedestrian area; 8) Input the fused feature maps at different scales with feature supplementation and enhancement into the detection module for pedestrian detection to obtain pedestrian detection results.

2. The multi-spectral pedestrian detection method based on an intelligent vehicle according to claim 1, wherein In step 2), when calculating the gradient of the infrared image and outputting the edge feature map, it is specifically as follows: 2.1) Arbitrarily select one of the 3 channels of the infrared image after data preprocessing as the original image P for gradient calculation; 2.2) Use the Sobel operator to calculate the horizontal and vertical gradients of the original image P; Let P represent the original image, G x and G y represent the horizontal and vertical edge features obtained by horizontal and vertical gradient calculations respectively. The formulas are as follows: The following formula is used to combine the horizontal and vertical edge features to obtain the edge feature map G: 2.3) After obtaining the edge feature map G, splice the edge feature map with the 3-channel infrared image after preprocessing to obtain a 4-channel infrared image with supplementary edge features; 2.4) The 4-channel infrared image passes through a 3×3 convolution to change the number of channels to 3 to adapt to the input of the feature extraction network.

3. The multispectral pedestrian detection method based on an intelligent vehicle according to claim 1, wherein In step 3), the ResNet50 network used is an improved ResNet50 network; In the {conv1, conv2_x, conv3_x, conv4_x, conv5_x} structure included in the original ResNet50, replace the standard convolutions included in the conv4_x and conv5_x stages with deformable convolutions; the deformable convolution adds an offset to the calculation position of the original standard convolution, so that the shape of the convolution is no longer square, and the attention of the convolution will change with the target of interest. The calculation method of the deformable convolution layer is as follows: Among them, is the sampling point on the convolutional regular grid, w is the weight of the convolution, x is the input feature map, p0 is the center of the grid, and p n traverses the positions of each sampling point in n is the offset of each sampling point learned by the network, and the offset sampling position is p n +Δp n In addition, by adding the weight coefficient Δm n to distinguish whether the introduced area is the area of the target of interest. If the convolution is not interested in the area of this sampling point, the weight coefficient is set to 0.

4. The multi-spectral pedestrian detection method based on an intelligent vehicle according to claim 1, wherein In step 5), perform high-level semantic feature extraction on the fused features of the highest layer to obtain high-level semantic information, specifically as follows: 5.1) Perform pyramid pooling on the fused features conv5_x_f of the highest layer to obtain high-level semantic features f1; 5.1.1) Input the fused features conv5_x_f into the pooling layer sub-branches of 4 scales in sequence. The pooling layer branches divide the feature map into different sub-regions and form different pooled feature representations. The spatial sizes of the feature maps output by the 4 pooling layer sub-branches are 1×1, 2×2, 4×4, and 6×6 respectively; 5.1.2) When the input feature channels are N, use a 1×1 convolution layer after each pooled feature to reduce the number of channels of the feature to N / 4; 5.1.3) Upsample the dimension-reduced features through bilinear interpolation to obtain features of the same size as the fused features conv5_x_f, and splice the features of different branches to obtain the final output feature f1 after pyramid pooling. The scale and number of channels of f1 are the same as those of the input features.

5. The multi-spectral pedestrian detection method based on an intelligent vehicle according to claim 1, wherein In step 5), perform high-level semantic feature extraction on the fused features of the highest layer to obtain high-level semantic information, specifically as follows: 5.2) Perform cross-attention processing on the fused features conv5_x_f of the highest layer to obtain high-level semantic features f2; 5.2.1) Input Feature Map Apply two 1×1 convolutional layers on conv5_x_f to generate two feature maps Q and K respectively, where C′ is the number of channels, and after dimensionality reduction calculation, the number of channels C′ is less than C; Among them, represents the input feature map, W represents the width of the input feature map, H represents the height of the input feature map, C and C′ are the number of channels, and after dimensionality reduction calculation, the number of channels C′ is less than the number of channels C; 5.2.2) At each position u in the spatial dimension of Q, obtain a vector Extract the feature vectors at the positions in K that are in the same row and the same column as position u to obtain a set Calculate the correlation degree of the two elements at the cross position; is the i-th element of Ω u The formula for calculating the correlation degree is as follows: where d i,u ∈ D is the correlation degree between feature Q u and Ω i,u and i = [1,..., H + W - 1], 5.2.3) Apply a softmax layer to the correlation degree set D to calculate the attention map A; 5.2.4) Apply another 1×1 convolutional layer on conv5_x_f to generate for adjusting features; at each position u in the spatial dimension of V, a vector and a set The set Φ u is the set of feature vectors in V that are in the same row or column as the position u, and the context information of the highest-level fused features can be collected by the following formula: Among them, [conv5_x_f]' u is the feature vector at position u, and A i,u is the attention scalar value at channel i and position u in A; adding the context information [conv5_x_f]' u to the feature [conv5_x_f] u can enhance the pixel-level representation of the feature; the high-level semantic feature finally output by the cross-attention module is [conv5_x_f]', denoted as f2.

6. The multi-spectral pedestrian detection method based on an intelligent vehicle according to claim 1, wherein In step 6), transfer the high-level semantic information to the shallow fused features of different scales and fuse them through a feature aggregation module. The feature aggregation module includes operations such as feature upsampling, feature fusion, and depth pooling. The specific steps are as follows: S61) Upsample the high-level semantic features to obtain the same scale as the shallow fused features; S62) Upsample the fused features of the previous stage by 2 times to obtain the same scale as the adjacent shallow fused features; S63) Fuse the upsampled high-level semantic features, the fused features of the previous stage after 2 times upsampling, and the shallow fused features of this stage by element-wise addition; S64) Perform depth pooling on the fused features obtained in S63 to reduce the interference caused by upsampling of high-level semantic features and the fused features of the previous stage; for the depth pooling process, first input the fused feature maps into the average pooling layer at different downsampling rates to convert them into different proportional spaces, and then input them into a 3×3 convolutional layer respectively for feature learning. Next, upsample the features after convolution, then merge the upsampled feature maps from different branches together, and finally process the fused features through a 3×3 convolutional layer; The depth pooling process allows each spatial position to view the context information of the feature map in different scale spaces, further refining the features after fusing high-level semantic information.

7. A multi-spectral pedestrian detection system based on intelligent vehicles, characterized in that, The system is used to deploy the multi-spectral pedestrian detection method based on an intelligent vehicle as described in any one of claims 1 to 6. The system includes a data acquisition module, a calculation and storage module, and a display output module; among them, the data acquisition module includes a visible light camera and an infrared camera; the calculation and storage module is implemented using a development board. By deploying the trained and optimized multi-spectral pedestrian detection model in the memory of the development board and running the model on the development board, the detection results can be obtained; the display output module displays the prediction boxes of pedestrians in the corresponding frame acquired by the camera through a display; the specific steps of the system are as follows: Step 1: Train and optimize a multi-spectral pedestrian detection model based on an intelligent vehicle using a multi-spectral pedestrian dataset; Step 2: Deploy and store the trained multi-spectral pedestrian detection model in the memory of the development board; Step 3: Acquire picture input data based on the visible light and infrared cameras in the data acquisition module and transmit the input data to the processor of the development board; Step 4: The processor preprocesses the input original picture data and simultaneously loads the detection inference code in the memory into the inference calculation module; Step 5: Input the data into the inference calculation module and run the detection model deployed in Step 2; Step 6: The inference calculation module outputs the detection results and outputs the detection boxes of pedestrians in the input picture through the display module.

8. A multi-spectral pedestrian detection device based on an intelligent vehicle, characterized in that, Including: An image preprocessing module for preprocessing paired visible light and infrared images; A gradient calculation module for calculating the gradient of the preprocessed infrared image, mining the significant gradient information in the infrared image, obtaining an edge feature map, and obtaining an infrared image with supplementary edge features; A feature extraction module for respectively inputting paired visible light and infrared images with supplementary edge features into the ResNet50 network for feature extraction and outputting the features at different stages in the feature extraction network; A feature fusion module that adds and fuses the output features of the same stage of the two modalities to obtain fused features with the same structure as ResNet50; among them, the output features select the highest-level features of each stage; A high-level semantic information extraction module for extracting high-level semantic features from the highest-level fused features of the last stage in the feature fusion module to obtain high-level semantic information; Obtain high-level semantic information as follows: 1) Perform pyramid pooling on the highest-level fused feature conv5_x_f to extract high-level semantic feature f1; 2) Perform cross-attention processing on the fused feature conv5_x_f at the highest level to extract the high-level semantic feature f2; 3) Add and fuse the information obtained by the two semantic information extraction sub-modules to obtain the fused high-level semantic feature f; A pedestrian position information enhancement module, which is used to perform pedestrian position information enhancement processing on the fused feature map supplemented with high-level semantic information to enhance the pedestrian position information in the fused feature map; The pedestrian position information enhancement processing is specifically as follows: use the existing pedestrian annotations to generate a binary image of 0 and 1 as the ground truth of the pedestrian position; predict the probability value of each position in the fused feature map being in the pedestrian area through supervised learning; multiply the predicted probability map element-wise with the fused features at different scales; use the product result as a residual to enhance the pedestrian position in the fused features; this process is expressed as: Among them, F is a set of fused feature maps of different scales for pedestrian detection, and F i refers to the i-th fused feature map; represents a convolution operation for predicting the possibility that each position in the feature map is in a pedestrian area; δ represents a sigmoid activation function for converting the predicted possibility into a probability value within 0 to 1; w i is a probability feature map, whose size is the same as that of F i is consistent; is a feature map with enhanced pedestrian area; A pedestrian detection module, which is used to perform pedestrian detection on the fused feature maps at different scales that have undergone feature supplementation and enhancement to obtain pedestrian detection results.