Multi-Spectral Pedestrian Detection Method Based on Cross-Modal Feature Enhancement and Confidence Fusion
By building a dual-popular person detection network in multi-spectral pedestrian detection, using the method of cross-modal feature enhancement and confidence fusion, the problems of limited expression ability of modal feature and unreliable fusion methods in the prior art are solved, and higher pedestrian detection accuracy and more reliable fusion effect are achieved.
Patent Information
- Application Number
- CN202310211007.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-03-07
AI Technical Summary
The existing multispectral infrared pedestrian detection methods have the shortcomings of limited modal feature expression capabilities, easy loss of key information, poor complementary capabilities between modals, and unreliable fusion methods, making it difficult to effectively deal with pedestrian targets in complex environments.
A multi-spectral pedestrian detection method based on cross-modal feature enhancement and confidence fusion is proposed. By building a dual-popular person detection network, using interactive shared attention module, pooled layer multi-scale adaptive fusion module and confidence fusion module, we effectively fuse multi-modal image features and enhance the modal feature expression ability.
It significantly improves the accuracy of pedestrian detection, effectively retains important information in pedestrian detection, suppresses and discards interference information, and improves the complementary ability and reliability of fusion between modes.
Smart Images

Figure CN116311364B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a multispectral pedestrian detection method based on cross-modal feature enhancement and confidence fusion, which can be used to process pedestrian targets under complex environmental conditions. Background Art
[0002] When a vehicle is driving, its environmental information is complex. It contains not only road information, other vehicle information, and pedestrian information, but also some uncertain information. As a direct protection object, the detection of pedestrians is of great significance. By detecting pedestrians within a certain distance, possible safety hazards can be perceived in advance, so as to take preventive actions according to the actual situation. At the same time, pedestrian detection is also an important technology in the field of unmanned driving and an indispensable part of unmanned driving.
[0003] Among many computer vision-based tasks, a large number of studies focus on visible light images. Visible light images are more favored by researchers because of their imaging mode that is more in line with the requirements of human visual perception. They are often used in tasks such as target detection, classification, and segmentation. Under normal circumstances, computer vision based on visible light images can meet the needs of tasks well. Early pedestrian detection was also based on visible light images. However, with the improvement of technical requirements, people gradually discovered that visible light images have many shortcomings, and ordinary pedestrian detection based on visible light images cannot meet actual needs well. When imaging, visible light images are based on a large number of photosensitive element arrays. By converting the light intensity in the environment into photoelectricity, analog electrical signals are obtained, and then the analog electrical signals are converted into common image pixel values through analog-to-digital conversion. Therefore, visible light images can form more detailed texture information, but this imaging method depends on external lighting information. Under weak external lighting conditions, the light sensitivity of visible light images will make it almost impossible for the image to capture effective information, thereby affecting the realization of specific tasks. In order to make up for the defects of light sensitivity in visible light images and improve safety, researchers have also proposed different solutions. Common methods include a combination of visible light imaging and infrared thermal imaging, a combination of visible light imaging and laser, etc.
[0004] Among the compensation methods in use, an infrared imaging sensor can capture the thermal radiation on the surface of an object, obtain the surface temperature distribution of the target object, and present it in the form of a digital image. Laser imaging measures the distance between the current position and the position of an obstacle based on the reflection time of a laser emission beam when it encounters an obstacle, and saves and utilizes it in the form of three-dimensional data. It can be seen from the imaging principle of the images that the above solutions do not rely on the external light intensity, and can well compensate for the defects caused by the light sensitivity of visible light images. As a living being with strong thermal radiation, pedestrians are more discriminative in infrared thermal imaging compared to the non-discriminative imaging of lasers, and are more suitable as supplementary information for visible light images. Therefore, the research on the fusion of visible light and infrared images has been increasing.
[0005] In "Multispectral pedestrian detection: Benchmark dataset and baseline" proposed by Hwang et al., while presenting the first dataset, they also gave their improved algorithm F+T+THOG on the visible light and infrared paired dataset based on the traditional single-modal object detection algorithm ACF.
[0006] In "Fusion of multispectral data through illumination-aware deep neural networks for pedestrian detection" proposed by Guan et al., starting from the characteristics of visible light and infrared data, they carefully designed an illumination awareness mechanism and reliably fused the input data under different illumination conditions according to the illumination intensity to improve the detection accuracy.
[0007] In "Multispectral pedestrian detection via simultaneous detection and segmentation" proposed by Li et al., with segmentation as an aid, the detection performance was further improved.
[0008] In "Improving multispectral pedestrian detection by addressing modality imbalance problems" proposed by Zhou et al., aiming at the modality imbalance problem existing between multi-modalities in visible light and infrared datasets, they proposed a modality balance network MBNet to solve this problem.
[0009] Although the methods proposed above are constantly improving to address existing problems, multi-spectral infrared detection still has drawbacks such as limited modal feature expression ability, easy loss of key information, poor complementary ability between modalities, unclear fusion focus, and unreliable fusion methods. How to more effectively process visible light and infrared data remains a question worthy of in-depth consideration. Summary of the Invention
[0010] The purpose of the present invention is to address the deficiencies of the above-mentioned existing technologies and propose a multi-spectral pedestrian detection method based on cross-modal feature enhancement and confidence fusion, which fully utilizes and fuses multi-modal image features to improve the accuracy of pedestrian detection.
[0011] To achieve the above purpose, the technical solution of the present invention is as follows:
[0012] (1) Obtain a dataset consisting of visible light images and infrared images, and divide it into a training set and a test set according to a ratio of 8:2.
[0013] (2) Construct a dual-stream pedestrian detection network α based on cross-modal feature enhancement and confidence fusion:
[0014] (2a) Establish an interactive common attention module composed of cross-cascaded multi-layer global pooling layers and multi-layer convolutional layers, which is used to calculate common attention weights and enhance features.
[0015] (2b) Establish a pooling layer multi-scale adaptive fusion module composed of one global pooling layer and two cascaded fully connected layers, which is used to encode the enhanced features to obtain fine-grained features.
[0016] (2c) Establish a modality-internal confidence module composed of four cascaded convolutional layers and a sigmoid function; establish a modality-intermediate confidence module composed of one convolutional layer and a sigmoid function, and connect the modality-internal confidence module and the modality-intermediate confidence module to form a confidence fusion module, which is used for confidence fusion of fine-grained features.
[0017] (2d) Add a common attention module between each convolutional layer of the existing VGG16 feature extraction backbone network, connect the existing RPN network after the last common attention module, and sequentially connect the multi-scale adaptive fusion module and the confidence fusion module after the RPN network to obtain a dual-stream pedestrian detection network based on cross-modal feature enhancement and confidence fusion.
[0018] (3) Input the training set into the dual-stream pedestrian detection network α, and train it using the mini-batch gradient descent method to obtain a trained prediction model.
[0019] (4) Input the test sets of visible light and infrared images into the prediction model to obtain the prediction results of pedestrian detection.
[0020] The present invention has the following advantages compared with the prior art:
[0021] 1) It effectively improves the expression ability of modal features
[0022] Due to the setting of the multi-spectral pedestrian detection network training framework for cross-modal feature enhancement and confidence fusion in the present invention, by using the interactive common attention module, the common information between multiple modalities is effectively fused, and the characteristics of single modalities are retained, enhancing the expression ability of the modal features of visible light features and infrared features; at the same time, due to the construction of a lightweight pooling layer multi-scale adaptive fusion module, taking the pooling layer as the processing object and giving an adaptive fusion method, the multi-scale information can be fused more flexibly, effectively improving the expression ability of the modal features of visible light and infrared.
[0023] 2) The pedestrian detection accuracy is higher
[0024] Aiming at the problem that important information of the input data is easily lost in the feature extraction process, the present invention constructs a confidence fusion module to constrain the texture information in the visible light image and the brightness information in the infrared image, retains the important information in pedestrian detection, suppresses and discards the interference information, can effectively evaluate the modality itself and the interaction between modalities, and obtains more reliable fusion information, making the pedestrian detection result of the present invention have a higher accuracy compared with other existing pedestrian detection algorithms. Brief Description of the Drawings
[0025] Figure 1 is the implementation flowchart of the present invention;
[0026] Figure 2 is the structure diagram of the dual-stream pedestrian detection network in the present invention;
[0027] Figure 3 is the structure diagram of the interactive common attention module in the dual-stream pedestrian detection network;
[0028] Figure 4 is the schematic diagram of the importance mask calculation of the interactive common attention module in the dual-stream pedestrian detection network;
[0029] Figure 5 is the structure diagram of the pooling layer multi-scale adaptive fusion module in the dual-stream pedestrian detection network;
[0030] Figure 6 is the structure diagram of the confidence fusion module in the dual-stream pedestrian detection network. Detailed Embodiment
[0031] The implementation process and effects of the present invention will be further described in detail below with reference to the accompanying drawings.
[0032] Refer to Figure 1, the implementation steps of the present invention are as follows:
[0033] Step 1. Obtain an image dataset and divide it into a training set and a test set.
[0034] (1.1) Obtain visible light and infrared images with the same resolution in different regions and different time periods, form pairs of visible light and infrared images, and use them as the image dataset;
[0035] (1.2) Randomly select pairs of visible light and infrared images from the image dataset according to a ratio of 8:2, and divide them into a training set and a test set.
[0036] Step 2. Construct a dual-stream pedestrian detection network based on cross-modal feature enhancement and confidence fusion.
[0037] Refer to Figure 2 , in this example, a common attention module is added between each convolutional layer of the existing VGG16 feature extraction backbone network. After the last layer of the interactive common attention module, the existing RPN network is connected. After the RPN network, a multi-scale adaptive fusion module and a confidence fusion module are connected in sequence to construct a dual-stream pedestrian detection network based on cross-modal feature enhancement and confidence fusion. The steps are as follows:
[0038] (2.1) Establish an interactive common attention module composed of cross-cascading of multiple layers of global pooling layers and multiple layers of convolutional layers. Among them, all the multiple layers of global pooling layers adopt global average pooling operations. The multiple layers of convolutional layers include two convolutional layers with a kernel size of 3, a stride of 1, an input channel of 2, and an output channel of 2, and one convolutional layer with a kernel size of 3, a stride of 1, an input channel of 2, and an output channel of 1;
[0039] (2.2) Establish a pooling layer multi-scale adaptive fusion module composed of one global pooling layer and two cascaded fully connected layers. Among them, the global pooling layer is implemented by global average pooling operation, and the output is a 512-dimensional vector. The fully connected layer is composed of a linear layer and a sigmoid function;
[0040] (2.3) Establish a confidence fusion module composed of a connection of an in-modal confidence module and an inter-modal confidence module;
[0041] The in-modal confidence module is an in-modal confidence module cascaded by four convolutional layers and a sigmoid function. Among them, the first convolutional layer has a kernel size of 3, a stride of 1, an input channel of 512, and an output channel of 128; the second convolutional layer has a kernel size of 3, a stride of 1, an input channel of 128, and an output channel of 32; the third convolutional layer has a kernel size of 3, a stride of 1, an input channel of 32, and an output channel of 8; the fourth convolutional layer has a kernel size of 1, a stride of 1, an input channel of 8, and an output channel of 1;
[0042] The inter-modal confidence module is an inter-modal confidence module cascaded by a convolutional layer and a sigmoid function. The convolutional layer has a kernel size of 3, a stride of 1, 2 input channels, and 1 output channel;
[0043] (2.4) Select the existing VGG16 feature extraction backbone network and the existing RPN network. Add the interactive common attention module between each convolutional layer of the VGG16 feature extraction network, connect the RPN network after the last interactive common attention module, and sequentially connect the multi-scale adaptive fusion module and the confidence fusion module after the RPN network to form a dual-stream pedestrian detection network based on cross-modal feature enhancement and confidence fusion.
[0044] Step 3. Use the mini-batch gradient descent algorithm to train the dual-stream pedestrian detection network to obtain a trained prediction model.
[0045] The specific implementation of this step is as follows:
[0046] (3.1) Input the training sets of visible light and infrared modality images into the existing VGG16 backbone network to extract the visible light image feature I rgb and the infrared image feature I t ;
[0047] (3.2) Enhance the two modality features through the interactive common attention module:
[0048] Refer to Figure 3 , and its enhancement steps are as follows:
[0049] (3.2.1) The interactive common attention module performs global average pooling on the infrared feature I t and the visible light feature I rgb respectively to obtain the infrared feature pooling vector V a and the visible light feature pooling vector V b ;
[0050] (3.2.2) Use the infrared feature pooling vector V a and the visible light feature pooling V b as inputs, and obtain the infrared feature weight vector ω1 and the visible light feature weight vector ω2 through multiple convolutional layers and a sigmoid function. Then, constrain these two feature weight vectors ω1 and ω2 between 0 and 1 through the Sigmoid function;
[0051] (3.2.3) Multiply the infrared feature I t by its weight vector ω1 in the channel direction to obtain the common feature O t extracted from the infrared feature; Multiply the visible light feature I rgbMultiply it with the weight vector ω2 in the channel direction to obtain the common feature O extracted from the visible light features rgb ;
[0052] (3.2.4) For the visible light feature I rgb and the infrared feature I t , first use the average operation in the channel direction to obtain the overall representation of a single channel, and then use the convolution and Sigmoid activation operations to obtain the importance mask ω of the feature map, as Figure 4 shown;
[0053] (3.2.5) Combine the importance mask ω with the infrared common feature O t and the visible light common feature O rgb respectively to obtain the enhanced infrared feature F rgb and the enhanced visible light feature F t :
[0054] F rgb = I rgb + O t × ω
[0055] F t = I t + O rgb × ω;
[0056] (3.3) Input the enhanced features into the existing RPN network to obtain the pooled features;
[0057] (3.4) Through the pooling layer multi-scale adaptive fusion module, obtain finer-grained features and achieve a better feature representation. Refer to Figure 5 , the specific implementation of this step is as follows:
[0058] (3.4.1) Take the two pooled features as the inputs of the pooling layer multi-scale adaptive fusion module, denoted as P1 and P2;
[0059] (3.4.2) Take the two pooled features of different scales P1 and P2 as inputs, and through the global pooling layer, output two 512-dimensional vectors;
[0060] (3.4.3) Concatenate the two vectors to obtain a 1024-dimensional vector, which contains the multi-scale information of the modality;
[0061] (3.4.4) Pass the 1024-dimensional vector through two independent fully connected layers to obtain two 512-dimensional small-scale weight vectors ω3 and large-scale weight vectors ω4, and use the Sigmoid activation function to constrain the small-scale weight vector ω3 and the large-scale weight vector ω4 between 0 and 1;
[0062] (3.4.5) Multiply the small-scale weight vector ω3 and the large-scale weight vector ω4 directly with the pooling layer features P1 and P2 of two different scales to weight the pooling features P1 and P2, and obtain the entire weighted fine-grained feature P3:
[0063]
[0064] where represents the activation function, represents the average pooling operation, and P3 is the output of the module, representing the fine-grained feature that fuses multi-scale features in an adaptive manner;
[0065] (3.5) Through the confidence fusion module, realize the confidence fusion of the fine-grained features of two modalities. Refer to Figure 6 , the specific implementation of this step is as follows:
[0066] (3.5.1) Input the visible light fine-grained feature into the built-in confidence module of the input modality to obtain the visible light confidence weight Input the infrared fine-grained feature F t (p) into the built-in confidence module of the input modality to obtain the infrared confidence weight
[0067] (3.5.2) Use the method of taking the mean to compress the visible light fine-grained feature and the infrared fine-grained feature F t (p) into a single channel in the channel direction to obtain the visible light pooling overall representation and the infrared overall representation;
[0068] (3.5.3) Input the visible light overall representation and the infrared overall representation into the inter-modal confidence module to obtain the interaction confidence C (IA) ;
[0069] (3.5.4) Calculate the reference modality confidence C (IA) according to the interaction confidence C , the visible light confidence weight and the infrared confidence weight t :
[0070]
[0071] (3.5.5) Calculate the confidence fusion result F according to the infrared fine-grained feature F t (p) , the reference modality confidence C t , the visible light fine-grained feature and the visible light confidence weight f :
[0072]
[0073] (3.6) Use the existing Faster_rcnn detection head to perform pedestrian target box prediction on the confidence fusion result F f to obtain the coordinates P of the predicted pedestrian target box r and the classification P c ;
[0074] (3.7) According to the coordinates P of the predicted pedestrian target box r and the classification P c and the coordinates G of the true pedestrian target box r and the classification G c , use the loss function to calculate the error L f ;
[0075] L f = (1 - 0.5·C f )[CE(P c , G c ) + S L1 (P r - G r )]
[0076] where C f represents the confidence of the fused feature, CE represents the cross-entropy loss, and S L1 represents the mean absolute error loss;
[0077] (3.8) According to the error value L f , use the backpropagation algorithm to update the network parameters;
[0078] (3.9) Repeat (3.1) to (3.8) until the loss function converges to obtain the trained prediction model.
[0079] Step 4. Test the trained pedestrian detection prediction model.
[0080] (4.1) Sequentially take c samples from the test set M and input them into the pedestrian detection network to obtain the detection result pred corresponding to each image Q c ; c ;
[0081] (4.2) Repeat (4.1) until all samples in the test set M have been detected to obtain the detection results of pedestrians.
[0082] The effects of the present invention can be further illustrated by the following simulation.
[0083] 1. Simulation data
[0084] The simulation data uses the KAIST dataset. The KAIST dataset is the first publicly available visible light and infrared image pedestrian detection dataset, which contains a total of 95,328 pairs of paired visible light and infrared images. The size of each image is 640×512, and a total of 103,128 targets are included. The test set of this data contains a total of 2,252 pairs of visible light and infrared data pairs, of which 1,455 pairs are from daytime and 797 pairs are from nighttime. The KAIST dataset contains daytime and nighttime scenes. Since the targets have the same absolute position in visible light and infrared images, it shows that the dataset has good inter-modal alignment conditions.
[0085] 2. Simulation content
[0086] Using the method of the present invention and the existing 9 methods of ACF, Halfway Fusion, Fusion RPN+BF, IAF R-CNN, IATDNN+IASS, MSDS-RCNN, MBNet, MLPD, and AR-CNN, the training set of the KAIST dataset is trained respectively to obtain trained prediction models. Then, the trained prediction models are used to test the multi-spectral pedestrian detection of the test set of the KAIST dataset, and the missing detection value MR-2 evaluation is carried out under 9 different settings of ALL pedestrians, Day pedestrians, Night pedestrians, Near pedestrians, Medium pedestrians, Far pedestrians, None-occluded pedestrians, Partial-occluded pedestrians, and Heavy-occluded pedestrians, as well as under different preset conditions IoU. The lower the MR-2, the better the detection ability. The comparison results are as follows in the table, where:
[0087] Table 1 shows the comparison results under the condition of IoU = 0.3.
[0088] Table 2 shows the comparison results under the condition of IoU = 0.5.
[0089] Table 3 shows the comparison results under the condition of IoU = 0.7.
[0090] Table 1 Comparison of the detection performance of the present invention and the existing 9 methods on the KAIST dataset (IoU = 0.3)
[0091]
[0092] Table 2 Comparison of the detection performance of the present invention and the existing 9 methods on the KAIST dataset (IoU = 0.5)
[0093]
[0094] Table 3 Comparison of Detection Performance between the Present Invention and Nine Existing Methods on the KAIST Dataset (IoU = 0.7)
[0095]
[0096] 3. Analysis of Simulation Results
[0097] As can be seen from Table 1, under the condition of IoU = 0.3, the present invention obtained performance indicators of 4.40%, 5.29%, and 2.75% respectively in the detection results of ALL, Day, and Night. Compared with the existing best MLPD method, the performance was improved by 0.52%, 0.15%, and 1.16% respectively.
[0098] As can be seen from Table 2, under the condition of IoU = 0.5, the present invention obtained performance indicators of 7.18%, 7.91%, and 5.75% respectively in the detection results of ALL, Day, and Night. Compared with the existing best MLPD method, the performance was improved by 0.3%, 0.26%, and 1.34% respectively.
[0099] As can be seen from Table 3, under the condition of IoU = 0.7, the present invention obtained performance indicators of 37.66%, 34.31%, and 43.35% respectively in the detection results of ALL, Day, and Night. Compared with the existing best MLPD method, the performance was improved by 5.16%, 7.74%, and 0.14% respectively.
[0100] The simulation results show that the method of the dual-stream pedestrian detection network based on cross-modal feature enhancement and confidence fusion of the present invention has the best pedestrian detection results and can effectively improve the accuracy of multi-spectral pedestrian detection.
[0101] The above description is only a specific example of the present invention and does not constitute any limitation to the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these modifications and changes based on the idea of the present invention are still within the protection scope of the claims of the present invention.
Claims
1. A multi-spectral pedestrian detection method based on cross-modal feature enhancement and confidence fusion, characterized in that Including: (1) Obtain visible light and infrared images to form a dataset, and divide it into a training set and a test set according to a ratio of 8:2; (2) Construct a dual-stream pedestrian detection network α based on cross-modal feature enhancement and confidence fusion: (2a) Establish an interactive common attention module composed of cross-cascading of multiple global pooling layers and multiple convolutional layers, which is used to calculate common attention weights and enhance features; (2b) Establish a pooling layer multi-scale adaptive fusion module composed of one global pooling layer and two cascaded fully connected layers, which is used to encode the enhanced features to obtain fine-grained features; (2c) Establish a modality-internal confidence module composed of four cascaded convolutional layers and a sigmoid function; establish a modality-intermediate confidence module composed of one convolutional layer and a sigmoid function, and connect the modality-internal confidence module and the modality-intermediate confidence module to form a confidence fusion module, which is used for confidence fusion of fine-grained features; (2d) Add a common attention module between each convolutional layer of the existing VGG16 feature extraction backbone network, connect the existing RPN network after the last common attention module, and sequentially connect the multi-scale adaptive fusion module and the confidence fusion module after the RPN network to obtain a dual-stream pedestrian detection network based on cross-modal feature enhancement and confidence fusion; (3) Input the training set into the dual-stream pedestrian detection network α, and train it using the mini-batch gradient descent method to obtain a trained prediction model; (4) Input the test sets of visible light and infrared images into the prediction model to obtain the prediction results of pedestrian detection.
2. The method according to claim 1, characterized in that In the common attention mechanism constructed in step (2a), the structures and parameters of each part are as follows: All of the multiple global pooling layers adopt global average pooling operations; The multiple convolutional layers include two convolutional layers with a convolutional kernel size of 3, a stride of 1, an input channel of 2, and an output channel of 2, and one convolutional layer with a convolutional kernel size of 3, a stride of 1, an input channel of 2, and an output channel of 1.
3. The method according to claim 1, characterized in that In step (2a), the interactive common attention module calculates common attention weights and enhances features as follows: (2a1) Infrared feature I t and visible light feature I rgb are respectively subjected to global average pooling to obtain an infrared feature pooling vector V a and a visible light feature pooling vector V b ; (2a2) Pool the infrared feature vector V a and the visible light feature pooling V b As inputs, pass through a multi-layer convolutional layer and a sigmoid function to obtain the infrared feature weight vector ω1 and the visible light feature weight vector ω2, and then constrain these two feature weight vectors ω1 and ω2 between 0 and 1 through the sigmoid function; (2a3) Multiply the infrared feature I t by its weight vector ω1 in the channel direction to obtain the common feature O extracted from the infrared feature t ; Multiply the visible light feature I rgb by its weight vector ω2 in the channel direction to obtain the common feature O extracted from the visible light feature rgb ; (2a4) Visible light feature I rgb and infrared feature I t , first use the average operation in the channel direction to obtain the overall representation of a single channel, and then use convolution and Sigmoid activation operations to obtain the importance mask ω of the feature map; (2a5) Combine the importance mask ω with the infrared common feature O t and the visible light common feature O rgb respectively to obtain the enhanced infrared feature F rgb and the enhanced visible light feature F t : F rgb = I rgb + O t × ω F t = I t + O rgb × ω。 4. The method according to claim 1, characterized in that, In the pooling layer multi-scale adaptive fusion module constructed in step (2b), the structures and parameters of each part are as follows: The global pooling is implemented by global average pooling operation, and the output is a 512-dimensional vector; The fully connected layer is composed of a linear layer and a sigmoid function.
5. The method according to claim 1, characterized in that, In step (2b), the enhanced features are encoded using the pooling layer multi-scale adaptive fusion module as follows: (2b1) Take the pooling layer features P1 and P2 of two different scales as inputs, and through the global pooling layer, the output is two 512-dimensional vectors; (2b2) Concatenate the two vectors to obtain a 1024-dimensional vector; (2b3) The 1024-dimensional vector passes through two independent fully connected layers to obtain two 512-dimensional small-scale weight vectors ω3 and large-scale weight vectors ω4, and the Sigmoid activation function is used to constrain the small-scale weight vector ω3 and the large-scale weight vector ω4 between 0 and 1; (2b4) Multiply the small-scale weight vector ω3 and the large-scale weight vector ω4 directly with the pooling layer features P1 and P2 of two different scales to weight the pooling features P1 and P2, and obtain the overall weighted pooling feature P3: Among them represents the activation function represents the average pooling operation, and P3 is the output of the module, representing the result of fusing multi-scale features in an adaptive manner.
6. The method according to claim 1, characterized in that, In the confidence fusion module constructed in step (2c), the structures and parameters of its various parts are as follows: In the intra-modal confidence module, the first convolutional layer has a kernel size of 3, a stride of 1, an input channel of 512, and an output channel of 128; the second convolutional layer has a kernel size of 3, a stride of 1, an input channel of 128, and an output channel of 32; the third convolutional layer has a kernel size of 3, a stride of 1, an input channel of 32, and an output channel of 8; the fourth convolutional layer has a kernel size of 1, a stride of 1, an input channel of 8, and an output channel of 1; The inter-modal confidence module consists of a convolutional layer with a kernel size of 3, a stride of 1, an input channel of 2, and an output channel of 1, and a sigmoid function.
7. The method according to claim 1, characterized in that, In step (2c), the confidence fusion module is used for the confidence fusion of fine-grained features, and the implementation is as follows: (2c1) Input the visible light fine-grained features and the infrared fine-grained feature F t (p) into the modality-built confidence module respectively to obtain the visible light confidence weight and the infrared confidence weight (2c2) Compress the visible light fine-grained features and the infrared fine-grained features F t (p) Use the method of taking the mean to compress them into a single channel in the channel direction to obtain the overall visible light representation and the overall infrared representation; (2c3) Input the overall visible light characterization and the overall infrared characterization into the inter-modal confidence module to obtain the interaction confidence C (IA) ; (2c4)According to the interaction confidence C (IA) , the optical confidence weight and the infrared confidence weight calculate the reference modal confidence C t : According to the infrared fine-grained feature F t (p) , the reference modal confidence C t , the visible light fine-grained feature and the visible light confidence weight calculate the confidence fusion result F f :
8. The method according to claim 1, wherein In step (3), the prediction model is trained using the mini-batch gradient descent method, and the implementation is as follows: (3a) Use the existing Faster_rcnn detection head to perform pedestrian target box prediction on the confidence fusion result F f to obtain the coordinates P r of the pedestrian target box and the classification P c ; (3b) According to the coordinates P of the predicted pedestrian target box r and the classification P c and the coordinates G of the true pedestrian target box r and the classification G c , use the loss function to calculate its error L f ; L f = (1 - 0.5·C f )[CE(P c , G c ) + S L1 (P r - G r )] Among them, C f represents the confidence of the fusion feature, CE represents the cross-entropy loss, and S L1 represents the mean absolute error loss; (3c) According to the error value L f , update the network parameters using the backpropagation algorithm (3d) Repeat (3a) to (3c) until the loss function converges to obtain the trained prediction model.
9. The method according to claim 1, wherein In step (4), the trained pedestrian detection prediction model is tested using the test set data M, and the implementation is as follows: (4a) Take c samples from the test set M in sequence and input them into the pedestrian detection network to obtain the detection result pred corresponding to each image Q c corresponding detection result pred c ; (4b) Repeat (4a) until all samples in the test set M are detected to obtain the detection results of multi-spectral pedestrians.
Citation Information
Patent Citations
Multi-modal pedestrian detection method based on improved YOLO model
CN111767882A
Multispectral pedestrian detection method based on multi-stage feature fusion information multiplexing
CN113361475A