Adaptive multispectral pedestrian detection method based on attention mechanism

By adopting an adaptive multi-spectral detection method based on attention mechanism in pedestrian detection, the visible light and infrared image characteristics are fused, and the problem of low detection accuracy under harsh light conditions is solved, achieving efficient and accurate pedestrian detection.

CN120107995APending Publication Date: 2025-06-06NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510036851.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art pedestrian detection accuracy is low under harsh lighting conditions, and the multi-spectral fusion method of deep neural networks consumes a lot of computing resources, which affects performance.

Method used

Adaptive multispectral pedestrian detection method based on attention mechanism is adopted, and the visible light and infrared image characteristics are fused through adaptive network technology and attention mechanism, and network paths are adaptively selected to achieve high detection rate and low leakage detection rate.

Benefits of technology

Ensure the accuracy of pedestrian identification under various light conditions, balance detection speed and accuracy, and improve the overall performance and robustness of multi-spectral pedestrian detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107995A_ABST
    Figure CN120107995A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive multispectral pedestrian detection method based on an attention mechanism, and belongs to the technical field of intelligent driving. Extracting a visible light preliminary feature and an infrared preliminary feature by using a first double-flow backbone network through an adaptive multispectral pedestrian detection model; obtaining a detection difficulty coefficient based on the visible light preliminary features and the infrared preliminary features, and judging a current detection scene; obtaining a pedestrian detection result by using the visible light preliminary features and the infrared preliminary features in a simple scene; in a difficult scene, visible light secondary features and infrared secondary features are extracted through a second double-flow backbone network, the visible light secondary features, the infrared secondary features, the visible light preliminary features and the infrared preliminary features are fused, and fused visible light features and fused infrared features are obtained and used for predicting pedestrian detection results. The pedestrian detection precision is improved by fusing the features of the two images; based on detection scene distinguishing, the detection speed and the detection precision are balanced, and the overall performance of pedestrian detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving technology, and in particular to an adaptive multi-spectral pedestrian detection method based on an attention mechanism. Background Art

[0002] Visible light imaging systems have become an important basis for the study of pedestrian detection algorithms due to their high resolution, good signal-to-noise ratio, strong contrast, and low cost. In daytime or well-lit environments, they can achieve relatively accurate detection results. However, they perform poorly in some harsh environments. For example, in low light, at night, or in complex weather conditions (such as rain, fog, and haze), the performance of visible light imaging systems will be greatly affected. The advantage of infrared imaging systems is that they can provide high-quality imaging under poor lighting conditions. Since any object with a temperature above absolute zero will release infrared radiation, infrared imaging systems can capture the thermal radiation information of the surrounding environment to identify and detect targets. Especially at night, in dense fog, smoke, or other complex environments, infrared imaging systems can provide stable and reliable performance, are not limited by visible light, and have strong adaptability and concealment. However, in environments such as strong light and high temperature, the imaging quality of infrared imaging systems will drop significantly, resulting in reduced detection accuracy. In addition, infrared imaging systems usually have weak detection capabilities for small targets at long distances, and are prone to problems such as image blur and loss of details. Visible light imaging systems have higher resolution and clarity in environments with good lighting conditions and are suitable for normal daytime scenes, while infrared imaging systems can play an important role in low light, bad weather or hidden environments. Therefore, combining these two imaging technologies to form a multi-sensor fusion system can effectively make up for the shortcomings of a single sensor and improve the robustness and adaptability of the overall detection system.

[0003] In the existing technology, the fusion of infrared imaging and visible light imaging is performed through deep neural networks to provide more contextual information, but it also means that more computing resources and time are required, and the added information may bring redundancy and overlap, thereby affecting network performance and may even cause performance degradation.

[0004] Therefore, a multispectral pedestrian detection method that can balance detection accuracy and speed is needed. Summary of the invention

[0005] In view of this, the present invention provides an adaptive multispectral pedestrian detection method based on attention mechanism, which is based on adaptive network technology, combined with attention mechanism, and integrated with multispectral feature analysis, so as to fully explore the advantages of visible light and infrared images in different lighting environments and their complementarity, and realize pedestrian detection with high detection rate and low missed detection rate.

[0006] To this end, the present invention provides the following technical solutions:

[0007] An adaptive multispectral pedestrian detection method based on attention mechanism, comprising:

[0008] The adaptive multispectral pedestrian detection model receives visible light images and infrared images;

[0009] Extracting preliminary visible light features and preliminary infrared features through the first two-stream backbone network;

[0010] Obtaining a detection difficulty coefficient based on the visible light preliminary features and the infrared preliminary features;

[0011] Determining whether the detection of the visible light image and the infrared image is an easy scene or a difficult scene according to the detection difficulty coefficient;

[0012] If it is an easy scene, the pedestrian detection result is obtained by using the visible light preliminary features and the infrared preliminary features through the first neck network and the first detection head;

[0013] If it is a difficult scene, the visible light secondary features and infrared secondary features are extracted through the second two-stream backbone network;

[0014] Fusing the visible light secondary feature, the infrared secondary feature, the visible light preliminary feature, and the infrared preliminary feature to obtain a fused visible light feature and a fused infrared feature;

[0015] Pedestrian detection results are obtained by using the fused visible light features and the fused infrared features through the second neck network and the second detection head.

[0016] Furthermore, obtaining the detection difficulty coefficient based on the visible light preliminary features and the infrared preliminary features includes:

[0017] Splicing the visible light preliminary features of each stage and the infrared preliminary features of each stage to obtain splicing features of each stage;

[0018] Perform global average pooling on the concatenated features of each stage to obtain the global feature vector of each stage;

[0019] The global feature vectors of each stage are concatenated to obtain the full-stage vector;

[0020] Input the full-stage vector into the fully connected layer to output the detection difficulty coefficient.

[0021] Furthermore, the fusing of the visible light secondary feature, the infrared secondary feature, the visible light preliminary feature and the infrared preliminary feature to obtain the fused visible light feature and the fused infrared feature includes:

[0022] Connecting the visible light secondary feature and the visible light preliminary feature to obtain the visible light connection feature;

[0023] Connecting the infrared secondary feature and the infrared preliminary feature to obtain the infrared connection feature;

[0024] Performing feature enhancement on the visible light connection feature to obtain an enhanced visible light feature;

[0025] Performing feature enhancement on the infrared connection feature to obtain an enhanced infrared feature;

[0026] The enhanced visible light features and the enhanced infrared features are interactively fused to obtain fused visible light features and fused infrared features.

[0027] Furthermore, the detection of the visible light image and the infrared image is judged as an easy scene or a difficult scene by the detection difficulty coefficient:

[0028] Preset detection difficulty coefficient threshold;

[0029] If the detection difficulty coefficient is greater than the detection difficulty coefficient threshold, it is a difficult scenario;

[0030] If the detection difficulty coefficient is less than or equal to the detection difficulty coefficient threshold, it is an easy scenario.

[0031] Furthermore, the visible light connection feature is obtained by connecting the visible light secondary feature and the visible light preliminary feature in a same-level connection manner.

[0032] Furthermore, the infrared connection feature is obtained by connecting the infrared secondary feature and the infrared preliminary feature in a same-level connection manner.

[0033] Furthermore, the visible light connection feature is enhanced by a deformable self-attention mechanism to obtain the enhanced visible light feature.

[0034] Furthermore, the infrared connection feature is enhanced by using a deformable self-attention mechanism to obtain the enhanced infrared feature.

[0035] Furthermore, the interactive fusion of the enhanced visible light feature and the enhanced infrared feature to obtain the fused visible light feature and the fused infrared feature includes:

[0036] Use the enhanced visible light feature as the query and the enhanced infrared feature as the key to calculate the visible light attention weight;

[0037] Use the enhanced infrared feature as query and the enhanced visible light feature as key to calculate the infrared attention weight;

[0038] Perform weighted summation of the values ​​according to the visible light attention weights to obtain updated visible light features;

[0039] Perform weighted summation of the values ​​according to the infrared attention weights to obtain updated infrared features;

[0040] The updated visible light features and the enhanced visible light features are connected through residual connection to obtain the fused visible light features;

[0041] The updated infrared features and enhanced infrared features are connected via residual connections to obtain fused features.

[0042] Furthermore, the training process of the adaptive multi-spectral pedestrian detection model includes:

[0043] In the first stage, the ability of the adaptive multi-spectral pedestrian detection model to detect pedestrians is optimized with the goal of minimizing the losses of the first backbone network and the second backbone network;

[0044] In the second stage, the adaptive multispectral pedestrian detection model is trained to distinguish between visible light image and infrared image detection scenes;

[0045] When the adaptive multi-spectral pedestrian detection model training is in the second stage, the training parameters of the first stage are not updated.

[0046] Advantages and positive effects of the present invention:

[0047] The present invention performs multispectral pedestrian detection by fusing infrared image and visible light image information, thereby ensuring the accuracy of pedestrian recognition under a variety of lighting conditions; and adopts an adaptive decision maker to make decisions based on the detection scenario, adaptively select the optimal network path, balance the detection speed and detection accuracy, and improve the overall performance of multispectral pedestrian detection; in difficult detection scenarios, the deformable self-attention mechanism and cross-attention mechanism are introduced for feature enhancement and fusion, which fully integrates the features of the two images, greatly reduces the missed detection rate of pedestrian detection, and improves the accuracy and robustness of multispectral pedestrian detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0049] Figure 1 is a flow chart of an adaptive multi-spectral pedestrian detection method based on the attention mechanism in this implementation;

[0050] Figure 2 This is a diagram of the backbone network structure in this implementation mode;

[0051] Figure 3 This is a structural diagram of the adaptive decision maker in this implementation mode;

[0052] Figure 4 This is a cascade structure diagram of the first dual-stream backbone network and the second dual-stream backbone network in this implementation manner;

[0053] Figure 5 This is the structural diagram of the attention mechanism in this implementation. DETAILED DESCRIPTION

[0054] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0055] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0056] The present invention provides an adaptive multi-spectral pedestrian detection method based on an attention mechanism. The method comprises the following steps: preliminarily extracting visible light and infrared features through a first dual-stream backbone network; splicing the preliminarily extracted visible light and infrared features and inputting them into an adaptive decision maker, which outputs a current detection difficulty coefficient and makes a decision; if the decision result is "easy", the visible light and infrared features are input into a first neck network and a head network to obtain a pedestrian detection result; if the decision result is "difficult", the visible light and infrared features are secondary extracted through a second dual-stream backbone network; splicing the secondary extracted visible light and infrared features with the preliminarily extracted features through an SLC (Same Level Composition) method; in order to make full use of the advantages of visible light and infrared images under different lighting environments and their complementarity, a deformable self-attention mechanism is used for feature enhancement so that the model focuses on pedestrians, and a cross-attention mechanism is used to complete the interactive fusion of visible light and infrared features; the fused features are input into a second neck network and a head network to obtain a pedestrian detection result.

[0057] The present invention provides an adaptive multi-spectral pedestrian detection method based on attention mechanism, the process is as follows Figure 1 As shown, specifically including:

[0058] S1. Obtain the visible light image and infrared image of the same area at the current moment. For the input visible light image X R and infrared image X T , using the first dual-stream backbone network B 1 Extract preliminary multi-scale features R of visible light images and infrared images respectively 1 , Y 1 ; The preliminary multi-scale features are divided into five stages, among which,

[0059] The first two-stream backbone network is composed of multiple modules. Among them, the CBS module consists of a convolutional layer, batch normalization BN, and SiLU activation function. The ELAN module is composed of a series of CBS modules, and its input feature map is the same size as the output feature map. The MP module consists of maximum pooling and convolutional layers. The upper branch uses maximum pooling to achieve spatial downsampling, followed by a CBS module to compress the channel; the lower branch first compresses the channel through the CBS module, and then uses a 3x3 convolution with a step size of 2 for downsampling. Finally, the outputs of the two branches are merged to produce a feature map with the same number of channels as the input but half the spatial resolution. The input image first passes through a series of convolutional layers, namely 4 CBS modules, to obtain a four-fold downsampled feature map; then the downsampled feature map is processed by the ELAN module. It is worth noting that ELAN does not change the size of the input and output feature maps. Therefore, after the ELAN operation, the size of the feature map obtained is still the four-fold downsampled feature map; then the MP module is used to perform downsampling again, and the ELAN and MP modules are repeatedly used for feature extraction to obtain feature maps of five sizes for subsequent use.

[0060] S2, the preliminary multi-scale feature R 1 、T 1 Send it to the adaptive decision maker Router to obtain the detection difficulty coefficient of the current detection frame

[0061] The structure of the adaptive decision maker is as follows: Figure 3 As shown, the first dual-stream backbone network B 1 Extracted visible light features of the same scale and infrared characteristics Splice to RT i , where i is the stage number, ranging from 0 to 5, and then compressed into a global feature vector through global average pooling Then the five stages Splice into a vector RT gap Then, this vector is passed through two learnable fully connected layers FC to obtain the detection difficulty coefficient of the current detection frame

[0062] S3. Set thresholds based on actual scenario requirements Set the detection difficulty coefficient of the current detection frame With threshold Make comparisons and determine the detection scenarios;

[0063] S4. If Less than or equal to The current detection frame is divided into an “easy” scene, and the multi-scale features R 1 , T 1 A simple addition is performed and the multi-scale features are input into the first neck network for fusion. Then the multi-scale features are input into the first head network. Three detection heads are used to perform pedestrian detection on the feature maps of the three scales respectively, and the final pedestrian detection results are obtained, including the target category and target position.

[0064] S5. If Greater than The current detection frame is classified as a "difficult" scene, and the second dual-stream backbone network B is used. 2 Extract the secondary multi-scale features R of visible light image and infrared image respectively 2 , T 2 ,The secondary multi-scale features are divided into five stages, among which,

[0065] Second dual-stream backbone network B 2 With the first two-stream backbone network B 1 The structure is the same.

[0066] Through the same level connection, the features R extracted by the first two-stream backbone network are respectively 1 and T 1 Add element-wise to R 2 and T 2 On the top, we get the connection feature R c and T c ;in,

[0067] Considering the difference between visible light and infrared spectrum in representing the real world, an efficient fusion mechanism is used to interact between the modalities, thereby gaining a gain in the final pedestrian detection effect. and and and Self-attention enhancement and cross-attention enhancement are performed on the features of three scales to obtain fusion features and

[0068] The fusion features and The input is processed by the second neck network, and the final pedestrian detection result, including the target category and target position, is obtained through the second head network detection.

[0069] Specifically, in and and and Self-attention enhancement and cross-attention enhancement are performed on features of three scales, including:

[0070] 1) In and and and Deformable self-attention is used on features of three scales to perform self-attention calculation to enhance feature expression, improve computational efficiency and focus on more relevant spatial regions.

[0071] Generate enhanced features R with feature dimension N×d through deformable self-attention mechanism dsa and T dsa ;in and (i=3, 4, 5).

[0072] The dimensions of visible light and infrared features are converted from H×W×C to N×d, and the spatial dimension H×W of the feature map is flattened into one dimension; H×W pixels are flattened into a sequence, and each pixel (spatial position) corresponds to a token. After flattening, the dimension of the feature map becomes N×C, where N=H×W, and the feature vector dimension of each token is C, which corresponds to the number of channels of the original feature map. Through a learned weight matrix W proj Convert C dimension to d dimension. Specifically, the linear transformation can be expressed as:

[0073] Map the input to Query, Key, and Value; through linear transformation, we get:

[0074]

[0075] Each query has a corresponding offset δ, which indicates the specific position that the query should focus on in the key space.

[0076] δ i =f offset (qi )

[0077] Among them, the offset δ i is a query q i The associated vector, f offset It is a fully connected layer that calculates the offset based on the query information.

[0078] Set query q i The position is i, and the offset δ i , the model determines the position j it needs to focus on as: j = i + δ i .

[0079] For each query q i , a small set is selected from the corresponding neighborhood (determined by the offset) to calculate the attention. This is a localized attention mechanism that reduces the amount of computation.

[0080] That is, for query q i , select neighborhood N i The key k in j , where j∈NW i , indicating that i The similarity between each query and the selected key is calculated using scaled dot product attention:

[0081]

[0082] The similarity between the query and the selected key is normalized by the softmax function to obtain the attention weight:

[0083] α ij =softmax(attn ij )

[0084] The weighted sum of the output values:

[0085]

[0086] Among them, v j is the corresponding value.

[0087] 2) Through cross-attention, the deep fusion of the two visual modalities is achieved, allowing the visible light features and infrared features to guide each other and enhance the multi-spectral perception ability.

[0088] Generate query (Q), key (K), and value (V) for visible and infrared modalities respectively;

[0089] For visible light features:

[0090] For infrared signatures:

[0091] in, is a learnable linear projection matrix.

[0092] Using the visible light feature as the query and the infrared feature as the key, calculate the visible light attention weight:

[0093]

[0094] Similarly, using the visible light feature as the query and the infrared feature as the key, the infrared attention weight is calculated:

[0095]

[0096] The values ​​are weighted summed according to the attention weights to obtain updated visible and infrared features:

[0097]

[0098] The updated features are connected to the original features through residual connection to obtain fusion features, namely:

[0099]

[0100] In order to further extract high-level semantic information, a feed-forward network FFN is added to the fused features to restore the feature dimension to H×W×C.

[0101] The training of the model is divided into two stages. In the first stage, the parameters of the adaptive decision maker are not trained. Only the two backbone networks and two detection heads are jointly trained. The training objective is defined as:

[0102]

[0103] Where x and y represent the input image and the true label respectively, θ i represents the learnable parameters of the i-th detection head, represents the training loss of the i-th detection head.

[0104] In the first stage, the dual backbone network is trained. The goal of the first stage training is to minimize the sum of the losses of the backbone networks B1 and B2. The loss function consists of three parts: classification loss, target confidence loss, and bounding box regression loss. After the first stage of training, the cascaded dual backbone network has the ability to detect. In the second stage of training the adaptive decision maker, θ 1 ,θ 2 will be frozen and not updated during the second phase of training.

[0105] After completing the first stage of training, the cascade detector has the ability to detect pedestrians. In the second stage, the parameters θ are frozen. 1,θ 2 , train the adaptive decision maker to automatically distinguish the scenes of the image. The loss function of this stage of training is:

[0106]

[0107] in, is the detection difficulty coefficient calculated by the adaptive decision maker, Δ represents the loss difference between two detectors checking pedestrians at the same time, If the loss difference Δ is small enough, the detection scenario of this image is classified as “easy”. On the contrary, if the loss difference is large enough, it should be classified as “difficult”.

[0108] According to the method of the present invention, during the detection process, the first detection head is used to immediately complete pedestrian detection for the easy images after extracting features using the first dual-stream backbone network, while the second detection head is used to complete pedestrian detection for the more difficult images after enhancement using the second dual-stream backbone network. This can complete adaptive multi-spectral pedestrian detection and achieve a balance between speed and accuracy.

[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An adaptive multispectral pedestrian detection method based on attention mechanism, characterized in that: include: The adaptive multispectral pedestrian detection model receives visible light images and infrared images; Extracting preliminary visible light features and preliminary infrared features through the first two-stream backbone network; Obtaining a detection difficulty coefficient based on the visible light preliminary features and the infrared preliminary features; Determining whether the detection of the visible light image and the infrared image is an easy scene or a difficult scene according to the detection difficulty coefficient; If it is an easy scene, the pedestrian detection result is obtained by using the visible light preliminary features and the infrared preliminary features through the first neck network and the first detection head; If it is a difficult scene, the visible light secondary features and infrared secondary features are extracted through the second two-stream backbone network; Fusing the visible light secondary feature, the infrared secondary feature, the visible light preliminary feature, and the infrared preliminary feature to obtain a fused visible light feature and a fused infrared feature; Pedestrian detection results are obtained by using the fused visible light features and the fused infrared features through the second neck network and the second detection head.

2. According to claim 1, an adaptive multispectral pedestrian detection method based on attention mechanism is characterized in that: The obtaining of the detection difficulty coefficient based on the visible light preliminary features and the infrared preliminary features includes: Splicing the visible light preliminary features of each stage and the infrared preliminary features of each stage to obtain splicing features of each stage; Perform global average pooling on the concatenated features of each stage to obtain the global feature vector of each stage; The global feature vectors of each stage are concatenated to obtain the full-stage vector; Input the full-stage vector into the fully connected layer to output the detection difficulty coefficient.

3. The adaptive multispectral pedestrian detection method based on attention mechanism according to claim 1, characterized in that: The step of fusing the visible light secondary feature, the infrared secondary feature, the visible light preliminary feature and the infrared preliminary feature to obtain the fused visible light feature and the fused infrared feature includes: Connecting the visible light secondary feature and the visible light preliminary feature to obtain the visible light connection feature; Connecting the infrared secondary feature and the infrared preliminary feature to obtain the infrared connection feature; Performing feature enhancement on the visible light connection feature to obtain an enhanced visible light feature; Performing feature enhancement on the infrared connection feature to obtain an enhanced infrared feature; The enhanced visible light features and the enhanced infrared features are interactively fused to obtain fused visible light features and fused infrared features.

4. The method for adaptive multispectral pedestrian detection based on attention mechanism according to claim 2, characterized in that: The detection difficulty coefficient is used to determine whether the detection of the visible light image and the infrared image is an easy scene or a difficult scene: Preset detection difficulty coefficient threshold; If the detection difficulty coefficient is greater than the detection difficulty coefficient threshold, it is a difficult scenario; If the detection difficulty coefficient is less than or equal to the detection difficulty coefficient threshold, it is an easy scenario.

5. The method of adaptive multispectral pedestrian detection based on attention mechanism according to claim 3, characterized in that: The visible light connection feature is obtained by connecting the visible light secondary feature and the visible light preliminary feature in a same-level connection manner.

6. The method of adaptive multispectral pedestrian detection based on attention mechanism according to claim 3, characterized in that: The infrared connection feature is obtained by connecting the infrared secondary feature and the infrared preliminary feature in a same-level connection manner.

7. The method of adaptive multispectral pedestrian detection based on attention mechanism according to claim 5, characterized in that: The enhanced visible light feature is obtained by performing feature enhancement on the visible light connection feature through a deformable self-attention mechanism.

8. The method of adaptive multispectral pedestrian detection based on attention mechanism according to claim 6, characterized in that: The infrared connection feature is enhanced by a deformable self-attention mechanism to obtain the enhanced infrared feature.

9. The method for adaptive multispectral pedestrian detection based on attention mechanism according to claims 7 and 8, characterized in that: The interactive fusion of the enhanced visible light feature and the enhanced infrared feature to obtain the fused visible light feature and the fused infrared feature includes: Use the enhanced visible light feature as the query and the enhanced infrared feature as the key to calculate the visible light attention weight; Use the enhanced infrared feature as query and the enhanced visible light feature as key to calculate the infrared attention weight; Perform weighted summation of the values ​​according to the visible light attention weights to obtain updated visible light features; Perform weighted summation of the values ​​according to the infrared attention weights to obtain updated infrared features; The updated visible light features and the enhanced visible light features are connected through residual connection to obtain the fused visible light features; The updated infrared features and enhanced infrared features are connected via residual connections to obtain fused features.

10. The method of adaptive multispectral pedestrian detection based on attention mechanism according to claim 1, characterized in that: The training process of the adaptive multi-spectral pedestrian detection model includes: In the first stage, the ability of the adaptive multi-spectral pedestrian detection model to detect pedestrians is optimized with the goal of minimizing the losses of the first backbone network and the second backbone network; In the second stage, the adaptive multispectral pedestrian detection model is trained to distinguish between visible light image and infrared image detection scenes; When the adaptive multi-spectral pedestrian detection model training is in the second stage, the training parameters of the first stage are not updated.