Target detection device fusing infrared and visible light information

Through an object detection device that fuses infrared and visible light information, global features are extracted from visible light and infrared images and combined light perception is performed, the problem of limited detection performance caused by changes in light conditions in the prior art is solved, and more accurate and robust object detection is achieved.

CN120298666AInactive Publication Date: 2025-07-11BEIJING TOPMOO TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510369056.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the face of changing lighting conditions, the prior art has failed to fully utilize the supplementary information provided by infrared images when reliantlying on visible light images, resulting in limited object detection performance in complex lighting scenarios.

Method used

The object detection device that fuses infrared and visible light information extracts global features from visible light and infrared images through the image global feature extraction module, combines the light information extraction module for joint light perception, uses visible light-infrared light feature interactive encoding to determine the light weight, and finally performs object detection.

Benefits of technology

It improves the accuracy and robustness of target detection, can better identify and locate targets under complex lighting conditions, and enhances the adaptability and reliability of the detector.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298666A_ABST
    Figure CN120298666A_ABST
Patent Text Reader

Abstract

The invention provides a target detection device fusing infrared and visible light information, and relates to the field of target detection, which comprises the following steps: firstly, extracting respective global features from a visible light image and an infrared image, and then carrying out illumination joint perception processing on the visible light image and the infrared image to extract illumination information; and finally, target detection is carried out based on the extracted global features and illumination information. Therefore, more accurate target detection can be realized by fully utilizing the complementary characteristic between the visible light image and the infrared image and the specific illumination condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of object detection, and more specifically, to an object detection device that fuses infrared and visible light information. Background Art

[0002] With the progress of technology, object detection technology plays a crucial role in fields such as autonomous driving, security monitoring, and industrial inspection. Visible light images can accurately identify target details with rich texture and color information under good lighting conditions. Infrared images, relying on thermal radiation characteristics, can still stably present the target contour and thermal distribution in harsh environments such as darkness, strong light interference, or occlusion, ensuring the reliability and stability of detection.

[0003] In this regard, the existing patent CN114882328B proposes an object detection method that combines visible light images and infrared images. First, it obtains the global features of visible light images and infrared images. Then, it only performs light perception on visible light images to extract light weights. Next, it uses a mutual information module to reduce the redundant information between the two image features and optimize the feature representation. Subsequently, it fuses the complementary information of these two images to generate a fused feature map. Finally, it combines the light information and the fused feature map and inputs them into a detector to complete the object detection task.

[0004] This patent only relies on visible light images to extract light information to determine light weights. However, in real-world scenarios, lighting conditions change constantly. For example, in the field of security monitoring, direct strong sunlight during the day may cause image overexposure and loss of key details; the low-light environment at night makes the image dim and blurry, making it difficult to identify targets. In the autonomous driving scenario, vehicles frequently encounter different lighting conditions during driving, such as strong light stimulation when driving out of a tunnel and low-light conditions at night. If completely relying on visible light images to extract light information and ignoring the supplementary information provided by infrared images, it will lead to incomplete light information. Moreover, different types of lighting changes (such as strong backlighting, low light, night lighting, etc.) may require different processing strategies. Using only light weights based on visible light images may not be able to effectively distinguish these complex lighting conditions.

[0005] Therefore, an optimized object detection scheme that fuses infrared and visible light information is desired. Summary of the Invention

[0006] To solve the above technical problems, this application is proposed. Embodiments of this application provide an object detection device that fuses infrared and visible light information.

[0007] According to one aspect of this application, an object detection device that fuses infrared and visible light information is provided, which includes:

[0008] An image global feature extraction module for extracting global features from visible light images and infrared images;

[0009] A lighting information extraction module for jointly perceiving the lighting of the visible light image and the infrared image to extract lighting information, wherein the lighting information extraction module includes: a lighting feature extraction unit for respectively extracting lighting features from the visible light image and the infrared image to obtain visible light lighting features and infrared lighting features; a visible-infrared light feature interaction unit for performing visible-infrared light implicit feature association interaction coding on the visible light lighting features and the infrared lighting features to obtain visible-infrared light feature interaction coding features; a lighting weight determination unit for determining a lighting weight based on the visible-infrared light feature interaction coding features;

[0010] A target detection module for performing target detection based on the global features and lighting information of the visible light image and the infrared image.

[0011] Compared with the prior art, the target detection device for fusing infrared and visible light information provided by the present application first extracts respective global features from the visible light image and the infrared image, then performs joint lighting perception processing on the visible light image and the infrared image to extract lighting information, and finally performs target detection based on the extracted global features and lighting information. In this way, by making full use of the complementary characteristics between the visible light image and the infrared image and the specific lighting conditions, it is beneficial to achieve more accurate target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0013] Figure 1 FIG. is a block diagram of a target detection device for fusing infrared and visible light information according to an embodiment of the present application.

[0014] Figure 2 FIG. is a block diagram of a lighting information extraction module in a target detection device for fusing infrared and visible light information according to an embodiment of the present application.

[0015] Figure 3 FIG. is a block diagram of a visible light lighting feature extraction unit in a target detection device for fusing infrared and visible light information according to an embodiment of the present application.

[0016] Figure 4Block diagram of the visible-infrared feature interaction unit in the target detection device that fuses infrared and visible light information according to an embodiment of the present application.

[0017] Figure 5 Block diagram of the illumination weight determination unit in the target detection device that fuses infrared and visible light information according to an embodiment of the present application. Detailed implementation manners

[0018] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0019] With the progress of technology, target detection technology plays an important role in fields such as autonomous driving, security monitoring, and industrial inspection. Visible light images can accurately identify target details under good lighting conditions, while infrared images can stably present the target contour by virtue of thermal radiation characteristics in harsh environments such as darkness, strong light interference, or occlusion, ensuring the reliability of detection. The existing patent CN114882328B proposes a target detection method that combines visible light images and infrared images. First, it obtains the global features of visible light images and infrared images, then performs illumination perception on the visible light images to extract illumination weights, then uses a mutual information module to reduce the redundant information between the features of the two images to optimize the feature representation, and finally uses the complementary information of the two images to generate a fused feature map and illumination information to complete target detection. However, this method has limitations in the face of changing lighting conditions. For example, in security monitoring, strong light during the day may cause the image to be overexposed, and low light at night will make the image dim and difficult to recognize; in autonomous driving, the vehicle will experience various lighting changes, from the moment of strong light when coming out of a tunnel to the weak light environment at night. Since this patent mainly relies on visible light images to extract illumination information and does not fully utilize the supplementary information provided by infrared images, it is difficult for it to comprehensively handle different lighting scenarios, thereby restricting the improvement of detection performance.

[0020] Based on this, the present application provides a target detection device that fuses infrared and visible light information. Figure 1 System block diagram of the target detection device that fuses infrared and visible light information according to an embodiment of the present application. As Figure 1 shown, in the target detection device 100 that fuses infrared and visible light information, it includes: an image global feature extraction module 110, which is used to extract global features from visible light images and infrared images; an illumination information extraction module 120, which is used to perform joint illumination perception on the visible light images and the infrared images to extract illumination information; a target detection module 130, which is used to perform target detection based on the global features and illumination information of the visible light images and the infrared images.

[0021] In an embodiment of the present application, the image global feature extraction module 110 is used to extract global features from visible light images and infrared images. It should be understood that considering the characteristics of the information contained in visible light images and infrared images. Visible light images can present rich texture details. Under good lighting conditions, information such as the surface material and color difference of objects can be clearly seen; infrared images are not restricted by lighting conditions and can highlight the contours and thermal radiation characteristics of target objects. For example, in a dark environment, the infrared characteristics of targets such as human bodies or heating devices are obvious. Extracting global features can comprehensively capture various information of these different modality images, avoid omitting important content caused by local information extraction, and provide a more complete data basis for subsequent processing.

[0022] In a specific example of the present application, an implementable way to extract global features from visible light images and infrared images is as follows:

[0023] First is the preprocessing link of the images. For visible light images and infrared images, normalization is an essential step. Since the pixel values of the two types of images have different scales and distribution ranges, the normalization operation can adjust them to a unified range, usually mapping the pixel values to the interval [0, 1]. This can not only eliminate the influence brought by the pixel value scale difference between different images, but also provide a stable and comparable data environment for the subsequent training of the neural network, making it easier for the network to learn effective features in the images. At the same time, in order to meet the input requirements of the backbone network based on the SSD algorithm, it is necessary to uniformly scale the visible light images and infrared images to a fixed size. This process generally uses mature image processing algorithms such as bilinear interpolation. By resampling and calculating the pixels of the image, while ensuring the integrity of the image content, the length and width of all input images are made to be the same, such as the common 300×300 pixels or 512×512 pixels, etc. The unified size facilitates the network to perform efficient batch processing, greatly improving the efficiency of feature extraction.

[0024] After preprocessing, feature extraction is carried out with the backbone network based on the SSD algorithm. In practical applications, network structures such as VGG-16 and ResNet are often selected as the backbone network of the SSD algorithm. Here, VGG-16 is taken as an example for illustration. VGG-16 is composed of multiple convolutional layers and pooling layers stacked in an orderly manner, and has powerful feature extraction capabilities. The preprocessed visible light image and infrared image are respectively input into the first convolutional layer of VGG-16. The convolutional kernels in the convolutional layer will slide pixel by pixel on the image for convolution operations. For the visible light image, the convolutional kernel can capture local features such as edges and textures in the image. For example, a 3×3 convolutional kernel can keenly detect tiny edge information in the image, such as the contour line of an object or the detailed changes in texture. For the infrared image, the convolutional kernel focuses on extracting features related to heat distribution, such as the contour of a thermal target and hot spot areas, which reflect the thermal radiation situation of the object.

[0025] As the convolutional layers are continuously stacked, the features of the image are gradually extracted and abstracted. In this process, the information contained in the feature maps output by the convolutional layer also gradually changes from the initial simple local features to more advanced and semantic information. After the convolutional layer, a pooling layer is usually connected, and the common one is the max pooling layer. Taking a 2×2 max pooling layer as an example, it will select the maximum value in each 2×2 pixel area as the output. In this way, the size of the feature map will be reduced by half. This operation can not only effectively reduce the dimension of the feature map, reduce the subsequent calculation amount, but also enhance the translational invariance of the features, so that the change of the position of the features in the image will not have too much impact on their extraction and recognition.

[0026] It should be noted that an important feature of the SSD algorithm is that object detection needs to be carried out on feature maps of multiple different scales. Therefore, in the VGG-16 network, feature maps of different levels are selected for subsequent processing. Among them, the shallow feature maps have obvious advantages for detecting small targets because they retain a high spatial resolution; while the deep feature maps contain more semantic information and are more conducive to detecting large targets. After a series of operations such as convolution and pooling, the feature maps of different scales obtained for the visible light image and the infrared image respectively will be combined. This process often draws on the idea of the Feature Pyramid Network to fuse feature maps of different scales and make full use of the advantages of each layer of feature maps. Finally, feature maps containing the global information of the visible light image and the infrared image are respectively generated.

[0027] In the embodiment of the present application, the illumination information extraction module 120 is configured to perform joint illumination perception on the visible light image and the infrared image to extract illumination information. Correspondingly, it is considered that the imaging quality of the visible light image is greatly affected by the illumination conditions. Under direct strong light, overexposure may occur, resulting in the loss of some details; while in a low-light environment, the image will become dim and the target is difficult to distinguish. In contrast, the imaging of the infrared image depends on the thermal radiation of the object itself and is basically not disturbed by changes in the external illumination conditions, and can stably provide information under the above-mentioned harsh illumination conditions. Performing joint illumination perception on the two can complement their limitations in obtaining illumination information and comprehensively grasp the illumination situation in the scene.

[0028] Based on this, the technical concept of the present application is to use image analysis and feature capture techniques based on machine learning to perform illumination feature extraction on the visible light image and the infrared image, so as to automatically judge the illumination condition category according to the implicit feature interaction representation between the visible light illumination feature and the infrared illumination feature, and determine the illumination weight based on this type. The present application combines the respective characteristics of the visible light image and the infrared image, not only improving the comprehensiveness and accuracy of the illumination information, but also enhancing the robustness and adaptability of the target detector under complex illumination conditions.

[0029] Specifically, Figure 2 is a block diagram of the illumination information extraction module in the target detection device that fuses infrared and visible light information according to the embodiment of the present application. As Figure 2 shown, the illumination information extraction module 120 includes: an illumination feature extraction unit 121, configured to perform illumination feature extraction on the visible light image and the infrared image respectively to obtain a visible light illumination feature and an infrared illumination feature; a visible light-infrared light feature interaction unit 122, configured to perform visible light-infrared light implicit feature association interaction coding on the visible light illumination feature and the infrared illumination feature to obtain a visible light-infrared illumination feature interaction coding feature; an illumination weight determination unit 123, configured to determine the illumination weight based on the visible light-infrared illumination feature interaction coding feature.

[0030] In the embodiment of the present application, the illumination feature extraction unit 121 is configured to perform illumination feature extraction on the visible light image and the infrared image respectively to obtain a visible light illumination feature and an infrared illumination feature. Specifically, Figure 3 is a block diagram of the visible light illumination feature extraction unit in the target detection device that fuses infrared and visible light information according to the embodiment of the present application. As Figure 3As shown, the illumination feature extraction unit 121 includes: a visible light illumination feature extraction subunit 1211, used to perform visible light illumination feature extraction on the visible light image to obtain a visible light illumination feature vector as the visible light illumination feature; an infrared light illumination feature extraction subunit 1212, used to perform infrared illumination feature extraction on the infrared image to obtain an infrared illumination feature vector as the infrared illumination feature.

[0031] In an embodiment of the present application, the visible light illumination feature extraction subunit 1211 is used to extract visible light illumination features from the visible light image to obtain a visible light illumination feature vector as the visible light illumination feature. Specifically, in an embodiment of the present application, the visible light illumination feature extraction subunit is used to: pass the visible light image through a visible light illumination feature extractor based on the ViT model to obtain the visible light illumination feature vector. Accordingly, considering that the visible light image contains rich information, not only the texture and color of the object, but also the illumination information exists in various forms. Therefore, in order to better characterize and extract the illumination information in the visible light, the present application passes the visible light image through a visible light illumination feature extractor based on the ViT model to obtain a visible light illumination feature vector. It should be understood that the ViT model, as a visual model based on the Transformer architecture, has unique advantages. It can effectively capture long-range dependencies in images, unlike traditional convolutional neural networks (CNNs) which are limited to local receptive fields. Specifically, when extracting visible light illumination features, the illumination conditions often have complex correlations in different regions of the image. The ViT model can globally analyze the image and integrate the illumination information of different regions. For example, in a scene image containing multiple light sources and shadows, the ViT model can better understand the mutual influence between illumination in different regions, fully capture the intrinsic connection of global illumination features, and dig out deeper and more subtle illumination features, which are of great value for accurately judging illumination conditions and subsequent target detection.

[0032] In a specific example of the present application, the implementation method of passing the visible light image through a visible light illumination feature extractor based on the ViT model to obtain the visible light illumination feature vector may be:

[0033] Before the visible light image is input into the visible light illumination feature extractor based on the ViT model, image preprocessing is required. The first is resizing. Since the ViT model has specific requirements for the input image size, the input visible light image should be uniformly adjusted to a suitable fixed size, such as 224×224 pixels. This can ensure the stability and efficiency of the model when processing images and avoid calculation confusion caused by image size differences. Then normalization is performed to normalize the image pixel values ​​to the range of [0,1 or [-1,1]. Usually, each pixel value is divided by 255 to convert the pixel value from [0,255] to [0,1]. The normalization operation can speed up the model training speed, stabilize the gradient, and improve the model convergence effect and accuracy. After completing the above processing, the normalized image is divided into multiple small image blocks of 16×16 pixels. Taking a 224×224 pixel image as an example, 14×14=196 image blocks will be obtained, which will become the basic units for subsequent feature extraction.

[0034] The preprocessed image blocks then enter the embedding layer processing stage. Each image block first passes through the linear projection layer. This fully connected layer converts the pixel information of the image block into a vector of fixed length. For example, a 16×16×3=768-dimensional image block may become a 768-dimensional vector after projection. In order for the model to capture the position information of the image block, it is also necessary to add a position encoding, which is a fixed vector with the same dimension as the image block vector. It is usually generated by sine and cosine functions and added to the image block vector to help the model understand the spatial structure of the image.

[0035] The vector sequence processed by the embedding layer enters the Transformer encoder for deep feature extraction. Each layer of the encoder contains a multi-head self-attention mechanism and a feedforward neural network. The multi-head self-attention mechanism divides the input vector into multiple heads to independently calculate the attention scores. These scores reflect the correlation between vectors. The context representation of each vector is obtained by weighted summation, which is extremely critical in illumination feature extraction and can capture the long-distance illumination dependency between image blocks. After that, the vector sequence is further transformed and features are extracted through a feedforward neural network composed of two fully connected layers and a ReLU activation function to enhance the model's expressiveness. At the same time, each layer of the Transformer encoder uses residual connections and layer normalization to alleviate the gradient vanishing problem and ensure stable training. Multiple Transformer encoder layers are stacked together, enabling the model to learn illumination features at different levels, from simple to complex, from concrete to abstract, and gradually mine more valuable illumination information.

[0036] After being processed by the Transformer encoder, it enters the feature vector generation stage. The special classification token added when inputting the vector sequence integrates the global information of the entire image after being processed by multiple layers of encoders. Extract the vector corresponding to this classification token from the final output, which is the illumination feature vector of the entire visible light image. This vector contains key information such as illumination intensity, direction, and distribution.

[0037] In the embodiment of the present application, the infrared light illumination feature extraction sub-unit 1212 is used to extract the infrared light illumination feature from the infrared image to obtain the infrared light illumination feature vector as the infrared light illumination feature. Specifically, in the embodiment of the present application, the infrared light illumination feature extraction sub-unit is used to: pass the infrared image through an infrared light illumination feature extractor based on a micro convolutional neural network to obtain the infrared light illumination feature vector. It should be understood that the infrared image mainly reflects the thermal radiation characteristics of an object and is basically not affected by changes in the external illumination conditions. However, in some cases (such as high-contrast scenes), it may lack detailed information. Based on this, in the technical solution of the present application, the infrared image is passed through an infrared light illumination feature extractor based on a micro convolutional neural network to obtain the infrared light illumination feature vector. That is, the micro convolutional neural network has unique advantages in processing such images. The convolutional layer of the convolutional neural network can perform convolutional operations by sliding the convolutional kernel on the image to automatically extract local features in the image. For the infrared image, by designing appropriate convolutional kernel parameters and network structure, the kernel specifically highlights the features related to thermal radiation. This local feature extraction method can effectively capture the local change information of the thermal radiation of different objects, such as the thermal radiation gradient change at the edge of an object. Moreover, the structure of the micro convolutional neural network is relatively simple and the amount of calculation is small. While ensuring the extraction of effective features, it can quickly process the infrared image and meet the requirements for processing speed in practical applications. In this way, the obtained infrared light illumination feature vector provides important supplementary information for judging the overall illumination conditions.

[0038] In the embodiment of the present application, the visible-infrared light feature interaction unit 122 is used to perform visible-infrared light implicit feature association interaction encoding on the visible light illumination feature and the infrared light illumination feature to obtain the visible-infrared light illumination feature interaction encoding feature. Specifically, Figure 4 It is a block diagram of the visible-infrared light feature interaction unit in the target detection device that fuses infrared and visible light information according to the embodiment of the present application. As Figure 4As shown, the visible light-infrared light feature interaction unit 122 includes: an infrared light implicit feature extraction subunit 1221, which is used to perform principal component analysis and implicit feature extraction on the infrared light feature vector to obtain a set of infrared light feature principal component implicit coding vectors; a visible light illumination implicit feature extraction subunit 1222, which is used to perform implicit feature extraction on the visible light illumination feature vector to obtain a visible light illumination feature implicit coding vector; an illumination feature implicit coding anchoring subunit 1223, which is used to perform implicit feature extraction on the visible light illumination feature vector to obtain a visible light illumination feature implicit coding vector; The set of vectors is anchored with implicit key clues to obtain the feature pair of {visible light illumination feature implicit coding vector, anchored infrared illumination feature principal component implicit coding vector}; the visible light-infrared illumination feature fine-grained interaction subunit 1224 is used to perform visible light-infrared feature fine-grained interaction on the visible light illumination feature vector and the infrared illumination feature vector based on the feature pair of {visible light illumination feature implicit coding vector, anchored infrared illumination feature principal component implicit coding vector} to obtain the visible light-infrared illumination feature interaction coding vector as the visible light-infrared illumination feature interaction coding feature.

[0039] It should be understood that the visible light illumination feature mainly reflects the reflection characteristics of the object to visible light, and contains rich color, texture and detail information, which is very critical for the identification and positioning of the target when the light is sufficient. The infrared illumination feature is based on the thermal radiation of the object, is not affected by the external visible light illumination conditions, and can stably present the contour and thermal distribution characteristics of the target in complex environments such as darkness, strong light interference or occlusion. Based on this, in order to make full use of their complementarity, make up for the shortcomings of the single modal illumination feature, and fully obtain the illumination information of the scene, the present application obtains the visible light-infrared illumination feature interactive coding feature by performing visible light-infrared light implicit feature association interaction coding on the visible light illumination feature and the infrared illumination feature. That is, the visible light and infrared illumination features are interactively coded by implicit feature association, which can mine potential key clues. In a complex environment, it can break through the limitations of surface feature matching, and understand the relationship between the two from a deeper level based on potential semantic association and structural consistency, thereby significantly improving the accuracy of the matching of the two illumination features, making the determined corresponding relationship more reliable, and optimizing the subsequent illumination condition judgment and target detection.

[0040] Specifically, in the embodiment of the present application, the infrared illumination implicit feature extraction subunit is used to: perform feature principal component analysis on the infrared illumination feature vector to obtain a set of infrared illumination feature principal component encoding vectors. This process can be expressed by the formula:

[0041]

[0042] Among them, V2 is the infrared illumination feature vector, PCA(·) is the principal component analysis operation of features, C2 is the infrared illumination feature sample covariance matrix calculated from V2, U2 is the infrared illumination feature principal component orthogonal matrix, v 21 , v 22 , v 2i and v 2m are respectively the 1st, 2nd, ith, and mth infrared illumination feature principal component coding vectors in the set of infrared illumination feature principal component coding vectors, Λ2 is the infrared illumination feature diagonal matrix, λ 21 , λ 2m are respectively the eigenvalues corresponding to v 21 and v 2m , and U2 T is the transpose matrix of U2;

[0043] Performing dot convolution implicit feature extraction on each infrared illumination feature principal component coding vector in the set of infrared illumination feature principal component coding vectors to obtain the set of infrared illumination feature principal component implicit coding vectors, this process can be expressed by the formula:

[0044]

[0045] Among them, Conv 1×1 is dot convolution coding, sigmoid is the sigmoid function, h 21 , h 22 , h 2i and h 2m are respectively the 1st, 2nd, ith, and mth infrared illumination feature principal component implicit coding vectors in the set of infrared illumination feature principal component implicit coding vectors, and H2 is the set of infrared illumination feature principal component implicit coding vectors.

[0046] Correspondingly, considering that the infrared illumination feature vector usually contains a large amount of dimensional information, this information not only increases the burden of data storage and processing, but may also introduce redundancy and noise, affecting the subsequent analysis efficiency and accuracy. Principal component analysis (PCA) of features transforms the original infrared illumination feature vector by constructing a new low-dimensional feature space. In this process, it sorts the features according to the variance of the data, concentrating the main information on a few principal components. For example, the original infrared illumination feature vector is a high-dimensional vector, possibly having dozens or even hundreds of dimensions. After PCA, it can be compressed into a low-dimensional vector with only a few principal components, greatly reducing the data dimension and facilitating data storage and processing. That is, the generated set of infrared illumination feature principal component implicit coding vectors has a low dimension and is rich in key information, ensuring the reliability of the data analysis results while maintaining efficient data processing.

[0047] It should be understood that although the set of principal component encoding vectors of the infrared light illumination features has concentrated key information through principal component analysis, there is still potential value to be explored. Point convolution, with its unique operation method, can effectively capture the local detailed information of each vector in the set. That is, by performing point convolution implicit feature extraction operations on each infrared light illumination principal component encoding vector in the set of infrared light illumination principal component encoding vectors, the subtle differences of the original infrared light principal components can be strengthened, thereby enhancing the feature representation ability. Moreover, there may be redundant information in the original set of infrared light illumination principal component encoding vectors, which will increase the computational amount and interfere with the model judgment. The point convolution implicit feature extraction operation screens and integrates features through carefully designed convolution kernel parameters. It can highlight the key features closely related to the target and illumination conditions and suppress or remove the redundant parts. This makes the generated set of infrared light illumination principal component implicit encoding vectors more focused on the core thermal features of the target, improving the quality and effectiveness of the features and enabling more efficient and accurate subsequent processing.

[0048] Next, perform implicit feature extraction on the visible light illumination feature vector to obtain a visible light illumination feature implicit encoding vector. The above process can be expressed by the formula:

[0049] H1 = sigmoid[Conv 1×1 (V1)]

[0050] where Conv 1×1 is point convolution encoding, sigmoid is the sigmoid function, V1 is the visible light illumination feature vector, and H1 is the visible light illumination feature implicit encoding vector.

[0051] It should be understood that the visible light image itself contains rich information such as color, texture, and shape, but the illumination feature vector directly analyzed from it may only cover the surface and more obvious features. By performing implicit feature extraction on the visible light illumination feature vector, the hidden and imperceptible feature information in the original vector can be deeply explored, and this feature information can provide a more valuable basis for subsequent illumination condition classification. In addition, the implicit feature extraction process will also perform a certain degree of abstraction and generalization on the original visible light illumination features. By removing some unimportant detail information and focusing on the representative feature patterns, the extracted visible light illumination feature implicit encoding vector has stronger generalization ability. This means that the model can better adapt to various complex and changeable actual environments, reduce the dependence on specific training data, and improve the generality and stability of the model.

[0052] Then, perform implicit key clue anchoring on the set of the implicit encoding vectors of the visible light illumination features and the implicit encoding vectors of the principal components of the infrared light illumination features to obtain a feature pair {implicit encoding vector of visible light illumination features, anchored implicit encoding vector of the principal components of infrared light illumination features}. The above process can be expressed by the formula:

[0053]

[0054] D bes tpa i r ={H1; h 2k}

[0055] where h 2i is the i-th implicit encoding vector of the principal components of the infrared light illumination features in the set of the implicit encoding vectors of the principal components of the infrared light illumination features, H1 is the implicit encoding vector of the visible light illumination features, <·> represents the inner product, ‖·‖ is the Euclidean norm for calculating vectors, ε is the modulation coefficient, argmax j (·) returns the j value corresponding to the maximum value, k is the position to find the maximum approximate matching value in the set of the implicit encoding vectors of the principal components of the infrared light illumination features, h 2k is the anchored implicit encoding vector of the principal components of the infrared light illumination features, and F bestpair is the feature pair {implicit encoding vector of visible light illumination features, anchored implicit encoding vector of the principal components of infrared light illumination features}.

[0056] It should be understood that in actual object detection tasks, visible light images present the appearance details of objects with rich color and texture information, while infrared images show the thermal distribution of objects based on their thermal radiation characteristics. The sets of latent encoding vectors of visible light illumination features and the principal component latent encoding vectors of infrared illumination features obtained through different image analyses describe objects from different physical characteristics, and the feature information they contain has different semantics. The latent key clue anchoring operation can build a bridge between these two feature spaces with different semantics, enabling these features to be associated and interact in a consistent manner. And considering that among numerous visible light and infrared features, not all features are equally important for object detection and illumination condition classification. By calculating the correlation between the set of latent encoding vectors of visible light illumination features and the set of principal component latent encoding vectors of infrared illumination features, latent key clue anchoring can screen out the most valuable feature pairs from a large number of feature combinations. It is worth mentioning that due to the different imaging principles of visible light and infrared images, there is often information ambiguity in the features they contain during fusion, which poses a great challenge to object detection. Through alignment operations, latent key clue anchoring gives a clear structural prior to subsequent feature interaction operations, eliminates the ambiguity between feature sources, enables the model to clearly identify the features of different objects, and thus improves the accuracy and stability of object detection.

[0057] Finally, based on the {latent encoding vector of visible light illumination features, anchored principal component latent encoding vector of infrared illumination features} feature pair, perform fine-grained interaction between the visible light illumination feature vector and the infrared illumination feature vector to obtain a visible light-infrared light feature interaction encoding vector as the visible light-infrared light feature interaction encoding feature. The above process can be expressed by the formula:

[0058]

[0059] where V2 is the infrared illumination feature vector, V1 is the visible light illumination feature vector, H1 is the latent encoding vector of visible light illumination features, h 2k is the anchored principal component latent encoding vector of infrared illumination features, h 2k T is the transposed vector of h 2k softmax is the activation function, S is the length of h 2k , is matrix multiplication, α and β are weighted hyperparameters, and V f is the visible light-infrared light feature interaction encoding vector.

[0060] It should be understood that based on the feature pair of {latent encoding vector of visible light illumination feature, anchored principal component latent encoding vector of infrared light illumination feature}, for the visible light illumination feature vector and the infrared light illumination feature vector, fine-grained interaction between visible light and infrared light features can be carried out. The multi-head self-attention mechanism can be used to deeply explore the complex non-linear relationship between visible light and infrared light illumination features, so as to enrich the expression of visible light and infrared light interaction features. Specifically, the multi-head self-attention mechanism can capture long-range dependencies and multi-dimensional semantic interactions, giving higher flexibility to feature fusion. For example, in a complex traffic scene, for a vehicle in the distance, the key visible light features such as its license plate number are far from the infrared feature of engine heat radiation in the image, but there is an important association between them, representing different attributes of the same target. The multi-head self-attention mechanism can simultaneously pay attention to these long-range features and dynamically adjust the attention to different features according to different task requirements. It can also handle multi-dimensional semantic interactions, such as the complex interaction relationship between the semantics of visible light dimensions such as the color and shape of the vehicle and the semantics of infrared dimensions such as the heat radiation intensity and distribution. In this way, feature fusion is no longer limited to a fixed mode, but can flexibly combine and optimize features of different modalities according to the specific scenario, thus improving the fusion effect. In particular, as an important supplement, the multi-stream architecture decomposes the granularity difference of feature interaction through parallel streams and dynamically adjusts the interaction weights to adapt to the context information, further optimizing the feature interaction process. This can more accurately capture the relationship between features in different scenarios, ensure the effective fusion of visible light and infrared features in various situations, and improve the quality and adaptability of feature fusion. Generally speaking, the visible light-infrared light illumination feature interaction encoding vector obtained through fine-grained interaction integrates the advantages of visible light and infrared images, contains rich and highly discriminative information, and significantly improves the accuracy of illumination condition judgment and target detection.

[0061] Preferably, in another embodiment of the present application, the visible light-infrared light illumination feature fine-grained interaction sub-unit is used for: performing prior correction semantic alignment on the eigenvalues of the feature pair of {latent encoding vector of visible light illumination feature, anchored principal component latent encoding vector of infrared light illumination feature} to obtain the corrected feature pair of {latent encoding vector of visible light illumination feature, anchored principal component latent encoding vector of infrared light illumination feature}; performing fine-grained interaction between visible light and infrared light features on the corrected feature pair of {latent encoding vector of visible light illumination feature, anchored principal component latent encoding vector of infrared light illumination feature} to obtain the visible light-infrared light illumination feature interaction encoding vector. The above process is expressed by the formula:

[0062]

[0063] where H 1i is the eigenvalue at the i-th position in H1, h2ki is h 2k is the eigenvalue at the i-th position in, H 1i ′ is H 1i the corrected eigenvalue, h 2ki ′ is h 2ki the corrected eigenvalue, H1′ is the implicit encoding vector of the corrected visible light illumination feature, h 2k ′ is the principal component implicit encoding vector of the corrected anchored infrared illumination feature, V f is the visible-infrared illumination feature interaction encoding vector.

[0064] That is, for the alignment ambiguity between the visible light illumination feature and the anchored infrared illumination feature caused by the source uncertainty among high-dimensional features from different sources, a weakening and blurring mechanism is introduced for weak blurring power-law expansion. That is, the 3 / 8 exponent is used as the precursor prior expansion to perform power-law prior responsive fuzzification convergence on the feature parameter boundary alignment condition, and the 1 / 4 exponent is used as the main body alignment expansion to strictly constrain the relaxation of the alignment attenuation of the feature power-law distribution. Thus, in the case where the alignment boundary condition constraint within the correlation effective range is poorly defined, the value correlation mechanism is used to avoid the prior ambiguity of the system behavior under a single mechanism, so as to correct the semantic distribution consistency fuzzification of the illumination feature within the alignment interval and enhance the intuitive mining of the implicit fine-grained interaction correlation of the principal component implicit encoding vector of the anchored infrared illumination feature.

[0065] In the embodiment of the present application, the illumination weight determination unit 123 is used to determine the illumination weight based on the visible-infrared illumination feature interaction encoding feature. Specifically, Figure 5 is the block diagram of the illumination weight determination unit in the target detection device that fuses infrared and visible light information according to the embodiment of the present application. As Figure 5 shown, the illumination weight determination unit 123 includes: an illumination condition category determination subunit 1231, which is used to obtain an illumination condition classification result based on the visible-infrared illumination feature interaction encoding vector, and the illumination condition classification result is used to represent the illumination condition category label; an illumination weight generation subunit 1232, which is used to determine the illumination weight based on the illumination condition classification result.

[0066] In the embodiment of the present application, the illumination condition category determination subunit 1231 is configured to obtain an illumination condition classification result based on the visible-infrared illumination feature interaction coding vector, and the illumination condition classification result is used to represent an illumination condition category label. Specifically, in the embodiment of the present application, the illumination condition category determination subunit is configured to: pass the visible-infrared illumination feature interaction coding vector through an illumination condition discriminator based on a classifier to obtain the illumination condition classification result, and the illumination condition classification result is used to represent an illumination condition category label. That is, classification processing is performed on the visible-infrared illumination feature interaction coding feature obtained by performing implicit feature association interaction on the visible light illumination feature vector and the infrared light illumination feature vector, so as to automatically determine the illumination condition category. It should be understood that the classifier has a powerful data classification ability, can process and analyze these complex coding vectors, can learn the distribution rules and feature patterns of the feature vectors under different illumination conditions, and convert them into illumination condition category labels that are easy to understand and use. The illumination condition category labels here can be "strong light", "weak light", "no light", etc.

[0067] In the embodiment of the present application, the illumination weight generation subunit 1232 is configured to determine an illumination weight based on the illumination condition classification result. It should be understood that under different illumination conditions, the effective information contained in the visible light image and the infrared image is different. When directly irradiated by strong light, the visible light image is prone to overexposure and some details are lost, but the overall scene information is rich; while the infrared image is not interfered by strong light at this time and can stably present the thermal radiation characteristics of the object. In a low-light environment, the visible light image is dim and the target is difficult to distinguish, and the advantage of the thermal radiation characteristics of the infrared image is more prominent. Therefore, determining the illumination weight according to the illumination condition classification result can reasonably allocate the information proportion of the two images in target detection according to different situations. In particular, the illumination weight in the present application refers to the proportion of the visible light modality and the infrared modality in target detection.

[0068] In a specific example of the present application, an implementable manner for determining the illumination weight based on the illumination condition classification result is as follows:

[0069] First, utilize the powerful capabilities of deep learning models to dynamically adjust weights. Construct a weight generation network based on a convolutional neural network (CNN), and use the classification results of lighting conditions in a suitable encoding form, such as one-hot encoding, as the input to this network. The network internally contains multiple convolutional layers and fully connected layers, which can conduct in-depth feature learning and processing on the input lighting condition information. Different lighting condition categories, such as "strong light", "weak light", "no light", etc., each have unique feature patterns. During the training process, the weight generation network learns from a large number of image data under different lighting conditions and the corresponding object detection results, gradually grasping the differences in the importance of visible light images and infrared images in object detection under different lighting conditions. For example, under strong light conditions, the visible light image may experience overexposure, resulting in the loss of some details, while the infrared image is not affected by strong light and can stably present the thermal radiation characteristics of objects. At this time, the network will learn that the infrared image is more crucial for object detection, and then output a lower weight for the visible light image and a higher weight for the infrared image; under weak light conditions, through the training and analysis of a large amount of data, the network will reasonably allocate the weights of the two according to the importance and relevance of the features. This dynamic weight adjustment method based on deep learning makes full use of the powerful feature extraction and learning capabilities of convolutional neural networks, can generate weights that match the corresponding lighting conditions in real time according to different lighting conditions, and greatly improves the intelligence and accuracy of weight determination.

[0070] Secondly, further optimize the weights by combining historical data and real-time feedback. Establish a complete historical data recording system that stores image data under different lighting conditions, the corresponding object detection results, and the lighting weights used at that time. When new lighting condition classification results are generated, the system will quickly search in the historical data to find similar lighting condition cases. If similar cases are found, the system will refer to the historical weight allocation scheme and make fine-tuning in combination with the real-time feedback information of the current scene. Taking the security monitoring scenario as an example, when it is detected that the current lighting condition is "weak light", the system will obtain the detection records and weight allocations in similar weak light scenarios from the historical data. However, each scene has certain particularities. If the thermal radiation characteristics of the target in the current scene are more obvious than those in the historical case, or there are some other special factors, such as the influence of reflectors in the environment on infrared light, the system will appropriately increase the weight of the infrared image according to these real-time feedback information to better adapt to the actual situation. At the same time, the current detection results are timely fed back to the system for further optimizing the weight generation model. Over time and with the accumulation of data, the model can continuously learn and adapt to various changes, making the determination of weights more intelligent and accurate, thereby improving the performance of object detection.

[0071] In addition, multi-modal information fusion can be used to help determine weights, making the determination of weights more comprehensive and reasonable. In addition to the classification results of the lighting conditions, other multi-modal information is also comprehensively considered. For example, scene semantic information also plays an important role in object detection. Taking the parking lot scene as an example, under low-light conditions at night, the metal material of the vehicle has unique reflection characteristics for infrared light, while key information such as license plates is more important in visible light images. To make full use of this multi-modal information, natural language processing technology is used to extract key information from the scene description, such as the scene type, possible features of the object, etc., and input it into a multi-modal fusion model together with the classification results of the lighting conditions. By fusing various information such as lighting categories and scene semantics, this model can more comprehensively analyze the value of different modal images in object detection. During the processing, the model will consider the mutual relationships and influences between various information, and comprehensively judge the relative importance of visible light images and infrared images for object detection under the current lighting conditions and scene, so as to generate more reasonable lighting weights. This way of multi-modal information fusion makes full use of the complementarity of different types of information, makes the determination of weights more scientific, and helps to improve the accuracy and reliability of object detection.

[0072] In summary, the illumination information extraction module 120 is clearly described. It uses machine learning-based image analysis and feature capture techniques to extract illumination features from visible light images and infrared images, automatically determines the illumination condition category based on the implicit feature interaction representation between the visible light illumination features and the infrared light illumination features, and determines the illumination weight based on this category. In this way, by combining the respective characteristics of visible light images and infrared images, not only the comprehensiveness and accuracy of the illumination information are improved, but also the robustness and adaptability of the object detector under complex illumination conditions are enhanced.

[0073] In the embodiment of the present application, the object detection module 130 is used to perform object detection based on the global features and illumination information of the visible light image and the infrared image.

[0074] Specifically, first, mutual information is captured for the global features of visible light images and infrared images to identify and eliminate redundant information. It should be understood that there is a certain overlap in the information expression between visible light images and infrared images. For example, the approximate position information of the target object is reflected in both images. This redundant information increases the data processing volume, reduces the detection efficiency, and may also interfere with subsequent analysis. Therefore, it is necessary to discover and process this redundancy through mutual information capture. In particular, as described in the patent, by using designed convolutional layers and fully connected layers, the image features are mapped into low-dimensional vectors. Based on this, the mutual information is calculated, and the parameters of the image global feature extraction module are optimized to reduce redundancy. Specifically, first, the global features of visible light images and infrared images are input into a mutual information calculation module. Within this calculation module, through two convolutional layers and one fully connected layer, the global features of visible light images and infrared images are mapped into a global feature vector of visible light images and a global feature vector of infrared images, denoted as Vv (from visible light images) and Vi (from infrared images), respectively. Based on these two low-dimensional feature vectors, two information distributions are created: a Gaussian distribution G_d1 with Vv as the mean and Vi as the variance; similarly, another Gaussian distribution G_d2 with Vi as the mean and Vv as the variance. Next, the mutual information M(G_d1, G_d2) between these two Gaussian distributions G_d1 and G_d2 is calculated. The formula for mutual information is M(G_d1, G_d2) = H(G_d1) + H(G_d2) - H(G_d1, G_d2), where H(G_d1) and H(G_d2) represent the information entropy of the two Gaussian distributions, and H(G_d1, G_d2) represents their joint entropy. Further analyzing the relationship between information entropy, cross entropy, and relative entropy, we can obtain H(G_d1) = C(G_d2, G_d1) - D(G_d2||G_d1) and H(G_d2) = C(G_d1, G_d2) - D(G_d1||G_d2), where C(G_d1, G_d2) is the cross entropy of the two Gaussian distributions, and D(G_d1||G_d2) is the relative entropy of the two Gaussian distributions. Based on the non-negative property of the joint entropy, H(G_d1, G_d2) can be removed, thus simplifying the mutual information formula to M'(G_d1, G_d2) = C(G_d2, G_d1) - D(G_d2||G_d1) + C(G_d1, G_d2) - D(G_d1||G_d2). The mutual information M'(G_d1, G_d2) obtained in this way quantifies the amount of redundant information between the features of visible light images and infrared images. Finally, through the optimization process, the parameters of the image global feature extraction module are adjusted to reduce the redundant information between these two features, ensuring that only the complementary information beneficial to target detection is retained in the feature fusion stage, rather than including redundant or duplicate information.

[0075] Next, fuse the global features of the two images and their complementary information to obtain a fused feature map. Correspondingly, considering the different imaging principles of visible light images and infrared images, the information carried by each has its own characteristics. To obtain more comprehensive scene information and enable more reliable operation in various complex scenarios, fuse the global features of the two images and their complementary information to obtain a fused feature map. In particular, the method of adding a channel attention mechanism in the patent can be used to obtain channel weights by analyzing feature differences, and perform complementary information fusion through operations with the original features, which can enhance the representativeness of the features and improve the accuracy of object detection. Specifically, fusing the global features of the two images and their complementary information to obtain a fused feature map includes: First, subtract the visible light image feature Fv and the infrared image feature Fi element by element, and the resulting difference feature represents their complementary information. Then, input this difference feature into the channel attention mechanism to generate channel weights. Finally, perform an element-wise multiplication operation using this channel weight and the corresponding original features (i.e., the visible light image feature Fv and the infrared image feature Fi), amplify the difference feature, and add it to the feature of the other modality to finally form the fused feature maps Fv' and Fi'. This process emphasizes the regional features that contribute more to object detection and also ensures that the fused feature maps contain the key complementary information from the two images. In addition, the calculation of the channel attention mechanism involves a series of operations, including processing data related to the feature map sizes H and W using the hyperbolic tangent activation function σ to further refine the importance of each channel, so that the final two fused feature maps can more accurately reflect the information provided by the two images.

[0076] Finally, combine the illumination weight and the fused feature Figure 1It is input into the detector for object detection. That is, the illumination weight is determined according to the classification result of the illumination condition, which can assist the detector to reasonably utilize the image information under different illuminations, enabling the detector to more accurately identify and locate the object under various illumination conditions. After obtaining the fused feature maps of the visible light image and the infrared image, these fused feature maps are used as inputs together with the obtained illumination weight and passed to the final detector. Specifically, this process can be represented by a function called d_m, which receives the illumination weight and the fused feature maps as input parameters and outputs the prediction results, including the class (cls), confidence (conf), and bounding box coordinates (box) of the object. This process can be simply expressed by the formula as follows: cls, conf, box = d_m(Ws, Wl, Fi', Fv'), where d_m refers to the detector, which utilizes the illumination information (Ws and Wl) provided by the illumination information extraction module and the fused feature maps (Fi' and Fv') generated in the previous steps to perform object detection in a more robust manner. This not only takes into account the complementarity between multi-modal data but also effectively combines the influence of the illumination condition on the detection result, which can effectively improve the performance and accuracy of the object detection algorithm under different illumination conditions.

[0077] In summary, the object detection device 100 that fuses infrared and visible light information according to the embodiments of the present application is elucidated. It first extracts the respective global features from the visible light image and the infrared image, then performs joint illumination perception processing on the visible light image and the infrared image to extract the illumination information, and finally performs object detection based on the extracted global features and illumination information. In this way, by making full use of the complementary characteristics between the visible light image and the infrared image and the specific illumination situation, it is beneficial to achieve more accurate object detection.

Claims

1. An object detection device that fuses infrared and visible light information, characterized in that, Including: An image global feature extraction module, configured to extract global features from visible light images and infrared images; A lighting information extraction module, configured to perform joint lighting perception on the visible light image and the infrared image to extract lighting information, wherein the lighting information extraction module includes: a lighting feature extraction unit, configured to perform lighting feature extraction on the visible light image and the infrared image respectively to obtain visible light lighting features and infrared lighting features; a visible-infrared light feature interaction unit, configured to perform visible-infrared light implicit feature association interaction coding on the visible light lighting features and the infrared lighting features to obtain visible-infrared light feature interaction coding features; a lighting weight determination unit, configured to determine a lighting weight based on the visible-infrared light feature interaction coding features; A target detection module, configured to perform target detection based on the global features and lighting information of the visible light image and the infrared image.

2. The target detection device for fusing infrared and visible light information according to claim 1, wherein, The lighting feature extraction unit includes: A visible light lighting feature extraction subunit, configured to perform visible light lighting feature extraction on the visible light image to obtain a visible light lighting feature vector as the visible light lighting feature; An infrared light lighting feature extraction subunit, configured to perform infrared light lighting feature extraction on the infrared image to obtain an infrared light lighting feature vector as the infrared light lighting feature.

3. The target detection device for fusing infrared and visible light information according to claim 2, characterized in that, The visible light lighting feature extraction subunit is configured to: pass the visible light image through a visible light lighting feature extractor based on a ViT model to obtain the visible light lighting feature vector.

4. The target detection device for fusing infrared and visible light information according to claim 3, characterized in that, The infrared light lighting feature extraction subunit is configured to: pass the infrared image through an infrared light lighting feature extractor based on a micro convolutional neural network to obtain the infrared light lighting feature vector.

5. The target detection device for fusing infrared and visible light information according to claim 4, characterized in that, The visible-infrared light feature interaction unit includes: An infrared light implicit feature extraction subunit, configured to perform principal component analysis and implicit feature extraction on the infrared light lighting feature vector to obtain a set of infrared light lighting feature principal component implicit coding vectors; A visible light lighting implicit feature extraction subunit, configured to perform implicit feature extraction on the visible light lighting feature vector to obtain a visible light lighting feature implicit coding vector; A lighting feature implicit coding anchor subunit, configured to perform implicit key clue anchoring on the visible light lighting feature implicit coding vector and the set of infrared light lighting feature principal component implicit coding vectors to obtain a {visible light lighting feature implicit coding vector, anchored infrared light lighting feature principal component implicit coding vector} feature pair; A visible-infrared light feature fine-grained interaction subunit, configured to perform visible-infrared light feature fine-grained interaction on the visible light lighting feature vector and the infrared light lighting feature vector based on the {visible light lighting feature implicit coding vector, anchored infrared light lighting feature principal component implicit coding vector} feature pair to obtain a visible-infrared light feature interaction coding vector as the visible-infrared light feature interaction coding feature.

6. The target detection device for fusing infrared and visible light information according to claim 5, characterized in that The infrared light implicit feature extraction subunit is configured to: Perform feature principal component analysis on the infrared light lighting feature vector to obtain a set of infrared light lighting feature principal component coding vectors; Performing dot convolution implicit feature extraction on each infrared illumination feature principal component coding vector in the set of the infrared illumination feature principal component coding vectors to obtain the set of the infrared illumination feature principal component implicit coding vectors.

7. The target detection device for fusing infrared and visible light information according to claim 6, wherein, The visible-infrared illumination feature fine-grained interaction subunit is configured to: Performing prior correction semantic alignment of eigenvalues on the {visible light illumination feature implicit coding vector, anchored infrared illumination feature principal component implicit coding vector} feature pair to obtain the corrected {visible light illumination feature implicit coding vector, anchored infrared illumination feature principal component implicit coding vector} feature pair; Based on the corrected {visible light illumination feature implicit coding vector, anchored infrared illumination feature principal component implicit coding vector} feature pair, performing visible-infrared light feature fine-grained interaction on the visible light illumination feature vector and the infrared illumination feature vector to obtain the visible-infrared illumination feature interaction coding vector.

8. The target detection device for fusing infrared and visible light information according to claim 7, characterized in that, The illumination weight determination unit includes: An illumination condition category determination subunit, configured to obtain an illumination condition classification result based on the visible-infrared illumination feature interaction coding vector, where the illumination condition classification result is used to represent an illumination condition category label; An illumination weight generation subunit, configured to determine an illumination weight based on the illumination condition classification result.

9. The target detection device for fusing infrared and visible light information according to claim 8, characterized in that, The illumination condition category determination subunit is configured to: pass the visible-infrared illumination feature interaction coding vector through an illumination condition discriminator based on a classifier to obtain the illumination condition classification result, where the illumination condition classification result is used to represent an illumination condition category label.