Multi-spectrum-based object detection method and related equipment

Through the multi-spectral object detection method, combined with feature extraction and modal fusion of infrared and visible images, the problem of poor target detection at night and inclement weather conditions is solved, all-weather object detection and positioning are achieved, and the judgment ability of the autonomous driving system is improved.

CN120220105APending Publication Date: 2025-06-27XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184603.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing object detection methods based on visible light images perform poorly at night or in severe weather conditions, resulting in reduced detection efficiency and increased error results, affecting the judgment ability of the autonomous driving system.

Method used

Multi-spectral object detection method is adopted to obtain visible light images and infrared images, feature extraction and modal fusion are performed, combined with self-attention mechanism and improved feature pyramid network, multi-scale feature fusion and classification regression are achieved, and the category and position information of objects are obtained.

Benefits of technology

In the night and inclement weather conditions, by fusing infrared and visible light images information, all-weather object detection and positioning can be achieved, the adaptability and accuracy of the object detection algorithm can be improved, and the judgment ability of the autonomous driving system can be enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220105A_ABST
    Figure CN120220105A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a multispectral-based object detection method and related equipment, and the method comprises the steps: obtaining a visible light image and an infrared image, carrying out the feature extraction of the visible light image and the infrared image, carrying out the modal fusion and multi-scale fusion of the extracted feature images, and obtaining feature enhancement images under different scales; and carrying out classification and regression on the feature enhancement images under different scales to obtain a detection result of the to-be-detected image data set. The infrared images are high in imaging penetrability under the night and severe weather conditions, the method is suitable for target detection and monitoring under the night and severe weather conditions, all-weather detection and positioning of any target under the night and severe weather conditions can be achieved by fusing the information of the infrared images and the information of the visible light images, and multi-scale information is fused, so that all-weather detection and positioning of any target under the night and severe weather conditions are achieved. The adaptability of a target detection algorithm under different scenes, illumination changes and scale changes can be improved, so that the accuracy of target detection is improved, and the ability of an automatic driving system to make correct judgment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multispectral target detection, and in particular, to an object detection method and related equipment based on multispectral.

Background Art

[0002] Target detection is one of the key technologies of an autonomous driving system, and its performance is closely related to driving safety. At present, the target detection method based on visible light images performs well in a constrained experimental environment, but has a poor application effect in a real scenario with a variable environment. For example, at night or under bad weather conditions, due to insufficient light, a single visible light image cannot provide enough reliable information, which will greatly reduce the detection efficiency of the target detection algorithm, obtain incorrect detection results, and seriously affect the ability of the autonomous driving system to make correct judgments.

Summary of the Invention

[0003] In view of this, the present invention provides an object detection method and related equipment based on multispectral.

[0004] The specific technical solution of the first embodiment of the present invention is: an object detection method based on multispectral, the method includes: obtaining a dataset of images to be detected; the dataset of images to be detected includes visible light images and infrared images; extracting features from the visible light images to obtain first feature maps of the visible light images at different feature extraction scales; extracting features from the infrared images to obtain second feature maps of the infrared images at different feature extraction scales; performing modal fusion and multi-scale fusion on all the first feature maps and all the second feature maps to obtain feature enhanced maps of the dataset of images to be detected at different scales; classifying and regressing the feature enhanced maps at different scales to obtain the detection results of the dataset of images to be detected; the detection results include the category information and position information of the objects included in the dataset of images to be detected.

[0005] Preferably, performing modal fusion and multi-scale fusion on all the first feature maps and all the second feature maps to obtain feature enhancement maps of the image dataset to be detected at different scales includes: inputting a first target feature map into a preset visible light attention branch structure for two different convolution operations to obtain a first visible light feature map and a second visible light feature map; the first target feature map is any one of all the first feature maps; inputting a second target feature map into a preset infrared attention branch structure for two different convolution operations to obtain a first infrared feature map and a second infrared feature map; the second target feature map is a feature map in all the second feature maps with the same scale as the first target feature map; obtaining a first target feature enhancement map according to the first visible light feature map, the second visible light feature map, the first infrared feature map, the second infrared feature map, the first target feature map and the second target feature map; all the first target feature enhancement maps constitute the feature enhancement maps at different scales.

[0006] Preferably, the obtaining a first target feature enhancement map according to the first visible light feature map, the second visible light feature map, the first infrared feature map, the second infrared feature map, the first target feature map and the second target feature map includes: obtaining visible light attention weights according to the first visible light feature map and the second visible light feature map; obtaining infrared attention weights according to the first infrared feature map and the second infrared feature map; obtaining a first target feature enhancement map according to the visible light attention weights, the infrared attention weights, the first target feature map and the second target feature map.

[0007] Preferably, the first target feature enhancement map is obtained by the following formula:

[0008] F = A R × F R + A T × F T

[0009] where F is the first target feature enhancement map, A R is the visible light attention weight, F R is the first target feature map, A T is the infrared attention weight, F T is the second target feature map.

[0010] Preferably, obtaining the visible light attention weight according to the first visible light feature map and the second visible light feature map includes: multiplying the first visible light feature map and the second visible light feature map in matrix; obtaining the visible light attention weight according to the visible light feature map after matrix multiplication, a preset convolutional layer, a preset layernorm layer, and a sigmoid function.

[0011] Preferably, the visible light attention weight is obtained by the following formula:

[0012] A R = sg[ln(conv3[sm(conv1(F m ))×conv2(F m )])]

[0013] where A R is the visible light attention weight, the function sg(·) is the sigmoid function, the function ln(·) is the preset layernorm layer, the function sm(·) is the softmax layer, the function conv(·) is the preset convolutional layer, the numbers 1 and 2 in the function conv(·) represent different positions of the preset convolutional layer, F m is the first target feature map, conv1(F m ) is the first visible light feature map, conv2(F m ) is the second visible light feature map.

[0014] Preferably, the detection result of the image dataset to be detected is obtained by the following formula:

[0015]

[0016] where D cls is the category information of the object in the detection result, D bbox is the position information of the object in the detection result, δ head is the detection head function, δ neck is the multi-scale feature fusion function, is the feature enhancement map at different scales, θ h is the preset detection parameter.

[0017] The specific technical solution of the second embodiment of the present invention is as follows: A multi-spectral based object detection system, the system comprising: a data acquisition module, a first feature extraction module, a second feature extraction module, a fusion module, and a detection module; the data acquisition module is used to acquire a dataset of images to be detected; the dataset of images to be detected includes visible light images and infrared images; the first feature extraction module is used to perform feature extraction on the visible light images to obtain first feature maps of the visible light images at different feature extraction scales; the second feature extraction module is used to perform feature extraction on the infrared images to obtain second feature maps of the infrared images at different feature extraction scales; the fusion module is used to perform modal fusion and multi-scale fusion on all the first feature maps and all the second feature maps to obtain feature enhanced maps of the dataset of images to be detected at different scales; the detection module is used to perform classification and regression on the feature enhanced maps at different scales to obtain the detection results of the dataset of images to be detected; the detection results include the category information and position information of the objects included in the dataset of images to be detected.

[0018] The specific technical solution of the third embodiment of the present invention is as follows: A multi-spectral based object detection device, comprising a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of the first embodiments of the present application.

[0019] The specific technical solution of the fourth embodiment of the present invention is as follows: A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of the first embodiments of the present application.

[0020] Implementing the embodiments of the present invention will have the following beneficial effects:

[0021] By acquiring visible light images and infrared images, performing feature extraction on the visible light images and infrared images, and performing modal fusion and multi-scale fusion on the extracted feature maps, the present invention obtains feature enhanced maps at different scales; performing classification and regression on the feature enhanced maps at different scales to obtain the detection results of the dataset of images to be detected. Infrared images have strong imaging penetration at night and in bad weather conditions and are suitable for target detection and monitoring at night and in bad weather conditions. By fusing the information of infrared images and visible light images, all-weather detection and positioning of any target can be achieved at night and in bad weather conditions. By fusing multi-scale information, the adaptability of the target detection algorithm under different scenarios, light changes, and scale changes can be improved, thereby improving the accuracy of target detection, and thus improving the ability of the autonomous driving system to make correct judgments.

Description of the Drawings

[0022] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0023] Figure 1 It is a flowchart of the steps of the multi-spectral object detection method in the present invention;

[0024] Figure 2 It is a structural diagram of the multi-spectral object detection network based on the self-attention mechanism fusion in the present invention;

[0025] Figure 3 It is a structural diagram of the self-attention mechanism fusion module network in the present invention;

[0026] Figure 4 It is a structural diagram of the existing feature pyramid network;

[0027] Figure 5 It is a structural diagram of the improved feature pyramid network adopted in the present invention;

[0028] Figure 6 It is a diagram showing the example of the effect on the FLIR dataset;

[0029] Figure 7 It is a diagram showing the example of the effect on the LLVIP dataset;

[0030] Figure 8 It is a schematic structural diagram of the multi-spectral object detection system;

[0031] Among them, 201, data acquisition module; 202, first feature extraction module; 203, second feature extraction module; 204, fusion module; 205, detection module.

Specific Embodiments

[0032] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0033] The terms "first", "second", etc. in the description, claims and drawings of this application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other steps or modules inherent to these processes, methods, products or devices.

[0034] Referring to "embodiments" herein means that specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0035] Please refer to Figure 1 , which is a flowchart of the steps of a multi-spectral based object detection method in the first embodiment of this application, so as to improve the accuracy of target detection. The method includes:

[0036] Step 101, obtain a dataset of images to be detected; the dataset of images to be detected includes visible light images and infrared images;

[0037] Step 102, perform feature extraction on the visible light image to obtain a first feature map of the visible light image at different feature extraction scales;

[0038] Step 103, perform feature extraction on the infrared image to obtain a second feature map of the infrared image at different feature extraction scales;

[0039] Step 104, perform modal fusion and multi-scale fusion on all the first feature maps and all the second feature maps to obtain a feature enhancement map of the dataset of images to be detected at different scales;

[0040] Step 105, perform classification and regression on the feature enhancement maps at different scales to obtain the detection results of the dataset of images to be detected; the detection results include the category information and location information of the objects included in the dataset of images to be detected.

[0041] Specifically, first, the visible light image and the infrared image are respectively input into two backbone networks for single-modal feature extraction. During the process of single-modal feature extraction, a self-attention mechanism fusion module is used for cross-modal feature fusion to establish long-range dependencies between the two modalities and obtain better fused feature information. An improved feature pyramid is used as the neck network for multi-scale feature fusion of the fused feature maps of different sizes, that is, the feature enhancement maps at different scales. The feature enhancement maps are sent to the head network for classification and regression, and the final detection results are output. The detection results include the category and location information of the target, realizing accurate detection of the target under low-light conditions.

[0042] The head network can predict the category and location of the target, achieving accurate detection of the target. At the same time, through technical means such as adaptive anchor box calculation and loss function optimization, the accuracy and generalization ability of target detection are further improved.

[0043] Specifically, feature extraction is performed on the visible light image and the infrared image. CSPDarknet is used as the backbone network. CSPDarknet combines the advantages of cross-stage partial connection and the Darknet network. By reducing computational redundancy and improving the model inference speed, the model has a fast processing speed while maintaining high accuracy.

[0044] Among them, single-modal feature extraction is respectively performed on the visible light image and the infrared image, and the feature extraction uses the following formula:

[0045]

[0046] Among them respectively represent the feature maps of the i-th layer (i = 1, 2, 3, 4, 5) of the infrared image and the visible light image. I R , I T respectively represent the input infrared image and visible light image, Φ backbone is the feature extraction function for the visible light image and the infrared image, and the parameters are θ R , θ T . The present invention uses CSPDarkNet as the backbone network to extract the main features, that is, as the Φ backbone function.

[0047] For feature maps (j = 3, 4, 5), cross-modal feature fusion is required in multi-spectral target detection to aggregate the features of the visible light and infrared branches. The cross-modal feature fusion is as follows:

[0048]

[0049] Among them represents the fused feature of the j-th layer, φfusion (·) represents the feature fusion function with parameter θ f The present invention adopts feature-level fusion in multi-spectral object detection to fuse multi-modal features of convolutional layers C3 to C5.

[0050] The method in this embodiment obtains a visible light image and an infrared image, extracts features from the visible light image and the infrared image, performs modal fusion and multi-scale fusion on the extracted feature maps to obtain feature enhancement maps at different scales, and classifies and regresses the feature enhancement maps at different scales to obtain the detection results of the image dataset to be detected. The infrared image has strong imaging penetration at night and in bad weather conditions and is suitable for object detection and monitoring at night and in bad weather conditions. By fusing the information of the infrared image and the visible light image, all-weather detection and positioning of any object can be achieved at night and in bad weather conditions. By fusing multi-scale information, the adaptability of the object detection algorithm under different scenarios, lighting changes and scale changes can be improved, thereby improving the accuracy of object detection and thus improving the ability of the autonomous driving system to make correct judgments.

[0051] In a specific embodiment, the performing modal fusion and multi-scale fusion on all the first feature maps and all the second feature maps to obtain the feature enhancement maps of the image dataset to be detected at different scales includes: inputting a first target feature map into a preset visible light attention branch structure to perform two different convolutional operations to obtain a first visible light feature map and a second visible light feature map; the first target feature map is any one of all the first feature maps; inputting a second target feature map into a preset infrared attention branch structure to perform two different convolutional operations to obtain a first infrared feature map and a second infrared feature map; the second target feature map is the feature map in all the second feature maps whose scale is the same as that of the first target feature map; obtaining a first target feature enhancement map according to the first visible light feature map, the second visible light feature map, the first infrared feature map, the second infrared feature map, the first target feature map and the second target feature map; all the first target feature enhancement maps constitute the feature enhancement maps at different scales.

[0052] Specifically, the self-attention mechanism has the advantages of long-range dependence modeling and capturing feature correlations. In this embodiment, this mechanism is introduced into multi-spectral feature fusion as φ fusion function, and the structure is as Figure 2As shown. This module has two attention branches with the same structure, forming a symmetric structure, including a visible light attention branch and an infrared attention branch. Before entering each branch, the input features of visible light and infrared features are connected through channels, and the first visible light feature map, the second visible light feature map containing visible light channel information, the first infrared feature map, and the second infrared feature map containing infrared channel information are output.

[0053] In a specific embodiment, obtaining the first target feature enhancement map according to the first visible light feature map, the second visible light feature map, the first infrared feature map, the second infrared feature map, the first target feature map, and the second target feature map includes: obtaining a visible light attention weight according to the first visible light feature map and the second visible light feature map; obtaining an infrared attention weight according to the first infrared feature map and the second infrared feature map; obtaining the first target feature enhancement map according to the visible light attention weight, the infrared attention weight, the first target feature map, and the second target feature map.

[0054] Specifically, taking the visible light attention branch as an example to illustrate the feature action process, and the same for infrared. As Figure 3 shown, the feature map F m enters two branches in the visible light attention branch. One of them changes the shape from 2C×H×W to HW×1×1 through a 1×1 convolutional layer and a reshape operation, and the obtained first visible light feature map is denoted by Q R and is enhanced through a softmax function. The other branch changes the shape to C×HW×1 through a 1×1 convolutional layer and a reshape operation, and the obtained second visible light feature map is denoted by V R . Multiply Q R and V R to obtain the visible light attention weight.

[0055] In a specific embodiment, obtaining the visible light attention weight according to the first visible light feature map and the second visible light feature map includes: multiplying the first visible light feature map and the second visible light feature map; obtaining the visible light attention weight according to the visible light feature map after matrix multiplication, a preset convolutional layer, a preset layernorm layer, and a sigmoid function.

[0056] Specifically, by combining a convolutional layer, a layernorm layer, and a sigmoid function, the model can learn more robust feature representations, thereby improving the generalization ability for unseen data. The layernorm layer can accelerate the convergence speed of the model, enabling the model to reach the optimal solution faster during training. The attention mechanism can improve the accuracy of the model for tasks by focusing on important parts of the input data. Especially in visible light image processing tasks, the attention mechanism can help the model better identify key information, thereby improving the overall performance.

[0057] In a specific embodiment, the visible light attention weights are obtained using the following formula:

[0058] A R = sg[ln(conv3[sm(conv1(F m ))×conv2(F m )])]

[0059] where A R is the visible light attention weight, the function sg(·) is the sigmoid function, the function ln(·) is a preset layernorm layer, the function sm(·) is the softmax layer, the function conv(·) is a preset convolutional layer, the numbers 1 and 2 in the function conv(·) represent preset convolutional layers at different positions, F m is the first target feature map, conv1(F m ) is the first visible light feature map, and conv2(F m ) is the second visible light feature map.

[0060] In a specific embodiment, the first target feature enhancement map is obtained using the following formula:

[0061] F = A R ×F R + A T ×F T

[0062] where F is the first target feature enhancement map, A R is the visible light attention weight, F R is the first target feature map, A T is the infrared attention weight, and F T is the second target feature map.

[0063] Specifically, the deep feature map has a low resolution but a large receptive field of the model, and is sensitive to objects of relatively large sizes. The shallow feature map, on the other hand, contains rich spatial detail information and is more conducive to the detection of small targets. The structure of the current Feature Pyramid Network (FPN) mainly consists of a bottom-up path and a top-down path, as Figure 4 shown. Among them, the bottom-up process is the forward feature extraction process of the deep convolutional network, and the top-down process is the process of upsampling the feature map of the last convolutional layer. The horizontal connection between the two paths is the process of fusing the features of the deep convolutional layer and the shallow convolutional features, thus constructing a deeper feature pyramid that integrates multi-layer feature information. However, when the depth of a network is large, a huge amount of detailed feature information will be lost during the feature extraction by FPN.

[0064] In this embodiment, based on FPN, the information path is shortened and the feature pyramid is enhanced with the accurate positioning information of the lower levels, creating a bottom-up path enhancement to transmit the semantic information once again from low dimension to high dimension. The improved feature pyramid is used as the neck network to perform multi-scale feature fusion on feature maps of different sizes and output more discriminative and effective features. As Figure 5 shown.

[0065] In a specific embodiment, the detection result of the image dataset to be detected is obtained by the following formula:

[0066]

[0067] where D cls is the category information of the object in the detection result, D bbox is the location information of the object in the detection result, δ head is the detection head function, δ neck is the multi-scale feature fusion function, is the feature enhancement map at different scales, and θ h is the preset detection parameter.

[0068] Specifically, the obtained multiple feature enhancement maps are further fed into the head network for classification and regression, and the final detection result including the category and location information of the target is output. The formula is as follows:

[0069] Compared with the traditional multi-spectral target detection and recognition algorithm, this method does not rely on hand-designed prior features, can adapt to complex backgrounds, has good robustness to environmental changes, good feature expression ability, and good detection accuracy.

[0070] Compared with the visible light object detection and recognition algorithm based on deep learning, this method introduces multi-spectral information and proposes a multi-spectral object detection method based on the fusion of self-attention mechanism. This method can simultaneously aggregate the complementary information of visible light and infrared images, realize the complementarity of feature information, improve the distinguishability between the target and the background, and achieve the rapid recognition of the target under low-light conditions.

[0071] This method adopts a self-attention mechanism fusion module, which uses the self-attention mechanism to capture the long-range correlation between different feature channels, adaptively generates weights representing the importance of information, so as to eliminate redundancy and enhance complementarity. This method can provide more comprehensive and robust information for the subsequent network, thereby improving the detection accuracy.

[0072] This method uses an improved feature pyramid as the neck network, and constructs semantic features and location information at different scales through upsampling, connecting elements, multiplying, and a hierarchical structure with lateral connections, which helps the detection and recognition of multi-scale targets.

[0073] In addition, in order to more intuitively evaluate the detection results, the proposed method is qualitatively compared with the baselines on the FLIR and LLVIP datasets, as shown respectively in Figure 6 、 Figure 7 From a visual perspective, even for densely occluded objects, the present invention can still detect all objects, while the baseline method has multiple false positives (FP) or false negatives (FN), that is, false detections. The first column: color image, the second column: thermal image. From the top row to the bottom row: ground truth, detection results of the baseline. The red inverted triangle indicates a false negative (FN).

[0074] In a specific embodiment, please refer to Figure 8, which is a schematic structural diagram of an object detection system based on multispectral in the second embodiment of the present application. The system includes: a data acquisition module 201, a first feature extraction module 202, a second feature extraction module 203, a fusion module 204, and a detection module 205; the data acquisition module 201 is used to acquire a dataset of images to be detected; the dataset of images to be detected includes visible light images and infrared images; the first feature extraction module 202 is used to extract features from the visible light images to obtain first feature maps of the visible light images at different feature extraction scales; the second feature extraction module 203 is used to extract features from the infrared images to obtain second feature maps of the infrared images at different feature extraction scales; the fusion module 204 is used to perform modal fusion and multi-scale fusion on all the first feature maps and all the second feature maps to obtain feature enhanced maps of the dataset of images to be detected at different scales; the detection module 205 is used to perform classification and regression on the feature enhanced maps at different scales to obtain the detection results of the dataset of images to be detected; the detection results include the category information and position information of the objects included in the dataset of images to be detected.

[0075] The system in this embodiment acquires visible light images and infrared images, extracts features from the visible light images and infrared images, and performs modal fusion and multi-scale fusion on the extracted feature maps to obtain feature enhanced maps at different scales; performs classification and regression on the feature enhanced maps at different scales to obtain the detection results of the dataset of images to be detected. Infrared images have strong imaging penetration at night and under bad weather conditions and are suitable for target detection and monitoring at night and under bad weather conditions. By fusing the information of infrared images and visible light images, all-weather detection and positioning of any target can be realized at night and under bad weather conditions. By fusing multi-scale information, the adaptability of the target detection algorithm under different scenarios, illumination changes, and scale changes can be improved, thereby improving the accuracy of target detection and thus improving the ability of the autonomous driving system to make correct judgments.

[0076] In a specific embodiment, the third embodiment of the present application provides an object detection device based on multispectral, including a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the steps of the method according to any one of the first embodiments of the present application.

[0077] In a specific embodiment, the fourth embodiment of the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of the first embodiments of the present application.

[0078] The above embodiments only represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

[0079] The above is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still belong to the protection scope of the technical solution of the present invention.

Claims

1. A multi-spectral object detection method, characterized in that: The method comprises: Acquire a data set of images to be detected; the data set of images to be detected includes visible light images and infrared images; Performing feature extraction on the visible light image to obtain a first feature map of the visible light image at different feature extraction scales; Performing feature extraction on the infrared image to obtain a second feature map of the infrared image at different feature extraction scales; Performing modality fusion and multi-scale fusion on all the first feature maps and all the second feature maps to obtain feature enhancement maps of the image data set to be detected at different scales; Classifying and regressing the feature enhancement images at different scales to obtain a detection result of the image data set to be detected; the detection result includes category information and position information of the objects included in the image data set to be detected.

2. The multi-spectral object detection method according to claim 1, characterized in that: The performing modality fusion and multi-scale fusion on all the first feature maps and all the second feature maps to obtain feature enhancement maps of the image data set to be detected at different scales includes: Inputting the first target feature map into a preset visible light attention branch structure to perform two different convolution operations to obtain a first visible light feature map and a second visible light feature map; the first target feature map is any feature map among all the first feature maps; Inputting the second target feature map into the preset infrared attention branch structure to perform two different convolution operations to obtain the first infrared feature map and the second infrared feature map; the second target feature map is a feature map of all the second feature maps whose feature map scale is the same as that of the first target feature map; A first target feature enhancement map is obtained according to the first visible light feature map, the second visible light feature map, the first infrared feature map, the second infrared feature map, the first target feature map and the second target feature map; all the first target feature enhancement maps constitute the feature enhancement maps at different scales.

3. The multi-spectral object detection method according to claim 2, characterized in that: The obtaining a first target feature enhancement map according to the first visible light feature map, the second visible light feature map, the first infrared feature map, the second infrared feature map, the first target feature map, and the second target feature map comprises: Obtaining a visible light attention weight according to the first visible light feature map and the second visible light feature map; Obtaining an infrared attention weight according to the first infrared characteristic image and the second infrared characteristic image; A first target feature enhancement map is obtained according to the visible light attention weight, the infrared attention weight, the first target feature map and the second target feature map.

4. The multi-spectral object detection method according to claim 3, characterized in that: The first target feature enhancement map is obtained using the following formula: F=A R ×F R +A T ×F T Among them, F is the first target feature enhancement map, A R is the visible light attention weight, F R is the first target feature map, A T is the infrared attention weight, F T is the second target feature map.

5. The multi-spectral object detection method according to claim 3, characterized in that: The obtaining a visible light attention weight according to the first visible light feature map and the second visible light feature map includes: Performing matrix multiplication on the first visible light characteristic map and the second visible light characteristic map; The visible light attention weight is obtained according to the visible light feature map after matrix multiplication, the preset convolution layer, the preset layernorm layer and the sigmoid function.

6. The multi-spectral object detection method according to claim 5, characterized in that: The visible light attention weight is obtained using the following formula: A R =sg[ln(conv3[sm(conv1(F m ))×conv2(F m )])] Among them, A R is the visible light attention weight, function sg(·) is the sigmoid function, function ln(·) is the preset layernorm layer, function sm(·) is the softmax layer, function conv(·) is the preset convolutional layer, the numbers 1 and 2 in function conv(·) represent the preset convolutional layers at different positions, respectively, and F m is the first target feature map, conv1(F m ) is the first visible light feature map, conv2(F m ) is the second visible light characteristic diagram.

7. The multi-spectral object detection method according to claim 1, characterized in that: The detection result of the image data set to be detected is obtained using the following formula: Among them, D cls is the category information of the object in the detection result, D bbox is the position information of the object in the detection result, δ head is the detection head function, δ neck is the multi-scale feature fusion function, is the feature enhancement map at different scales, θ h To preset the detection parameters.

8. A multi-spectral object detection system, characterized in that: The system comprises: a data acquisition module, a first feature extraction module, a second feature extraction module, a fusion module and a detection module; The data acquisition module is used to acquire a data set of images to be detected; the data set of images to be detected includes visible light images and infrared images; The first feature extraction module is used to perform feature extraction on the visible light image to obtain a first feature map of the visible light image at different feature extraction scales; The second feature extraction module is used to extract features from the infrared image to obtain a second feature map of the infrared image at different feature extraction scales; The fusion module is used to perform modality fusion and multi-scale fusion on all the first feature maps and all the second feature maps to obtain feature enhancement maps of the image data set to be detected at different scales; The detection module is used to classify and regress the feature enhancement images at different scales to obtain the detection results of the image data set to be detected; the detection results include category information and position information of the objects included in the image data set to be detected.

9. A multi-spectral object detection device, comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 2 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 2 to 8.