Multi-modal target detection method, multi-modal target detection device, medium and equipment

By combining multimodal object detection methods with infrared and visible light images, dual-stream feature extraction and context-guided pyramid networks are used to solve the robustness and accuracy of occlusion target detection in complex environments, and efficient detection in strong interference occlusion environments is achieved.

CN120339575APending Publication Date: 2025-07-18XIDIAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510294549.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing occlusion target detection methods are poor in the face of complex environments, making it difficult to accurately distinguish occlusion targets, and the detection performance of single-modal image detection algorithms in severe weather and occlusion situations is degraded.

Method used

Multimodal object detection method is adopted, combined with infrared images and visible light images, and features of different scales are extracted and fused through the dual-stream feature extraction module, context-guided pyramid network and prediction module. The cross-modal fusion and context-guided pyramid network enhance feature representation to improve detection accuracy.

Benefits of technology

It improves the accuracy and robustness of occlusion target detection, can maintain efficient detection in a strong interference occlusion environment, reduce missed detection and false alarm rates, and improves the model's perception of multi-scale features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339575A_ABST
    Figure CN120339575A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal target detection method, a multi-modal target detection device, a medium and electronic equipment, and relates to the field of image processing. The multi-modal target detection method comprises the following steps: acquiring an infrared image and a visible light image of the same scene; the infrared image and the visible light image are input into a multi-modal target detection model, the multi-modal target detection model comprises a double-flow feature extraction module, a context-guided pyramid network and a prediction module, the double-flow feature extraction module comprises two processing flows, and feature maps of different scales are extracted from the infrared image and the visible light image respectively; the context-guided pyramid network carries out feature fusion on the feature maps from a large scale to a small scale to obtain a fused feature map; the prediction module is used for determining a prediction frame where a target is located in the fusion feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular, to a multi-modal object detection method, a multi-modal object detection device, a medium, and a device. Background Art

[0002] The task of object detection is to quickly and accurately locate the objects in an image, and occluded object detection is an important topic in the field of object detection. Occluded objects refer to the situation where there is a dense scene in the image, and objects are occluded by each other or by non-objects. Currently, the occluded object algorithms are divided into two types: traditional algorithms and deep learning algorithms. The traditional occluded object detection methods mainly rely on manually designed feature extraction algorithms to describe and detect objects, and locate and identify objects by combining feature extraction and classifiers. Typical traditional occluded object detection methods include methods based on object features, methods based on background features, methods based on image data structures, etc. Most of these methods rely on prior information and have poor robustness and generality in the face of strong occlusion backgrounds in complex environments.

[0003] The object detection method based on deep learning can use training to obtain the features of each layer of image data, and then learn information to output results through a classifier. In recent years, with the development of artificial intelligence technology, deep learning algorithms have achieved good results in the field of occluded object detection. According to different design ideas, deep learning object detection methods can be divided into three categories: candidate boxes, regression, and generative adversarial networks. Deep learning methods have better robustness and accuracy than traditional algorithms, and can also effectively handle multiple objects and scale changes. However, occluded objects may have very similar features to each other, resulting in the detection model being unable to accurately distinguish occluded objects. Moreover, the occlusion situation is diverse and complex, and existing detectors cannot learn all occlusion models, and deep convolutional networks are not robust in the face of partial occlusion. Most deep learning-based occluded object detection methods are basically based on two-dimensional visual information, that is, the color and texture information provided by visible light images or infrared images for detection, and cannot avoid the problem of reduced detection performance caused by similar color and texture features between the object and the background area. Summary of the Invention

[0004] The present application provides a multi-modal object detection method, a multi-modal object detection device, a medium, and a device, which can improve the accuracy and robustness of object detection algorithms.

[0005] In a first aspect, the present application provides a multi-modal object detection method, including:

[0006] Obtain an infrared image and a visible light image of the same scene;

[0007] Input the infrared image and the visible light image into a multi-modal object detection model, which includes a two-stream feature extraction module, a context-guided pyramid network, and a prediction module.

[0008] The two-stream feature extraction module includes two processing streams, which respectively extract feature maps of different scales from the infrared image and the visible light image.

[0009] The context-guided pyramid network performs feature fusion on the feature maps from large scale to small scale to obtain fused feature maps.

[0010] The prediction module is used to determine the prediction boxes where the targets are located in the fused feature maps.

[0011] According to the multi-modal object detection method in the embodiments of the present application, the two-stream backbone network is used to obtain information of different modal images for complementarity, extract features of different scales, and integrate these multi-scale features together. It has good representation ability on the basis of feature extraction and enhancement, helps to capture multi-level features of the image, and can improve the model's perception of multi-scale features. The mid-term fusion strategy is adopted to fuse the extracted dual-modal image features after the feature extraction network. This fusion strategy can ensure that the backbone network can retain the independent key features of each modality. The context-guided pyramid network is used for feature fusion, so that it has both position information from the shallow layer and semantic information from the deep layer, thereby enhancing the effectiveness of feature representation and effectively guiding the model to learn information about the detection targets, thus improving the detection accuracy of the model.

[0012] In an exemplary embodiment, each processing stream of the two-stream feature extraction module includes a feature extraction module, and the feature extraction module includes two GBS modules, an MSEE module connected in sequence, and alternating GBS modules and MSEE modules.

[0013] The GBS module is used to perform ghost convolution, batch normalization, and SiLU activation function processing on the infrared image or the visible light image to obtain the processed feature map.

[0014] The MSEE module extracts local features of different scales through multi-scale pooling operations; performs edge enhancement on the local features of different scales, interpolates the enhanced local features, and splices the interpolated features, and processes the spliced features through convolution operations.

[0015] The feature maps obtained after the processing of two GBS modules and one MSEE module are successively passed through multiple GBS modules and MSEE modules to obtain the feature maps of small targets, medium targets, and large targets.

[0016] In an exemplary embodiment, the feature extraction module further includes a cross-modal fusion module.

[0017] The cross-modal fusion module splices the feature maps corresponding to the infrared image and the visible light image respectively to obtain a bimodal feature map; the feature maps of the infrared image and the visible light image include the feature maps of large targets, medium targets and small targets;

[0018] The cross-modal fusion module performs a convolution operation on the bimodal feature map to obtain a correlation matrix; the correlation matrix is respectively input into two independent GBS modules to obtain an infrared corrected feature and a visible light corrected feature; the feature map of the infrared image is spliced with the infrared corrected feature to obtain a first spliced feature, the feature map of the visible light image is spliced with the visible light corrected feature to obtain a second spliced feature, and then the first spliced feature and the second spliced feature are spliced to obtain a first fusion feature map.

[0019] In an exemplary embodiment, it further includes:

[0020] The first fusion feature map is sequentially processed by a residual module and a SimAM attention module to obtain a processed feature map, which is used as the output of the cross-modal fusion module.

[0021] In an exemplary embodiment, the feature extraction module further includes a spatial pyramid pooling layer, which is connected after the MSEE module,

[0022] The spatial pyramid pooling layer includes two branches. The first branch adjusts the number of channels of the feature map output by the MSEE module;

[0023] The second branch performs max-pooling processing on the feature map output by the MSEE module through different max-pooling layers, and splices the feature maps obtained by processing each max-pooling layer;

[0024] The feature map processed by the first branch is spliced with the feature map processed by the second branch as the output of the spatial pyramid pooling layer.

[0025] In an exemplary embodiment, the context-guided pyramid network includes a plurality of context-guided fusion modules;

[0026] The context-guided fusion module performs fusion from large scale to small scale. The first fusion feature map of large targets is fused with the first fusion feature map of medium targets to obtain a first fusion map; the first fusion map is fused with the first fusion feature map of small targets to obtain a second fusion map;

[0027] The fusion of the first fusion feature map of large targets and the first fusion feature map of medium targets to obtain a first fusion map specifically includes:

[0028] The first fusion feature map of the large target is spliced with the first fusion feature map of the medium target, and the SE attention module is used to enhance the spliced feature map. The enhanced feature map is segmented to obtain a first segmentation map and a second segmentation map. Multiply the first segmentation map by the first fusion feature map of the large target, and add the multiplied feature map to the first fusion feature map of the medium target to obtain a first merged map. Multiply the second segmentation map by the first fusion feature map of the medium target, and add the multiplied feature map to the first fusion feature map of the large target to obtain a second merged map. The first merged map and the second merged map are spliced to obtain a first fusion map.

[0029] In an exemplary embodiment, the context-guided pyramid network further includes a C2F module, which is connected to the context-guided fusion module.

[0030] After the C2F module processes the first fusion map output by the context-guided fusion module, a first object detection map is obtained.

[0031] The first object detection map and the first fusion feature map of the small target are jointly input into the context-guided fusion module for fusion to obtain a second fusion map, and then the second fusion map is input into the C2F module to obtain a second object detection map.

[0032] The second object detection map is input into the GBS module and then jointly input into the context-guided fusion module with the first object detection map for fusion to obtain a third object detection map.

[0033] The first object detection map, the second object detection map, and the third object detection map are respectively input into the occlusion attention module for depthwise separable convolution processing to obtain corresponding object feature maps.

[0034] In a second aspect, the present application provides a multi-modal object detection device, including:

[0035] An image acquisition module for acquiring an infrared image and a visible light image of the same scene.

[0036] A model processing module for inputting the infrared image and the visible light image into a multi-modal object detection model, where the multi-modal object detection model includes a two-stream feature extraction module, a context-guided pyramid network, and a prediction module.

[0037] The two-stream feature extraction module includes two processing streams, which respectively extract feature maps of different scales from the infrared image and the visible light image.

[0038] The context-guided pyramid network performs feature fusion on the feature maps from large scale to small scale to obtain fusion feature maps.

[0039] The prediction module is used to determine the prediction box where the object is located in the fusion feature map.

[0040] In a third aspect, the present application provides an electronic device, which includes a memory and one or more processors. Among them, one or more computer programs are stored in the memory, and the computer programs include instructions. When the instructions are executed by the processor, the electronic device can execute the multi-modal target detection method in the first aspect.

[0041] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions run on an electronic device, the electronic device is caused to execute the multi-modal target detection method in the first aspect.

[0042] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device is caused to execute the multi-modal target detection method described in the first aspect.

[0043] It can be understood that for the beneficial effects that can be achieved by the above-provided multi-modal target detection device, electronic device, computer-readable storage medium, and computer program product, reference can be made to the beneficial effects in the first aspect, and details are not elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic flowchart of the multi-modal target detection method provided by an embodiment of the present application;

[0045] Figure 2 It is a schematic diagram of the model structure in the multi-modal target detection method provided by an embodiment of the present application Figure 1 ;

[0046] Figure 3 It is a schematic diagram of the model structure in the multi-modal target detection method provided by an embodiment of the present application Figure 2 ;

[0047] Figure 4 It is a schematic diagram of the model structure in the multi-modal target detection method provided by an embodiment of the present application Figure 3 ;

[0048] Figure 5 It is a schematic diagram of the model structure in the multi-modal target detection method provided by an embodiment of the present application Figure 4 ;

[0049] Figure 6 It is a schematic diagram of the model structure in the multi-modal target detection method provided by an embodiment of the present application Figure 5 ;

[0050] Figure 7 It is a schematic diagram of the structure of the electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] In order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. For example, the first chip and the second chip are only used to distinguish different chips, and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily limit differences. It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way. In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more.

[0052] It should be noted that "when... " in the embodiments of the present application can be at the instant when a certain situation occurs, or within a period of time after a certain situation occurs. The embodiments of the present application do not make specific limitations on this.

[0053] To overcome the situation that in the existing deep learning-based object detection methods, when facing occluded objects, the objects will be occluded by each other or by obstacles, resulting in the detection model being unable to accurately distinguish the objects. At the same time, most of the current object detection algorithms are mainly single-modal detection algorithms based on visible light or infrared images, but single-modal images have their respective defects. Visible light images are difficult to obtain complete object information in the face of complex situations such as bad weather, occlusion, and low light, and cannot achieve precise detection tasks. Infrared images are formed by the thermal radiation energy emitted by objects and can better reflect the thermal radiation characteristics in the images. However, compared with visible light images, infrared images are not sensitive to the characteristics of scene brightness changes, and have low clarity, blurred visual effects, and no linear relationship between the gray distribution and the target reflection characteristics.

[0054] To address the above problems, this embodiment provides a multi-modal object detection method. The implementation manner of this embodiment will be described in detail below with reference to the accompanying drawings.

[0055] Exemplarily, the multi-modal object detection method can be applied to various electronic devices such as a computer (PC), a tablet computer, a virtual reality / augmented reality device, a wearable device, an industrial computer, a vehicle-mounted computer, etc.; it can also be applied to a server, the cloud, a server cluster, etc. The embodiments of the present application do not make special limitations on this.

[0056] Figure 1 The flowchart of the multi-modal object detection method provided by the embodiments of the present application is shown.

[0057] As Figure 1 shown, the multi-modal object detection method may include the following steps:

[0058] Step 101: Obtain an infrared image and a visible light image of the same scene.

[0059] Exemplarily, first, the number of channels of the infrared image is upsampled to 3 dimensions to align with the dimension of the visible light image for subsequent operations.

[0060] Step 102: Input the infrared image and the visible light image into a multi-modal object detection model, where the multi-modal object detection model includes a two-stream feature extraction module, a context-guided pyramid network, and a prediction module.

[0061] Step 103: The two-stream feature extraction module includes two processing streams, which respectively extract feature maps of different scales from the infrared image and the visible light image.

[0062] Step 104: The context-guided pyramid network performs feature fusion on the feature maps from large scale to small scale to obtain a fused feature map.

[0063] Step 105: The prediction module is used to determine the prediction box where the target is located in the fused feature map.

[0064] Figure 2 shows the structural diagram of the multi-modal object detection model. Refer to Figure 2 , each processing stream of the two-stream feature extraction module includes a feature extraction module, and the feature extraction module includes two GBS modules, an MSEE module connected in sequence, and alternating GBS modules and MSEE modules; the GBS module is used to perform ghost convolution, batch normalization, and SiLU activation function processing on the infrared image or the visible light image to obtain a processed feature map; the MSEE module extracts local features of different scales through multi-scale pooling operations; performs edge enhancement on the local features of different scales, interpolates the enhanced local features, and splices the interpolated features, and processes the spliced features through convolution operations; the feature maps obtained after the processing of two GBS modules and one MSEE module are successively passed through multiple GBS modules and MSEE modules to obtain feature maps of small targets, medium targets, and large targets.

[0065] The feature extraction module further includes a cross-modal fusion module; the cross-modal fusion module splices the feature maps corresponding to the infrared image and the visible light image respectively to obtain a bimodal feature map; the feature maps of the infrared image and the visible light image include the feature maps of large targets, medium targets and small targets; the cross-modal fusion module performs a convolution operation on the bimodal feature map to obtain a correlation matrix; the correlation matrix is respectively input into two independent GBS modules to obtain an infrared corrected feature and a visible light corrected feature; the feature map of the infrared image is spliced with the infrared corrected feature to obtain a first spliced feature, the feature map of the visible light image is spliced with the visible light corrected feature to obtain a second spliced feature, and then the first spliced feature and the second spliced feature are spliced to obtain a first fusion feature map.

[0066] The cross-modal feature fusion module also processes the first fusion feature map through a residual module and a SimAM attention module in sequence to obtain a processed feature map as the output of the cross-modal fusion module.

[0067] The feature extraction module further includes a spatial pyramid pooling layer connected after the MSEE module. The spatial pyramid pooling layer includes two branches. The first branch adjusts the number of channels of the feature map output by the MSEE module; the second branch performs max-pooling processing on the feature map output by the MSEE module through different max-pooling layers, and splices the feature maps obtained by each max-pooling layer processing; the feature map processed by the first branch and the feature map processed by the second branch are spliced as the output of the spatial pyramid pooling layer.

[0068] Exemplarily, the first fusion feature map can be processed through a residual module and a SimAM attention module in sequence to obtain a processed feature map as the output of the cross-modal fusion module.

[0069] The feature extraction module adopts Ghost Convolutions (Ghost Conv). The specific operation is as follows: First, use ordinary convolution f′ (half the number of normal convolution kernels) to obtain an eigen feature map Y′, where X is the input data. As shown in Equation (2), use a linear operation Φ (where Φ is a 3*3 or 5*5 depth convolution) to operate on each feature map y′ of Y′ i to obtain a feature map y i,j . Finally, the eigen feature map Y′ and the feature map y i,jPerform splicing to obtain the output. As shown in Equation (3) for ordinary convolution, when undergoing s transformations, the computational cost of Ghost Conv (excluding the BN (Batch Normalization) layer and activation function) will decrease by approximately s times. Among them, the data input format is c*h*w, representing the input channel, image height, and width respectively. n is the number of convolution kernels, k is the convolution kernel size, d is the size of the linear transformation convolution kernel, s is the number of transformations, and m is the feature map identifier.

[0070] Y′ = X * f′ (1)

[0071]

[0072] Y = [y 11 , y 12 , … y 1s , … y ms (3)

[0073]

[0074] The infrared image and the visible light image respectively pass through the GBS module twice and then are combined with the MSEE module to perform multi-scale feature extraction, edge information enhancement, and convolution operations. The GBS module is a network composed of Ghost Conv processing, Batch Normalization layer (BN), and SiLU activation function. The structure of the Multi-scale edge enhancement (MSEE) module is as Figure 3 shown. The main purpose of this module is to extract features from different scales, highlight edge information, integrate these multi-scale features together, and finally output enhanced features through the convolutional layer. Specifically, the MSEE module first performs multi-scale pooling through adaptive pooling to extract local information of different sizes, which helps to capture multi-level features of the image. Among them, the edge enhancement module is specifically used to extract edge information, making the network more sensitive to edges, which plays an important role in the target detection task. By aligning the features extracted at different scales to the same scale through interpolation operations, then splicing them together, and finally fusing them into a unified feature representation through the convolutional layer, the model's perception of multi-scale features can be improved. At this time, the visible light and infrared images respectively obtain the visible light feature map y Vis1 and the infrared feature map y IR1 . The size of the feature map is 1 / 16 of the original image.

[0075] The feature map y Vis1 and the feature map y IR1 respectively pass through a GBS layer and the MSEE module to obtain the visible light small target feature map y Vis2 and the infrared small target feature map y IR2 , and the size of the feature map is 1 / 8 of the original image. The feature map yVis2 With the feature map y IR2 After passing through a GBS layer and an MSEE module respectively, the target feature map y in visible light is obtained Vis3 And the target feature map y in infrared IR3 , the size of the feature map is 1 / 16 of the original image. After that, the target feature map y in visible light Vis3 And the target feature map y in infrared IR3 After passing through a GBS layer and an MSEE module respectively, they pass through a spatial pyramid pooling layer, which consists of two branches. The first branch only changes the number of channels, and the second branch connects through different max-pooling layers to obtain different receptive fields, thereby increasing the receptive field. Then, through convolution, the feature maps of different pooling layers are concatenated. Finally, the feature maps of the first branch and the second branch are concatenated to obtain the large target feature map y in visible light Vis4 And the large target feature map y in infrared IR4 , the size of the feature map is 1 / 32 of the original image.

[0076] According to the visible light and infrared feature maps obtained in the previous steps, they are input into the cross-modal fusion module (RTF) for feature fusion. The RTF module is as Figure 4 shown. In this embodiment, a mid-term fusion strategy is adopted to fuse the extracted bimodal image features after the feature extraction network. This fusion strategy can ensure that the backbone network can retain the independent key features of each modality. The small target feature map y in visible light Vis2 And the small target feature map y in infrared IR2 Are input into the cross-modal fusion module; the medium target feature map y in visible light Vis3 And the medium target feature map y in infrared IR3 Are input into the cross-modal fusion module; the large target feature map y in visible light Vis4 And the large target feature map y in infrared IR4 Are input into the cross-modal fusion module for feature fusion. The cross-modal feature fusion module is divided into two parts. In the first stage, the visible light feature map and the infrared feature map are corrected through the interaction between the feature layers, and in the second stage, key information complementary fusion is carried out. In the first stage, the visible light feature map and the infrared feature map extracted by the two-stream backbone network are subjected to a concatenation operation to obtain a bimodal feature map, as shown in Equation (5).

[0077] y RGB-T =Concat(y Vis ; y IR ) (5)

[0078] After that, through the structure of serial connection of the GBS layer and the SimAM attention module for y RGB-TPerform a convolution operation to integrate visible light information and infrared information, capture the key features unique to each modality and learn the relationships between these features, construct a correlation representation, and decompose the correlation matrix into two corrected features containing visible light information and infrared information respectively through two independent GBS modules, which are called the visible light corrected feature y RGB and the infrared corrected feature y TH , as shown in Equation (6).

[0079] y RGB ,y TH =Split(y RGB-T ) (6)

[0080] Concatenate the corrected feature y RGB with the original visible light feature map y Vis and use the GBS module to correct the original visible light feature map. Supplement the information of the visible light image with information such as the target edge information in the infrared image to obtain a new feature map (i.e., the first concatenated feature). Similarly, correct the infrared feature map through the visible light feature map to make its details and texture information have higher distinguishability, and obtain a new infrared feature map (i.e., the second concatenated feature). Perform a concatenation operation on the two corrected feature maps to obtain a feature map containing less misleading information (i.e., the first fused feature map). The correction process is shown in the following formula.

[0081]

[0082] In the second stage, construct a residual connection structure and introduce SimAM (A Simple, Parameter-Free Attention Module) at the tail of the residual block to further filter redundant information and highlight key information, as shown in Equation (10)

[0083]

[0084] Among them, CBS refers to performing ordinary convolution operations Conv, batch normalization BN, and SiLU activation function processing.

[0085] Use the feature map obtained by Equation (10) as the output of the two-stream feature extraction network, specifically including the first fused feature maps corresponding to small targets, medium targets, and large targets respectively.

[0086] Then, the context-guided pyramid network further fuses and enhances the features extracted by the two-stream feature extraction network. Specifically, the context-guided pyramid network includes multiple context-guided fusion modules; the context-guided fusion modules perform fusion from large scales to small scales, fuse the first fusion feature map of large objects with the first fusion feature map of medium objects to obtain a first fusion map; fuse the first fusion map with the first fusion feature map of small objects to obtain a second fusion map; the step of fusing the first fusion feature map of large objects with the first fusion feature map of medium objects to obtain a first fusion map specifically includes: splicing the first fusion feature map of large objects with the first fusion feature map of medium objects, using an SE attention module to enhance the spliced feature map, segmenting the enhanced feature map to obtain a first segmentation map and a second segmentation map; multiplying the first segmentation map by the first fusion feature map of large objects, adding the multiplied feature map to the first fusion feature map of medium objects to obtain a first merged map; multiplying the second segmentation map by the first fusion feature map of medium objects, adding the multiplied feature map to the first fusion feature map of large objects to obtain a second merged map, and splicing the first merged map with the second merged map to obtain a second fusion map.

[0087] The feature maps of different target scales obtained by the feature extraction module are input into the context-guided pyramid network for feature fusion, so that it has both the position information from the shallow layer and the semantic information from the deep layer at the same time. For the context-guided feature pyramid network, one direction is top-down, which transmits and enhances the strong semantic features of the high layer, but ignores the position information of the low layer. The other direction is the reverse pyramid network, whose direction is bottom-up, which transmits the position information of the low layer. At the same time, the feature maps from different layers are fused, so that it has the position information from the low layer and the semantic information from the high layer, which is beneficial to extracting the effective features of the target and improving the detection accuracy. Among them, the context-guided fusion module is used to aggregate the information of the two-layer feature maps. Through the SE attention mechanism, the module can capture and utilize important context information during the feature fusion process, thereby enhancing the effectiveness of the feature representation and effectively guiding the model to learn the information of the detection target, thus improving the detection accuracy of the model. Through the weighted feature recombination operation, the module can enhance important features while suppressing unimportant features, and improve the discriminative ability of the feature map. The module structure is relatively simple and does not introduce too much computational overhead, making it suitable for application in real-time object detection tasks.

[0088] Fused feature map of the large object layer Through upsampling and the feature map from the medium object layer Input into the context-guided fusion module to obtain the feature map F1 (i.e., the first fusion map). The context-guided fusion module is as Figure 5As shown below. First, the double-layer feature maps are concatenated, and then the SE attention module is used to enhance the channel features of the input feature maps, which can effectively improve the accuracy of model detection. After the slicing operation and multiplying with the original feature maps, the results are added to the feature maps of different feature layers to enhance important features and suppress unimportant features, thereby improving the discriminative ability of the feature maps. As shown in the following formula

[0089]

[0090] F1 = Concat(y RGB-T2F +y RGB-T1 ×y RGB-T1F ,y RGB-T1F +y RGB-T2 ×y RGB-T2F ))) (13)

[0091] where Concat is the concatenation operation, SE is the channel attention mechanism, Spilt is the slicing operation, is the feature map of the fused large target layer, is the feature map from the medium target layer, y RGB-T1 , y RGB-T2F are the sliced feature maps, and F1 is the output feature map of the context-guided fusion module.

[0092] The context-guided pyramid network further includes a C2F module, which is connected to the context-guided fusion module. After processing the first fusion map output by the context-guided fusion module, the C2F module obtains the first object detection map; the first object detection map and the first fusion feature map of small objects are jointly input into the context-guided fusion module for fusion to obtain the second fusion map, and then the second fusion map is input into the C2F module to obtain the second object detection map; after passing the second object detection map through the GBS module, it is jointly input into the context-guided fusion module with the first object detection map for fusion to obtain the third object detection map; the first object detection map, the second object detection map, and the third object detection map are respectively input into the occlusion attention module for depthwise separable convolution processing to obtain the corresponding object feature maps.

[0093] The first fusion map F1 passes through the C2F module and upsampling to obtain the first object detection map, which is then combined with the feature map from the small object layer to obtain the second fusion map F2 through the context-guided fusion module.

[0094] After F2 passes through the C2F module and the GBS layer, the feature map of F1 passing through the C2F and the context-guided fusion module obtain F3. After the feature map F3 passes through the C2F module and the GBS layer, it combines with the feature map to obtain F4 through the context-guided fusion module.

[0095] In this embodiment, the model further includes a detection head, which is divided into three feature maps with different scales, corresponding to large targets, medium targets, and small targets respectively. The image feature information contained in the feature maps of different sizes is different. The small target feature map contains more position information from the lower layer, while the large target feature map contains more semantic information from the higher layer.

[0096] The small target detection feature map is derived from the feature map F2 through the C2F module. The medium target detection feature map F4 is derived from the map F3 through the C2F module. The feature map for large target detection is obtained by passing the feature map F4 through one layer of the C2F module.

[0097] The occlusion attention module is added after the feature maps obtained in the above steps respectively. The module is as Figure 6 shown. The first part of this module is a depthwise separable convolution with residual connections. The depthwise separable convolution operates depthwise one by one, that is, separating the convolution channel by channel. Although the depthwise separable convolution can learn the importance of different channels and reduce the number of parameters, it ignores the information relationship between channels. To make up for this loss, the outputs of different depth convolutions are then combined through pointwise convolution. Then a two-layer fully connected network is used to fuse the information of each channel, so that the network can strengthen the connections between all channels. The output learned by the fully connected layer is processed by an exponential function to expand the value range from [0,1] to [1,e]. This exponential normalization provides a monotonic mapping relationship, making the result more tolerant of position errors. Finally, the output of this module is used as the attention to multiply the original feature, enabling the model to handle occlusion more effectively.

[0098] Taking the input image of 640×640 as an example, the small target feature map y S has an image size of 80×80; the medium target feature map y M has an image size of 40×40; the large target feature map y L has an image size of 20×20.

[0099] The feature map obtained after being processed by the occlusion attention module is used as the output of the context-guided pyramid network and input into the prediction module. The prediction module includes a prediction head, which adopts an anchor-base method. Among them, the traditional intersection over union (IoU) loss function is prone to misjudgment for occluded targets because occluded targets may overlap with each other in the image and lack sufficient information. And the IoU loss function is very sensitive to the position change of the target, resulting in poor detection effect. The loss function is replaced with a repulsive loss function, making the prediction box closer to the true target box it is responsible for and farther away from the surrounding targets.

[0100] During the training process, the parameters are updated by regression using the loss function for occluded targets.

[0101] The formula is shown as follows and is divided into three parts.

[0102] L = L Attr + α * L RepGT + β * L RepBox (14)

[0103]

[0104]

[0105] Among them, L Attr is the loss value (attraction term) generated by the predicted box and the true target box; L Repgt is the loss value (repulsion term (RepGT)) generated by the predicted box and the adjacent true target box; L RepBox is the loss value (repulsion Box (RepBox)) generated by the predicted box and the adjacent predicted boxes that do not predict the same true target. α and β are correlation coefficients, IoU (Intersection over Union) is the intersection over union, P is the candidate box, G is the true box, P+ is the set of positive candidate boxes, B P is obtained by regression adjustment according to the predicted box P, is the set of true boxes with the largest IoU with the predicted box P, Smooth ln is the distance matrix, and Ι is the identity function.

[0106] We set P as the candidate box, G as the true box, and P+ as the set of positive candidate boxes. The meaning of a positive candidate box is that the IoU with at least one true box is greater than a certain threshold. B P is obtained by regression adjustment according to the predicted box P, It is the set of ground truth boxes with the largest IoU with the prediction box P. The first part is the loss value (attraction term) generated by the prediction box and the ground truth target box; the second part is the loss value (repulsion term (RepGT)) generated by the prediction box and the adjacent ground truth target boxes; the third part is the loss value (repulsionBox (RepBox)) generated by the prediction box and the adjacent prediction boxes that do not predict the same ground truth target. Two correlation coefficients α and β are used to balance the two repulsive loss values. In the subsequent decoder, the denoising idea of DINO HEAD is adopted, and two hyperparameters are used to generate positive and negative samples. The noise scale of the positive samples is smaller than the smaller hyperparameter, and the noise scale of the negative samples is between the two hyperparameters. Positive samples are expected to predict the presence of objects, while negative samples are expected to have no objects. By generating hard negative samples with a small difference from the positive samples, repeated predictions are avoided and confusion is suppressed. During the detection process, the predicted results of the feature map are decoded and the prediction boxes are drawn on the original image. The prediction boxes greater than the threshold function are found, and different categories are looped through. The objects of this category are sorted according to the confidence score, and the NMS (Non-Maximum Suppression) algorithm is used to filter out redundant detection boxes to obtain the final detection results. The occluded object detection results are output.

[0107] This network is an end-to-end network, and the final output is:

[0108] output = net(input1, input2) (18)

[0109] Among them, input1 is the visible light image, and input2 is the infrared image.

[0110] The present invention has the following advantages compared with the existing methods:

[0111] (1) Compared with traditional occluded object detection algorithms, this method does not rely on prior information such as the saliency of objects and the consistency of the background. The feature extraction module based on the multi-scale feature enhancement module of the present invention has a powerful feature extraction ability and has strong robustness and generalization ability in the face of a strongly interfering occlusion environment.

[0112] (2) Compared with general occluded object detection methods, this method uses a two-stream backbone network and a cross-modal fusion module to obtain information from different modal images for complementarity. At the same time, the context-guided pyramid network layer is used to expand the receptive field to represent features more fully and can be efficiently fused with the strong semantic information at the high level, which is beneficial to improving the accuracy of occluded object detection. The occluded object loss function adopted by the present invention is less sensitive to the changes in the occlusion scene compared with the traditional loss function, making the network converge faster when facing occluded objects and the detection of occluded objects more stable, reducing the missed detection and false alarm rates of the model.

[0113] (3) Compared with specific occlusion target detection methods, the multi-scale feature enhancement module adopted in this method can extract features from different scales compared with ordinary convolution modules, highlight edge information, integrate these multi-scale features together, and has good representation ability on the basis of feature extraction and edge enhancement, which helps to capture multi-level features of images and can improve the model's perception of multi-scale features. Through the context-guided pyramid network and the occlusion target loss function, the network can effectively utilize image information of different modalities for information complementarity to improve the model detection accuracy.

[0114] Furthermore, this embodiment also provides a multi-modal target detection device, which can be used to execute the above multi-modal target detection method. The multi-modal target detection device specifically includes an image acquisition module for acquiring infrared images and visible light images of the same scene;

[0115] a model processing module for inputting the infrared image and the visible light image into a multi-modal target detection model. The multi-modal target detection model includes a two-stream feature extraction module, a context-guided pyramid network, and a prediction module. The two-stream feature extraction module includes two processing streams, which respectively extract feature maps of different scales from the infrared image and the visible light image; the context-guided pyramid network performs feature fusion on the feature maps from large scale to small scale to obtain a fused feature map; the prediction module is used to determine the prediction box where the target is located in the fused feature map.

[0116] In an exemplary implementation, each processing stream of the two-stream feature extraction module includes a feature extraction module. The feature extraction module includes two GBS modules, an MSEE module connected in sequence, and alternating GBS modules and MSEE modules; the GBS module is used to perform ghost convolution, batch normalization, and SiLU activation function processing on the infrared image or the visible light image to obtain a processed feature map; the MSEE module extracts local features of different scales through multi-scale pooling operations; performs edge enhancement on the local features of different scales, interpolates the enhanced local features, and splices the interpolated features, and processes the spliced features through convolution operations; the feature maps obtained after the processing of two GBS modules and one MSEE module are successively passed through multiple GBS modules and MSEE modules to obtain feature maps of small targets, medium targets, and large targets.

[0117] In an exemplary embodiment, the feature extraction module further includes a cross-modal fusion module; the cross-modal fusion module splices the feature maps corresponding to the infrared image and the visible light image respectively to obtain a bimodal feature map; the feature maps of the infrared image and the visible light image include the feature maps of large targets, medium targets and small targets; the cross-modal fusion module performs a convolution operation on the bimodal feature map to obtain a correlation matrix; the correlation matrix is respectively input into two independent GBS modules to obtain an infrared corrected feature and a visible light corrected feature; the feature map of the infrared image is spliced with the infrared corrected feature to obtain a first spliced feature, the feature map of the visible light image is spliced with the visible light corrected feature to obtain a second spliced feature, and then the first spliced feature and the second spliced feature are spliced to obtain a first fusion feature map.

[0118] In an exemplary embodiment, the cross-modal fusion module is further configured to: process the first fusion feature map through a residual module and a SimAM attention module in sequence to obtain a processed feature map as the output of the cross-modal fusion module.

[0119] In an exemplary embodiment, the feature extraction module further includes a spatial pyramid pooling layer, which is connected after the MSEE module. The spatial pyramid pooling layer includes two branches. The first branch adjusts the number of channels of the feature map output by the MSEE module; the second branch performs max-pooling processing on the feature map output by the MSEE module through different max-pooling layers, and splices the feature maps obtained by processing each max-pooling layer; the feature map processed by the first branch and the feature map processed by the second branch are spliced as the output of the spatial pyramid pooling layer.

[0120] In an exemplary embodiment, the context-guided pyramid network includes a plurality of context-guided fusion modules; the context-guided fusion modules perform fusion from large scale to small scale, fuse the first fusion feature map of large objects with the first fusion feature map of medium objects to obtain a first fusion map; fuse the first fusion map with the first fusion feature map of small objects to obtain a second fusion map; the step of fusing the first fusion feature map of large objects with the first fusion feature map of medium objects to obtain a first fusion map specifically includes: splicing the first fusion feature map of large objects with the first fusion feature map of medium objects, enhancing the spliced feature map by using an SE attention module, segmenting the enhanced feature map to obtain a first segmentation map and a second segmentation map; multiplying the first segmentation map by the first fusion feature map of large objects, adding the multiplied feature map to the first fusion feature map of medium objects to obtain a first merged map; multiplying the second segmentation map by the first fusion feature map of medium objects, adding the multiplied feature map to the first fusion feature map of large objects to obtain a second merged map, and splicing the first merged map and the second merged map to obtain a first fusion map.

[0121] In an exemplary embodiment, the context-guided pyramid network further includes a C2F module, which is connected to the context-guided fusion module. After processing the first fusion map output by the context-guided fusion module, the C2F module obtains a first object detection map; inputs the first object detection map and the first fusion feature map of small objects into the context-guided fusion module for fusion to obtain a second fusion map, and then inputs the second fusion map into the C2F module to obtain a second object detection map; inputs the second object detection map through a GBS module and then inputs it together with the first object detection map into the context-guided fusion module for fusion to obtain a third object detection map; respectively input the first object detection map, the second object detection map and the third object detection map into an occlusion attention module for depthwise separable convolution processing to obtain corresponding object feature maps.

[0122] The specific details of each module or unit in the above multi-modal object detection device have been described in detail in the corresponding multi-modal object detection method, so they will not be elaborated here.

[0123] An embodiment of the present application further provides an electronic device. Figure 7 The structural schematic diagram of the electronic device suitable for implementing the embodiments of the present disclosure is shown. Figure 7 The electronic device 600 shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0124] As Figure 7As shown, the electronic device 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for system operation are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0125] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required so that a computer program read from it can be installed into the storage section 608 as required.

[0126] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above functions defined in the embodiments of the present application are executed.

[0127] It should be noted that the computer-readable medium shown in this disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0129] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware, and the described units may also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0130] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and the one or more programs include instructions that, when executed by the electronic device, cause the electronic device to implement the methods described in the above embodiments.

[0131] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.

[0132] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-modal object detection method, characterized in that, Including: Obtaining an infrared image and a visible light image of the same scene; Inputting the infrared image and the visible light image into a multi-modal object detection model, the multi-modal object detection model including a two-stream feature extraction module, a context-guided pyramid network, and a prediction module, The two-stream feature extraction module includes two processing streams, respectively extracting feature maps of different scales from the infrared image and the visible light image; The context-guided pyramid network performs feature fusion on the feature maps from large scale to small scale to obtain a fused feature map; The prediction module is used to determine the prediction box where the target is located in the fused feature map.

2. The multimodal object detection method according to claim 1, wherein Each processing stream of the two-stream feature extraction module includes a feature extraction module, and the feature extraction module includes two GBS modules, an MSEE module connected in sequence, and alternating GBS modules and MSEE modules; The GBS module is used to perform ghost convolution, batch normalization, and SiLU activation function processing on the infrared image or the visible light image to obtain a processed feature map; The MSEE module extracts local features of different scales through multi-scale pooling operations; enhances the edges of the local features of different scales, performs interpolation processing on the enhanced local features, and splices the interpolated features, and processes the spliced features through convolution operations; The feature maps obtained after the processing of two GBS modules and one MSEE module are successively passed through multiple GBS modules and MSEE modules to obtain feature maps of small targets, medium targets, and large targets.

3. The multimodal object detection method according to claim 2, wherein The feature extraction module further includes a cross-modal fusion module; The cross-modal fusion module splices the feature maps corresponding to the infrared image and the visible light image respectively to obtain a bimodal feature map; the feature maps of the infrared image and the visible light image include feature maps of large targets, medium targets, and small targets; The cross-modal fusion module performs convolution operations on the bimodal feature map to obtain a correlation matrix; inputs the correlation matrix into two independent GBS modules respectively to obtain an infrared corrected feature and a visible light corrected feature; splices the feature map of the infrared image with the infrared corrected feature to obtain a first spliced feature, splices the feature map of the visible light image with the visible light corrected feature to obtain a second spliced feature, and then splices the first spliced feature and the second spliced feature to obtain a first fused feature map.

4. The multimodal object detection method according to claim 3, wherein Also including: The first fused feature map is successively processed through a residual module and a SimAM attention module to obtain a processed feature map as the output of the cross-modal fusion module.

5. The multimodal object detection method according to claim 2, wherein The feature extraction module further includes a spatial pyramid pooling layer, connected after the MSEE module, The spatial pyramid pooling layer includes two branches, and the first branch adjusts the number of channels of the feature map output by the MSEE module; The second branch performs max-pooling processing on the feature map output by the MSEE module through different max-pooling layers, and splices the feature maps obtained by the processing of each max-pooling layer; The feature map processed by the first branch and the feature map processed by the second branch are spliced as the output of the spatial pyramid pooling layer.

6. The multimodal object detection method according to claim 3, wherein, The context-guided pyramid network includes a plurality of context-guided fusion modules; The context-guided fusion modules perform fusion from large scale to small scale, fuse the first fusion feature map of large objects with the first fusion feature map of medium objects to obtain a first fusion map; fuse the first fusion map with the first fusion feature map of small objects to obtain a second fusion map; The step of fusing the first fusion feature map of large objects with the first fusion feature map of medium objects to obtain a first fusion map specifically includes: Concatenate the first fusion feature map of large objects with the first fusion feature map of medium objects, use the SE attention module to enhance the concatenated feature map, segment the enhanced feature map to obtain a first segmentation map and a second segmentation map; multiply the first segmentation map by the first fusion feature map of large objects, add the multiplied feature map to the first fusion feature map of medium objects to obtain a first merged map; multiply the second segmentation map by the first fusion feature map of medium objects, add the multiplied feature map to the first fusion feature map of large objects to obtain a second merged map, and concatenate the first merged map with the second merged map to obtain a first fusion map.

7. The multimodal object detection method according to claim 6, wherein The context-guided pyramid network further includes a C2F module, which is connected to the context-guided fusion module, After processing the first fusion map output by the context-guided fusion module, the C2F module obtains a first object detection map; Input the first object detection map and the first fusion feature map of small objects into the context-guided fusion module for fusion to obtain a second fusion map, and then input the second fusion map into the C2F module to obtain a second object detection map; Input the second object detection map through the GBS module and then input it together with the first object detection map into the context-guided fusion module for fusion to obtain a third object detection map; Input the first object detection map, the second object detection map and the third object detection map into the occlusion attention module respectively, and perform depthwise separable convolution processing to obtain corresponding object feature maps.

8. A multimodal object detection device, characterized in that, Including: An image acquisition module for acquiring infrared images and visible light images of the same scene; A model processing module for inputting the infrared images and visible light images into a multi-modal object detection model, and the multi-modal object detection model includes a two-stream feature extraction module, a context-guided pyramid network, and a prediction module, The two-stream feature extraction module includes two processing streams, which respectively extract feature maps of different scales from infrared images and visible light images; The context-guided pyramid network performs feature fusion on the feature maps from large scale to small scale to obtain a fused feature map; The prediction module is used to determine the prediction box where the object is located in the fused feature map.

9. A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the multi-modal object detection method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, It includes a processor and a memory, and one or more computer programs are stored in the memory. The one or more computer programs include instructions which, when executed by the electronic device, cause the electronic device to execute the multimodal object detection method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method and system for detecting nonferrous metal target of scraped car

    CN120953758A

  • A method and system for detecting non-ferrous metal targets in scrapped automobiles

    CN120953758B