Method and apparatus for detecting multi-modal object using cross-guide attention

The multi-modal object detection method employing cross-guided attention effectively addresses the limitations of single-image object detection by generating cross-attention maps that combine RGB and infrared image features, resulting in improved detection accuracy and robustness across different environments.

WO2025110496A1PCT designated stage expired Publication Date: 2025-05-30PUSAN NAT UNIV IND UNIV COOPERATION FOUND
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016150
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-10-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing object detection techniques using single RGB images are vulnerable to noise and lighting changes, while infrared images offer stability but suffer from low resolution and accuracy issues due to distance. Multi-modal approaches combining RGB and infrared images face performance degradation due to inadequate reflection of feature differences.

Method used

A multi-modal object detection method using cross-guided attention, which involves receiving RGB and infrared images, extracting feature maps, generating cross-attention maps based on correlations between feature maps, and synthesizing these maps to create a combined feature map for improved object detection.

Benefits of technology

This approach enhances object detection performance by accurately reflecting correlations between RGB and infrared feature maps, leading to robust detection even in adverse weather conditions or at varying distances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016150_30052025_PF_FP_ABST
    Figure KR2024016150_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method and an apparatus for detecting a multi-modal object using cross-guide attention. The method for detecting a multi-modal object using cross-guide attention according to one aspect of the present invention is an object detection method performed by an electronic device, and comprises the steps of: receiving an RGB image and an infrared image; extracting an RGB feature map for the RGB image and an infrared feature map for the infrared image using a neural network model, generating a cross-attention map of each of the RGB feature map and the infrared feature map by means of a cross-attention operation, and generating a combined feature map by synthesizing and combining same with each of the RGB feature map and the infrared feature map; and detecting an object in the combined feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for multimodal object detection using cross-guided attention

[0001] The present invention relates to a technology for detecting objects in an image, and more particularly, to a multi-modal object detection technology using RGB images and infrared images.

[0002] Object detection is utilized in various technological fields, including autonomous driving. Accurately recognizing and responding to the surrounding environment is essential for the safety of autonomous driving. Accurate object detection is crucial for enhancing the safety of autonomous driving.

[0003] Existing single-image object detection techniques, such as those using RGB images, suffer from performance fluctuations due to noise and lighting changes, making them vulnerable to environmental changes. Infrared images are less sensitive to lighting changes than RGB images, allowing for accurate object detection even in changing environments. However, they suffer from low resolution and temperature, which affects the distance from the object, causing accuracy to deteriorate with distance.

[0004] In this regard, a multi-modal object detection technology combining RGB and infrared images has been disclosed. Existing techniques that simply combine single-image detection techniques into a two-stream structure suffer from the problem of poorly reflecting the characteristics of RGB and infrared images, resulting in lower performance compared to object detection using a single image.

[0005] Recently proposed cross-modal object detection techniques that incorporate the importance of RGB and infrared images through attention mechanisms offer superior performance compared to conventional single-image object detection methods. However, these techniques simply link features from RGB and infrared images using a self-attention mechanism, limiting the interaction between RGB and infrared images in the calculation of attention maps.

[0006] The purpose of the present invention is to solve the above problem, and to improve object detection performance by combining RGB images and infrared images by generating a cross attention map by reflecting weights according to correlations between feature maps for each of the RGB images and infrared images.

[0007] The purpose of the present invention is not limited to the purposes mentioned above, and other purposes not mentioned can be clearly understood from the description below.

[0008] A multi-modal object detection method using cross-guided attention according to one aspect of the present invention for achieving the above-described object object detection method is performed by an electronic device, comprising the steps of: receiving an RGB image and an infrared image as input; extracting an RGB feature map for the RGB image and an infrared feature map for the infrared image using a neural network model; generating cross-attention maps of each of the RGB feature map and the infrared feature map through a cross-attention operation; generating a combined feature map by synthesizing and combining the RGB feature map and the infrared feature map, respectively; and detecting an object in the combined feature map.

[0009] A multi-modal object detection device using cross-guided attention according to another aspect of the present invention includes a memory for storing commands, and a processor for extracting an RGB feature map for an RGB image and an infrared feature map for an infrared image by processing an RGB image and an infrared image input from the outside in a plurality of input locations in the input order by executing the commands stored in the memory, generating cross-attention maps of each of the RGB feature maps and the infrared feature maps through a cross-attention operation, generating a combined feature map by synthesizing and combining the RGB feature map and the infrared feature map, and implementing a neural network model for detecting an object from the combined feature map.

[0010] According to the present invention, there is an effect of improving object detection performance by generating a combined feature map that clearly reflects the correlation between feature maps for RGB images and infrared images.

[0011] Therefore, it has the advantage of enabling accurate object detection even in environments with poor weather conditions such as snow or rain, or in noisy environments with poor lighting at night.

[0012] Additionally, it has the effect of enabling object detection with high performance regardless of the distance to the object.

[0013] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description of the claims.

[0014] FIG. 1 is a block diagram of a multi-modal object detection device using cross-guide attention according to one embodiment of the present invention.

[0015] FIG. 2 is a flowchart of a multi-modal object detection method using cross-guide attention according to another embodiment of the present invention.

[0016] FIG. 3 is a diagram showing the structure of a neural network model implemented according to embodiments of the present invention.

[0017] FIG. 4 is a diagram showing a detailed structure of a cross-guide attention block included in a neural network model implemented according to embodiments of the present invention.

[0018] Figure 5 is a drawing showing the results of detecting an object using a conventional object detection technology and an embodiment of the present invention, respectively.

[0019] A multi-modal object detection method using cross-guided attention according to one aspect of the present invention for achieving the above-described object object detection method is performed by an electronic device, and includes the steps of: receiving an RGB image and an infrared image as input; extracting an RGB feature map for the RGB image and an infrared feature map for the infrared image using a neural network model; generating cross-attention maps of each of the RGB feature map and the infrared feature map through a cross-attention operation; generating a combined feature map by synthesizing and combining the RGB feature map and the infrared feature map, respectively; and detecting an object in the combined feature map.

[0020] The step of generating the combined feature map includes the steps of extracting an RGB feature map for the RGB image and an infrared feature map for the infrared image by inputting the RGB image and the infrared image into a neural network model, performing a cross-attention operation on the RGB feature map and the infrared feature map based on a weight calculated from the infrared feature map and a weight calculated from the RGB feature map, respectively, thereby generating an RGB cross-attention map and an infrared cross-attention map, and generating a combined feature map by combining feature maps calculated from a result of logical combination of the RGB feature map and the RGB cross-attention map, and the infrared feature map and the infrared cross-attention map.

[0021] The above neural network model has a symmetrical structure and includes a pair of cross-attention modules that receive input for a base and input for a guide, apply a query determined from the base and a key and value derived from the guide to a preset attention operation formula, and produce a cross-attention map.

[0022] The step of calculating the RGB cross attention map and the infrared cross attention map is performed by inputting the RGB feature map as a base and the infrared feature map as a guide to the pair of cross attention modules, respectively, and calculating the RGB cross attention map guided by the infrared feature map and the infrared cross attention map guided by the RGB feature map.

[0023] A multi-modal object detection device using cross-guided attention according to another aspect of the present invention comprises: a memory for storing commands; and a processor for executing the commands stored in the memory, thereby extracting an RGB feature map for the RGB image and an infrared feature map for the infrared image by processing an RGB image and an infrared image input from the outside in the input order at each of a plurality of input locations, generating cross-attention maps of each of the RGB feature map and the infrared feature map through a cross-attention operation, generating a combined feature map by synthesizing and combining the RGB feature map and the infrared feature map, and implementing a neural network model for detecting an object from the combined feature map.

[0024] The neural network model includes a plurality of encoder blocks that encode the RGB image and the infrared image in the input order and output a feature map for the input, a cross-guide attention block that receives the RGB feature map and the infrared feature map output from the encoder block for the RGB image and the infrared image, respectively, and performs a cross-attention operation based on a weight calculated from each other to produce an RGB cross-attention map and an infrared cross-attention map, and a combining block that combines feature maps calculated from the result of the logical combination of the RGB feature map and the RGB cross-attention map, the infrared feature map and the infrared cross-attention map, to generate a combined feature map.

[0025] The above cross-guide attention block has a symmetrical structure and receives one of the RGB feature map and the infrared feature map extracted from the encoder block as a base and the other as a guide, and applies a query determined from the base and a key and value derived from the guide to a preset attention operation formula to produce a cross-attention map, including a pair of cross-attention models, to produce an RGB cross-attention map and an infrared cross-attention map.

[0026] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below, but may be implemented in various different forms. These embodiments are provided only to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the description of the claims. Meanwhile, the terminology used in this specification is for the purpose of describing the embodiments and is not intended to limit the present invention. In this specification, the singular includes the plural unless specifically stated otherwise.

[0027] The present invention relates to a multi-modal object detection technology for detecting objects from data combining RGB images and infrared images.

[0028] In particular, the present invention is characterized by a technical feature of improving object detection performance through a combined feature map by generating a combined feature map by combining an RGB image and an infrared image by calculating a cross-attention map according to the correlation between an RGB image and an infrared image through a cross-attention operation and sequentially transmitting the combined feature map to a lower layer through a hierarchical structure.

[0029] These technical features can be achieved by a configuration in which feature maps extracted from each image, including RGB images and infrared images, are subjected to attention operations using weights based on feature maps extracted from other images, and cross-attention maps reflecting different forms of correlation are generated and used to combine the feature maps for each image.

[0030] Referring to the attached drawings below, a multi-modal object detection method and device using cross-guide attention according to an embodiment of the present invention will be described in detail.

[0031] A multi-modal object detection method using cross-guide attention according to another embodiment of the present invention can be performed by a multi-modal object detection device using cross-guide attention according to one embodiment of the present invention.

[0032] For the convenience of the following explanation, the drawing symbols are matched for functionally identical contents and duplicate explanations are avoided.

[0033] Referring to FIG. 1, a multi-modal object detection device (10) using cross-guide attention may include a memory (11), a processor (12), an input / output interface (13) connected to an input / output device, and a communication interface (14) for communicating with an external network.

[0034] Here, the input / output device is for receiving user input and outputting results according to the user input, and may include, but is not limited to, a mouse, keyboard, touch display, or display.

[0035] The memory (11), processor (12), input / output interface (13), and communication interface (14) can be connected to each other through a communication bus.

[0036] The memory (11) can store instructions for performing a multi-modal object detection method using cross-guide attention according to another embodiment of the present invention.

[0037] The processor (12) can execute commands stored in the memory (11).

[0038] The processor (12) can receive an RGB image and an infrared image from a user (S100) and generate a combined feature map by combining the RGB image and the infrared image input from the user (S200) by executing commands stored in the memory (11).

[0039] The processor (12) extracts an RGB feature map for the RGB image and an infrared feature map for the infrared image by processing the RGB image and the infrared image input from the user in the input order at each of a plurality of input locations, generates a cross-attention map of each of the RGB feature map and the infrared feature map through a cross-attention operation, and generates a combined feature map by synthesizing and combining the RGB feature map and the infrared feature map, respectively (S200).

[0040] A neural network model implemented by executing commands stored in a memory (11) in a processor (12) may include a plurality of encoder blocks, at least one cross-guide attention block, a plurality of logical OR operators, and a combination block.

[0041] The processor (12) encodes an RGB image and an infrared image through an encoder block of a neural network model, extracts an RGB feature map for the RGB image and an infrared feature map for the infrared image, and performs a cross-attention operation based on a weight calculated from the infrared feature map and a weight calculated from the RGB feature map on the RGB feature map and the infrared feature map, respectively, thereby calculating an RGB cross-attention map and an infrared cross-attention map, and can generate a combined feature map by combining feature maps calculated from the logical combination of the RGB feature map and the RGB cross-attention map, and the infrared feature map and the infrared cross-attention map (S200).

[0042] Specifically, referring to FIG. 3, a neural network model (1000) implemented as a processor (12) executes commands stored in a memory (11) may include a first encoder block (111, 121), a second encoder block (112, 122), a third encoder block (113, 123), a fourth encoder block (114, 124), a first cross-guide attention block (210), a second cross-guide attention block (220), a third cross-guide attention block (230), a third cross-guide attention block (230), a first logical OR operator (311, 321), a second logical OR operator (312, 322), a third logical OR operator (313, 323), and a combination block (400).

[0043] The first encoder block (111, 121) encodes the RGB image and the infrared image, respectively, and generates the first RGB feature map for the RGB image ( ) and the first infrared feature map for the infrared image ( ) may be extracted.

[0044] The first cross-guide attention block (210) outputs the first RGB feature map ( ) and the first infrared feature map ( ) may be input and a cross-attention operation may be performed based on the weights derived from each other to generate a first RGB cross-attention map and a first infrared cross-attention map.

[0045] The first cross-guide attention block (210), the second cross-guide attention block (220), and the third cross-guide attention block (230) can receive RGB feature maps and infrared feature maps from the first encoder block (111, 121), the second encoder block (112, 122), and the third encoder block (113, 123), respectively, and generate an RGB cross-attention map and an infrared cross-attention map.

[0046] At this time, the internal configuration of the first cross guide attention block (210), the second cross guide attention block (220), and the third cross guide attention block (230) can be configured as shown in FIG. 4.

[0047] Referring to FIG. 4, the cross-guide attention block may include a pair of cross-attention models (201, 202) that have a symmetrical structure and receive one of an RGB feature map and an infrared feature map input from the outside as a base and the other as a guide, and apply a query determined from the base and a key and a value derived from the guide to a preset attention operation formula to generate a cross-attention map.

[0048] Here, the query may be derived by performing layer normalization and window partitioning on the feature map input as the base, and then linearly projecting it.

[0049] Additionally, the keys and values ​​can be derived by performing Layer Normalization and Window Partitioning on the feature map input as a guide and then performing Linear Projection on it.

[0050] Each cross-attention model (201, 202) may receive one of the RGB feature map and the infrared feature map as a base and the other as a guide, and may use a query determined from the base and a key and value derived from the guide to multiply the query by the transpose matrix of the key, divide the value by the dimension of the key, add a preset bias value to the result, input the result into a softmax function to produce a weight, and multiply the value by the produced weight to produce a cross-attention map of the base in which parts similar to the guide are emphasized.

[0051] The first cross-attention model (201) receives an RGB feature map as a guide based on an infrared feature map and performs a query ( ) and determine the key ( from the guide RGB feature map) ) and value( ) and extract the infrared cross attention map from RGB to infrared as shown in the mathematical formula below. ) can be produced.

[0052]

[0053] Here, is a query determined from an infrared feature map, and Keys and values ​​extracted from the RGB feature map, is the dimension of the key, refers to the bias value among the preset learning parameters.

[0054] The second cross-attention model (202) receives an infrared feature map as a guide based on an RGB feature map and performs a query based on the RGB feature map as a base. ) is determined, and the key () is determined from the infrared feature map as a guide. ) and value( ) and extract the RGB cross attention map from infrared to RGB as shown in the mathematical formula below. ) can be produced.

[0055]

[0056] Here, is a query determined from the RGB feature map, and Keys and values ​​extracted from the infrared feature map, is the dimension of the key, refers to the bias value among the preset learning parameters.

[0057] The cross-guide attention block is composed of a pair of cross-attention models (201, 202) that have a symmetrical structure and receive one of the RGB feature map and the infrared feature map as a base and the other as a guide, and multiply each feature map by a weight according to the similarity with the other feature map to produce a cross-attention map, so that a cross-attention map in which the similar parts with the other feature map are emphasized can be generated for each of the RGB feature map and the infrared feature map.

[0058] That is, the cross-guide attention block according to the present invention reflects the similarity with the infrared feature map when calculating the attention map of the RGB feature map, and reflects the similarity with the RGB feature map when calculating the attention map of the infrared feature map, and can generate two cross-attention maps, one for each of the RGB feature map and the other for each of the infrared feature maps as a guide.

[0059] Afterwards, the cross attention map generated based on the infrared feature map is combined with the infrared feature map, and the combined features can be combined with the processed features through additional layered normalization and feed-forward processing.

[0060] Additionally, the cross attention map generated based on the RGB feature map is combined with the RGB feature map, and the combined features can be combined with processed features through additional layered normalization and feed-forward processing.

[0061] According to the present invention, rather than independently applying the attention mechanism to the RGB feature map and the infrared feature map like the existing self-attention mechanism, the attention mechanism is cross-applied using feature maps of different formats as a guide to generate a cross-attention map.

[0062] Accordingly, there is an advantage of generating a cross-attention map that improves the problem of difficulty in reflecting correlations between different modalities (RGB, infrared) in the existing self-attention mechanism by reflecting correlations with feature maps of different modalities.

[0063]

[0064] The first logical OR operator (311, 321) can receive the first RGB feature map and the first infrared feature map extracted from the first encoder block (111, 121) between the first encoder block (111, 121) and the second encoder block (112, 122), and the first RGB cross attention map and the first infrared cross attention map output from the cross guide attention block (200) for the first RGB feature map and the first infrared feature map.

[0065] The first logical OR operator (311, 321) can output the result of logical ORing the first RGB feature map and the first RGB cross attention map and the result of logical ORing the first infrared feature map and the first infrared cross attention map as inputs to the second encoder block (112, 122) that is configured as a pair corresponding to the RGB image and the infrared image.

[0066] The second encoder block (112, 122) encodes the result of logically combining the first RGB feature map and the first RGB cross attention map output from the first logically combining operator (311, 321) and the result of logically combining the first infrared feature map and the first infrared cross attention map to generate the second RGB feature map ( ) and the second infrared feature map ( ) can be extracted.

[0067] The second cross-guide attention block (220) outputs the second RGB feature map ( ) and the second infrared feature map ( ) may be input and a cross-attention operation may be performed based on the weights derived from each other, thereby producing a second RGB cross-attention map and a second infrared cross-attention map.

[0068] The second cross-guide attention block (220) may combine the first RGB cross-attention map and the first infrared cross-attention map output from the first cross-guide attention block (210) into the second RGB feature map and the second infrared feature map output from the second encoder block (112, 122), respectively, and input one of them as a base and the other as a guide into the internal cross-attention model (201, 202), thereby producing the second RGB cross-attention map and the second infrared cross-attention map.

[0069] The second cross-guide attention block (220) receives the first RGB cross-attention map and the first infrared cross-attention map as input, down-samples them to have the same size as the output of the second encoder block (112, 122), adds them to the second RGB feature map and the second infrared feature map output from the second encoder block (112, 122), and inputs the resulting output as a base or guide to the cross-attention model (201, 202) to produce the second RGB cross-attention map and the second infrared cross-attention map.

[0070] The second logical OR operator (312, 322) can receive the second RGB feature map and the second infrared feature map extracted from the second encoder block (112, 122) between the second encoder block (112, 122) and the third encoder block (113, 123), and the second RGB cross attention map and the second infrared cross attention map output from the cross guide attention block (200) for the second RGB feature map and the second infrared feature map.

[0071] The second logical OR operator (312, 322) can output the result of logical ORing the second RGB feature map and the second RGB cross attention map and the result of logical ORing the second infrared feature map and the second infrared cross attention map as inputs to the third encoder block (113, 123) that is configured as a pair corresponding to the RGB image and the infrared image.

[0072] The third encoder block (113, 123) encodes the result of logically combining the second RGB feature map and the second RGB cross attention map output from the second logically combining operator (312, 322) and the result of logically combining the second infrared feature map and the second infrared cross attention map to generate the third RGB feature map ( ) and third infrared feature map ( ) can be extracted.

[0073] The third cross-guide attention block (230) outputs the third RGB feature map ( ) from the third encoder block (113, 123). ) and third infrared feature map ( ) may be input and a cross-attention operation may be performed based on the weights derived from each other, thereby producing a third RGB cross-attention map and a third infrared cross-attention map.

[0074] The third cross-guide attention block (230) receives the second RGB cross-attention map and the second infrared cross-attention map generated from the second cross-guide attention block (220), down-samples them to have the same size as the output of the third encoder block (113, 123), and adds the results to the third RGB feature map and the third infrared feature map output from the third encoder block (113, 123), and inputs the results as a base or guide to the cross-attention model (201, 202) to produce the third RGB cross-attention map and the third infrared cross-attention map.

[0075] After the first cross-guide attention block (210), the feature maps for RGB input as base and guide to the cross-guide attention blocks (220, 230) ) and feature maps for infrared can be expressed as the following mathematical formula.

[0076]

[0077] Here, i={2,3}, is the RGB feature map input from the i-th encoder block according to the input order, is a first scale feature map that downsamples the size of the RGB cross-attention map output from the i-1th cross-guide attention block to be the same as the output of the i-th encoder block, is the infrared feature map input from the i-th encoder block, is a second scale feature map that downsamples the size of the infrared cross attention map output from the i-1th cross guide attention block to be the same as the output of the i-th encoder block.

[0078] That is, a plurality of cross-guide attention blocks (210, 220, 230) are connected in a hierarchical structure to apply a cross-attention map received from an upper layer to an RGB feature map and an infrared feature map input from the outside, and to produce a cross-attention map from the RGB feature map and the infrared feature map.

[0079] According to the present invention, a cross-guide attention block generates a cross-attention map by interacting feature maps for images of different formats, and a plurality of cross-attention maps are hierarchically connected so that a cross-attention map received in a previous input sequence is further reflected in the current cross-attention generation, thereby enhancing the object detection effect.

[0080] The third logical sum operation unit (313, 323) can receive the third RGB feature map and the third infrared feature map extracted from the third encoder block (113, 123) between the third encoder block (113, 123) and the fourth encoder block (114, 124), and the third RGB cross attention map and the third infrared cross attention map output from the cross guide attention block (200) for the third RGB feature map and the third infrared feature map.

[0081] The third logical OR operator (313, 323) can output the result of logical ORing the third RGB feature map and the third RGB cross attention map and the result of logical ORing the third infrared feature map and the third infrared cross attention map as inputs to the fourth encoder block (114, 124) that is configured as a pair corresponding to the RGB image and the infrared image.

[0082] The fourth encoder block (114, 124) encodes the result of logically combining the third RGB feature map and the third RGB cross attention map output from the third logical operator (313, 323) and the result of logically combining the third infrared feature map and the third infrared cross attention map to generate the fourth RGB feature map ( ) and the fourth infrared feature map ( ) can be extracted.

[0083] The combination block (400) outputs the second RGB feature map ( ) from the second encoder block (112, 122). ) and the second infrared feature map ( ), the third RGB feature map output from the third encoder block (113, 123) ) and third infrared feature map ( ), the 4th RGB feature map output from the 4th encoder block (114, 124) ) and the fourth infrared feature map ( ) can be combined to create a combined feature map.

[0084] Thereafter, the processor (12) can detect an object from the combined feature map generated through the combined block (400) (S300).

[0085] According to the present invention, attention maps are generated from feature maps for each of an RGB image and an infrared image through a cross-guide attention block (200) of a symmetrical structure based on an attention mechanism, and the attention maps are synthesized into a feature map and encoded and combined, thereby enabling accurate object detection even in images with a lot of noise due to darkness or bad weather conditions.

[0086] In particular, the present invention is characterized in that it strengthens the interaction between the RGB image and the infrared image in the generation of a combined feature map by generating two cross attention maps using inputs of different forms as a base and a guide by applying a weight based on the similarity between the infrared feature map and the RGB feature map to the RGB feature map, and having a structure that is symmetrical to each other based on the existing self-attention mechanism, and the other applying a weight based on the similarity between the infrared feature map and the RGB feature map to the RGB feature map.

[0087] That is, according to the present invention, in a cross-modal object detection technology that detects an object by combining features of an RGB image and an infrared image, a feature map in which a portion with high similarity is further strengthened according to the correlation between the RGB feature map for the RGB image and the infrared feature map for the infrared image is combined.

[0088] In addition, multiple cross-guide attention blocks that enhance interaction are provided in a hierarchical structure so that the previously produced cross-attention map is reflected in the current cross-attention map production, which can contribute to improving object detection performance.

[0089] Accordingly, there is an advantage of improving object detection accuracy compared to existing cross-modal object detection techniques that individually reflect the attention maps of RGB images and infrared images or produce the attention map as a single layer.

[0090]

[0091] Hereinafter, an experiment was conducted to confirm that object detection performance is improved compared to existing object detection techniques when object detection is performed according to an embodiment of the present invention.

[0092] For the experiment, the neural network model according to the embodiment of the present invention was trained for 200 epochs using the SGD optimizer on a system with four NVIDIA A100 GPUs and the PyTorch framework. At this time, the learning rate was The batch size was set to 16.

[0093] The table below shows the evaluation indices derived from performing object detection using the FLIR-aligned dataset using existing object detection technologies such as YOLOv5, two-stream YOLOv5, +CFT, and a neural network model according to an embodiment of the present invention.

[0094] As evaluation indicators, mAP50, which is the average of the average precision at IoU (Interestion Over Union) = 0.50, mAP75, which is the average of the average precision at IoU (Interestion Over Union) = 0.75, and mAP (mean Average Precision) were used.

[0095] modality method mAP50 mAP75 mAPRGB YOLO v56 7.8 25.93 1.8 infrared YOLO v57 3.9 35.73 9.5 RGB+infrared two-stream YOLO v57 3.03 2.03 7.4 RGB+infrared CFT 78.7 35.54 0.2 RGB+infrared invention 79.3 38.54 2.1

[0096] Referring to the table above, it can be seen that two-stream YOLOv5, which receives both RGB images and infrared images in a two-stream structure, has lower performance in mAP50 and mAP overall compared to YOLOv5, which detects objects with a single image, when infrared images are used. Meanwhile, the existing cross-modal object detection technology (CFT) shows higher performance in mAP50 and mAP than the existing YOLOv5 and two-stream YOLOv5 technology, but according to the present invention, it can be seen that it shows a performance that is 0.6 higher in mAP50 and 1.9 higher in mAP than the existing +CFT.

[0097] The table below shows the evaluation indices derived from performing object detection using the LLVIP dataset using existing object detection technologies such as YOLOv5, two-stream YOLOv5, CFT, and a neural network model according to an embodiment of the present invention.

[0098] As an evaluation index, mAP75, which is the average of the average precision at IoU=0.75, was used.

[0099] modality method mAP75 mAPRGBYOLOv551.950.0 infraredYOLOv572.261.9 RGB+infraredtwo-streamYOLOv571.462.3 RGB+infraredCFT72.963.6 RGB+infraredThis invention75.465.1

[0100] Referring to the table above, when performing object detection according to the present invention, it can be confirmed that the performance is further improved as mAP75 and mAP are each 2.5 higher than CFT, which shows the best performance among existing technologies. In addition, the table below shows the recall derived from the results of object detection based on the previously disclosed MLPD (multi-label pedestrian detector), ARCNN, MBNet, ProbEn, and the present invention using the KAIST multi-spectral pedestrian dataset.

[0101] Method RecallMLPD96.70ARCNN97.25MBNet98.42ProbEn98.90Invention99.04

[0102] Referring to the table above, it can be seen that although existing methods also provide recall rates exceeding 96, the recall rate is the highest when objects are detected according to the present invention. In other words, it can be seen that object detection performance is improved compared to existing techniques when object detection is performed according to an embodiment of the present invention.

[0103] Figure 5 shows the results of object detection using RGB images (a), object detection using a conventional single-image object detection technique (YOLOv5), object detection using cross-modal object detection technique (CFT), and object detection using the present invention (c) for RGB-IR pair data provided in the KAIST multispectral pedestrian dataset.

[0104] Here, the green line represents pedestrians, the yellow line represents cars, and the red line represents bicycles.

[0105] Referring to FIG. 5, it can be confirmed that when the present invention is used, pedestrians, cars, and bicycles that were not detected by existing object detection technology are further detected, and objects that were not distinguishable by existing object detection technology are successfully detected.

[0106] That is, according to the present invention, by interacting with the feature maps of the RGB image and the infrared image to extract features multiple times in the input order and combine them to perform object detection, object detection performance can be improved compared to existing object detection technologies.

[0107] Meanwhile, the blocks in the attached block diagram and the steps in the flowchart may be implemented as computer instructions that are loaded into the processor or memory of an electronic device capable of data processing (e.g., a general-purpose computer, a special-purpose computer, a portable laptop computer, a network computer) and perform designated functions. Since these computer program instructions can be stored in a computer-readable memory, the functions described in the blocks in the block diagram or the steps in the flowchart may also be produced as a product that includes instruction means for performing them.

[0108] Those skilled in the art will appreciate that the present invention can be implemented in other specific forms without altering its technical spirit or essential characteristics. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims below rather than the detailed description above, and all changes or modifications derived from the claims and their equivalents should be construed as being included within the scope of the present invention.

[0109] The present invention relates to a technology for detecting objects in an image, and more particularly, to a multi-modal object detection technology using RGB images and infrared images.

Claims

1. A method for detecting an object performed by an electronic device, Step of receiving RGB images and infrared images; A step of extracting an RGB feature map for the RGB image and an infrared feature map for the infrared image using a neural network model, generating cross-attention maps of each of the RGB feature maps and the infrared feature maps through a cross-attention operation, and generating a combined feature map by synthesizing and combining the RGB feature maps and the infrared feature maps respectively; and A multi-modal object detection method using cross-guided attention, comprising: a step of detecting an object in the above combined feature map.

2. In paragraph 1, The step of generating the above combined feature map is A step of extracting an RGB feature map for the RGB image and an infrared feature map for the infrared image by inputting the RGB image and the infrared image into a neural network model; A step of generating an RGB cross-attention map and an infrared cross-attention map by performing a cross-attention operation based on the weights calculated from the infrared feature map and the weights calculated from the RGB feature map, respectively; A step of generating a combined feature map by combining feature maps derived from the logical combination of the RGB feature map and the RGB cross attention map, the infrared feature map and the infrared cross attention map; including A multi-modal object detection method using cross-guided attention.

3. In paragraph 2, The above neural network model It includes a pair of cross-attention modules that have a symmetrical structure and receive inputs for the base and the guide, apply the query determined from the base and the key and value derived from the guide to a preset attention operation formula to produce a cross-attention map. The step of calculating the RGB cross attention map and the infrared cross attention map is By inputting the RGB feature map as a base and the infrared feature map as a guide to each of the pair of cross-attention modules, and by inputting the infrared feature map as a base and the RGB feature map as a guide, an RGB cross-attention map guided by the infrared feature map and an infrared cross-attention map guided by the RGB feature map are generated. A multi-modal object detection method using cross-guided attention.

4. Memory for storing commands; and A multi-modal object detection device using cross-guided attention, comprising: a processor for extracting an RGB feature map for the RGB image and an infrared feature map for the infrared image by processing an RGB image and an infrared image input from the outside in the input order at each of a plurality of input locations by executing commands stored in the memory; generating cross-attention maps of each of the RGB feature maps and the infrared feature maps through a cross-attention operation; generating a combined feature map by synthesizing and combining the RGB feature maps and the infrared feature maps respectively; and implementing a neural network model for detecting an object from the combined feature map.

5. In paragraph 4, The above neural network model A plurality of encoder blocks each encoding the RGB image and the infrared image in the input order and outputting a feature map for the input, A cross-guide attention block that receives the RGB feature map and the infrared feature map output for each of the RGB image and the infrared image from the encoder block, performs a cross-attention operation based on the weights calculated from each other, and thereby calculates an RGB cross-attention map and an infrared cross-attention map, A combination block that generates a combined feature map by combining feature maps derived from the logical combination of the RGB feature map and the RGB cross attention map, and the infrared feature map and the infrared cross attention map. A multi-modal object detection device using cross-guided attention.

6. In paragraph 5, The above cross-guide attention block A method of producing an RGB cross-attention map and an infrared cross-attention map, including a pair of cross-attention models that have symmetrical structures and take one of the RGB feature maps and the infrared feature maps extracted from the encoder block as a base and the other as a guide, and apply a query determined from the base and a key and value derived from the guide to a preset attention operation formula to produce a cross-attention map. A multi-modal object detection device using cross-guided attention.

Citation Information

Patent Citations

  • Method and device for identifying object from video

    CN116030387A

  • Smartphone accessories having fragrance capsules and manufacturing method thereof

    KR1020240173259A

  • Apparatus and method for fusing visible light image and infrared image based on multi-scale network

    KR102565989B1

  • Visible light and infrared fusion image-based object detection method and apparatus

    WO2020171281A1