Cross-Domain Information Fusion Target Detection Method and Device
By registering modal A and modal B images and processing of dual-stream object detection models, multimodal interaction modules are used to generate multimodal features, solving the problem of poor object detection stability in complex scenarios, and achieving high-accurate object detection effect.
Patent Information
- Application Number
- CN202411907227.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The existing target detection methods have poor detection stability in complex scenarios such as heavy fog, heavy rain, and night. Especially in dark scenes, visible light images cannot provide sufficient foreground target information. The multimodal fusion method fails to fully explore the correlation between different modes, resulting in a large room for improvement in the detection results.
The object detection method of cross-domain information fusion is adopted, and the modal A image and modal B image are registered, and the dual-stream object detection model is input. The multimodal features are generated and object detection is carried out using modules such as feature extraction, multi-scale enhancement, spatial-level multi-modal interaction and channel-level multi-modal interaction.
Accurate detection of various targets in complex scenarios is achieved, and the accuracy and stability of detection results are improved through the complementary multimodal information and the filtering of redundant information.
Smart Images

Figure CN119360174B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a target detection method and device for cross-domain information fusion. Background Art
[0002] As a popular research direction in the field of computer vision, target detection has been widely applied in many fields such as intelligent driving, medical diagnosis, and security monitoring, greatly promoting the development of productivity. According to the candidate prior information, the current target detection methods can be divided into three types: anchor box-based target detection methods, anchor point-based target detection methods, and query-based target detection methods.
[0003] The anchor box-based target detection method usually consists of two steps. The first step is to generate a series of regions of interest, and the second step is to fine-tune the positions of the regions of interest and determine their categories. The anchor point-based target detection method is an improvement on the method based on anchor boxes with more manual components. The query-based target detection method is a more flexible detection method that does not rely on anchor points and anchor boxes and searches the entire image for the targets to be concerned through a predefined fixed number of queries.
[0004] However, most of the above-mentioned target detection methods are based on visible light images and it is difficult to ensure the stability of the detection system in various practical application scenarios such as heavy fog, heavy rain, and night. Especially in the night scene, the visible light image cannot provide sufficient foreground target information. Although technologies such as low-light enhancement are constantly developing, the detection based on single-light images alone is difficult to meet the application requirements in complex scenarios. Therefore, researchers have tried to alleviate this problem by introducing multi-modal fusion methods that incorporate images of other modalities.
[0005] However, due to the huge semantic information gap between different modalities, the existing multi-modal fusion methods have not fully exploited the correlation between different modalities, resulting in a large room for improvement in the target detection results of the existing multi-modal fusion methods. Summary of the Invention
[0006] The present invention is made to solve the above problems, and aims to provide a target detection method and device for cross-domain information fusion.
[0007] The present invention provides an object detection method for cross - domain information fusion, which is used to obtain the detection result of an object based on the modality A image and modality B image corresponding to the object, and has the following features, including the following steps: Step S1, register the modality A image and modality B image to obtain the registered modality A image and modality B image; Step S2, input the registered modality A image and registered modality B image into a two - stream object detection model to obtain the detection result. Among them, the two - stream object detection model includes: a feature extraction module, which extracts high - level semantic feature A from the registered modality A image and extracts high - level semantic feature B from the registered modality B image; a multi - scale enhancement module, which performs multi - scale enhancement on the feature pair composed of high - level semantic feature A and high - level semantic feature B to obtain an enhanced input feature pair; a fusion module, which performs spatial - level multi - modal interaction and channel - level multi - modal interaction on the enhanced input feature pair to obtain multi - modal features; an object detection module, which generates a detection box and the confidence of the corresponding category as the detection result according to the multi - modal features.
[0008] In the object detection method for cross - domain information fusion provided by the present invention, it may also have the following features: Among them, the feature extraction module includes a backbone network, and the calculation expressions of high - level semantic feature A and high - level semantic feature B are: , , where in the formula is high - level semantic feature A, is the modality A image, is high - level semantic feature B, is the modality B image.
[0009] In the object detection method for cross - domain information fusion provided by the present invention, it may also have the following features: Among them, in the multi - scale enhancement module, the feature pair is upsampled by multiple different multiples in the width and height dimensions respectively to obtain corresponding feature pairs of multiple different scales, and all feature pairs of different scales are expanded into sequences and spliced in the width and height dimensions to obtain an enhanced input feature pair.
[0010] In the object detection method for cross - domain information fusion provided by the present invention, it may further have the following features: Among them, the fusion module includes a plurality of sequentially connected high - order collaborative interaction units. The high - order collaborative interaction unit includes a spatial - level multimodal interaction subunit and a channel - level multimodal interaction subunit. The spatial - level multimodal interaction subunit performs a compromise split on the input in the channel dimension, and fuses the split results in the spatial dimension through a deformable attention mechanism and a residual connection to obtain the corresponding output. The channel - level multimodal interaction subunit sequentially performs channel - dimension concatenation, sequence - dimension split, three - dimensional reshaping, and channel - level attention processing on the input to obtain the corresponding output. In the high - order collaborative interaction unit, the output of the spatial - level multimodal interaction subunit of this high - order collaborative interaction unit serves as the input of the channel - level multimodal interaction subunit of this high - order collaborative interaction unit, and the output of this channel - level multimodal interaction subunit serves as the input of the spatial - level multimodal interaction subunit of the next high - order collaborative interaction unit. The input of the spatial - level multimodal interaction subunit of the first high - order collaborative interaction unit is the enhanced input feature pair.
[0011] In the object detection method for cross - domain information fusion provided by the present invention, it may further have the following features: Among them, the spatial - level multimodal interaction subunit splits the input into feature pairs with the same shape and feature pairs The spatial - level multimodal interaction subunit fuses the feature pairs and feature pairs in the spatial dimension to obtain the output and The calculation expressions are: , , where in the formula is the number of attention heads in the deformable attention mechanism, and are fully - connected layers respectively, is the number of different scales of the features corresponding to the enhanced input feature pair, is the number of sampling points of the th attention head, is the attention weight of the th attention head, is the sampling point corresponding to on the th attention head updated based on the offset .
[0012] In the object detection method for cross - domain information fusion provided by the present invention, it may further have the following features: Among them, the channel - level multimodal interaction subunit generates the output and for the input The calculation expression is: , , , , , , where is the channel-level concatenation operation, is the splitting operation, is the reshaping operation, is the number of different scales corresponding to the features in the enhanced input feature pair, feature height, feature width, feature coordinate point corresponding value in, is a learnable parameter.
[0013] In the object detection method for cross-domain information fusion provided by the present invention, it may also have the following features: Among them, the construction and training process of the dual-stream object detection model includes the following steps: the training data construction step, constructing a training data set according to existing multiple different-modal existing object images; the model construction step, constructing a dual-stream object detection model; the model training step, training the dual-stream object detection model according to the training data set to obtain a trained dual-stream object detection model. In the model training step, the loss function for training the dual-stream object detection model includes a classification loss and an object localization loss. The calculation expression of the classification loss is: , where is a regulation factor, is the predicted label generated by the dual-stream object detection model according to the image, is the true classification label corresponding to the image, is the calculation result of the classification loss. The calculation expression of the object localization loss is: , where is the predicted bounding box, is the predicted bounding box corresponding annotation bounding box, is the minimum bounding box enclosing the predicted bounding box and the annotation bounding box , is the predicted bounding box and the annotation bounding box intersection over union ratio between, is the calculation result of the object localization loss.
[0014] In the object detection method for cross-domain information fusion provided by the present invention, it may also have the following features: In the training data construction step, at each scenario, at least two images of different modalities of an existing object are collected at the same shooting position as the existing object images. All the existing object images corresponding to the same shooting position and the same existing object are paired pairwise to construct image pairs. The object category and position information of the corresponding existing object are respectively labeled for the two existing object images in the image pair, and the two object images are registered to eliminate misalignment.
[0015] The present invention also provides an object detection device for cross-domain information fusion, which is used to obtain the detection result of an object according to the modality A image and modality B image corresponding to the object, and has the following features, including: a registration unit for registering the modality A image and modality B image to obtain the registered modality A image and modality B image; a detection unit including a two-stream object detection model for inputting the registered modality A image and the registered modality B image into the two-stream object detection model to obtain the detection result. Among them, the two-stream object detection model includes: a feature extraction module for extracting high-level semantic feature A from the registered modality A image and extracting high-level semantic feature B from the registered modality B image; a multi-scale enhancement module for performing multi-scale enhancement on the feature pair composed of high-level semantic feature A and high-level semantic feature B to obtain an enhanced input feature pair; a fusion module for performing spatial-level multi-modal interaction and channel-level multi-modal interaction on the enhanced input feature pair to obtain multi-modal features; an object detection module including a detection head for generating a detection box and the confidence of the corresponding category as the detection result according to the multi-modal features.
[0016] According to the object detection method and device for cross-domain information fusion involved in the present invention, on the one hand, through the spatial-level multi-modal interaction sub-unit, the deformable attention mechanism is used to adaptively capture the complementary feature information of the two modalities at the spatial level; on the other hand, through the channel-level multi-modal interaction sub-unit, the channel-level attention block is used to model the channel interaction between the two different modalities, so as to realize the fusion of complementary information and the filtering of redundant information between multi-modalities. Therefore, the object detection method and device for cross-domain information fusion of the present invention can achieve accurate detection of various objects in complex scenarios. Description of the Drawings
[0017] Figure 1 It is a block diagram of the object detection device in the embodiment of the present invention.
[0018] Figure 2 It is a block diagram of the two-stream object detection model in the embodiment of the present invention.
[0019] Figure 3 It is a schematic flow chart of the construction and training of the two-stream object detection model in the embodiment of the present invention.
[0020] Figure 4 It is a schematic flowchart of the object detection method for cross - domain information fusion in the embodiments of the present invention. Specific embodiments
[0021] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the following embodiments will specifically elaborate on the object detection method and device for cross - domain information fusion of the present invention in conjunction with the accompanying drawings.
[0022] In this embodiment, an object detection device for cross - domain information fusion is provided, hereinafter referred to as the object detection device, which is used to obtain the detection result of the object according to the modality A image and modality B image corresponding to the object. In this embodiment, modality A and modality B are different modalities, and visible - light images, infrared imaging images, lidar imaging images, and synthetic aperture radar imaging images are all images of different modalities.
[0023] Figure 1 It is a block diagram of the object detection device in the embodiments of the present invention.
[0024] As Figure 1 shown, the object detection device 100 includes a registration unit 11 and a detection unit 12.
[0025] The registration unit 11 is used to register the modality A image and the modality B image to obtain the registered modality A image and modality B image.
[0026] Due to the difference in imaging principles, even if two different modality images corresponding to the object are continuously captured at the same shooting position, there are generally different degrees of misalignment between these two images. If there is offset or deformation between the image pair, then artifacts will inevitably appear in the subsequent fusion result. To address this problem, an adaptive symmetric registration algorithm is used in this embodiment to register the object images. Hereinafter, the infrared image and the visible - light image are used as the modality A image and the modality B image respectively to illustrate the registration process:
[0027] Input the infrared image into the trained infrared - to - visible - light register to obtain the registered infrared image, and input the visible - light image into the trained visible - light - to - infrared register to obtain the registered visible - light image.
[0028] Among them, when training the infrared - to - visible - light register, an affine transformation is performed on the infrared image to obtain a transformation matrix. Then, the visible - light image and the transformed infrared image are used as the input of this register, and the parameters of the transformation matrix are used as the annotation that this register wants to predict, so as to learn the infrared image before transformation. The purpose is that when there is misalignment between the input visible - light image and the infrared image, the infrared image can be corrected according to this register to obtain an infrared image that accurately corresponds to the position of the visible - light image.
[0029] When training a visible light to infrared registrator, an affine transformation is performed on the visible light image to obtain a transformation matrix. Then, the infrared image and the transformed visible light image are used as the input of the registrator, and the parameters of the transformation matrix are used as the annotation that the registrator wants to predict, so as to learn the visible light before transformation. The purpose is that when there is a misalignment between the input infrared image and the visible light image, the visible light image can be corrected according to the registrator to obtain a visible light image that accurately corresponds to the position of the infrared image.
[0030] The detection unit 12 includes a two-stream object detection model, which is used to input the registered modality A image and the registered modality B image into the two-stream object detection model to obtain a detection result.
[0031] Figure 2 It is a block diagram of the two-stream object detection model in the embodiment of the present invention.
[0032] As Figure 2 shown, the two-stream object detection model 200 includes a feature extraction module 21, a multi-scale enhancement module 22, a fusion module 23, and an object detection module 24.
[0033] The feature extraction module 21 includes two parallel backbone networks, which respectively extract high-level semantic features A from the registered modality A image and high-level semantic features B from the registered modality B image. In this embodiment, the size of the input image of the feature extraction module 21 is 1024. The feature extraction module 21 first processes the input image through a 16-fold downsampling layer to obtain a feature map of 64×64 size, and then passes through 24 ViT blocks to obtain the backbone features of different modalities, that is, high-level semantic features. Among them, the random dropout rate is 0.3, and the hidden layer dimension is set to 1024.
[0034] Among them, the calculation expressions of the high-level semantic feature A and the high-level semantic feature B are:
[0035] ,
[0036] ,
[0037] In the formula is the high-level semantic feature A, is the modality A image, is the high-level semantic feature B, is the modality B image. Among them, is the number of feature channels, and represent the height and width of the last layer.
[0038] The multi-scale enhancement module 22 performs multi-scale enhancement on the feature pair composed of the high-level semantic feature A and the high-level semantic feature B to obtain an enhanced input feature pair.
[0039] Among them, the multi-scale enhancement module 22 upsamples the feature pair by multiple different multiples in the width and height dimensions respectively to obtain corresponding feature pairs of multiple different scales, unfolds all the feature pairs of different scales into sequences in the width and height dimensions and splices them to obtain an enhanced input feature pair.
[0040] In this embodiment, the multiples include 2 times, 4 times, 8 times and 16 times. For the multiple multi-scale features obtained by upsampling, they are further strengthened through several convolutional layers respectively to obtain an enhanced input feature pair. Among them, the number of output channels is set to 256, layer normalization is used for regularization, and the GeLU activation function is selected for non-linear transformation. In this embodiment, the enhanced input feature pair is composed of the feature pair wherein, is the sequence dimension after splicing of the multi-scale features, when, the feature pair is the feature pair input to the multi-scale enhancement module 22, when, the feature pair is the feature pair formed by the multi-scale features obtained by upsampling by 2 times, 4 times, 8 times and 16 times.
[0041] The fusion module 23 performs spatial-level multi-modal interaction and channel-level multi-modal interaction on the enhanced input feature pair to obtain a multi-modal feature.
[0042] Among them, the fusion module 23 includes multiple sequentially connected high-order collaborative interaction units 231. In this embodiment, 3 high-order collaborative interaction units 231 are sequentially connected. The high-order collaborative interaction unit 231 includes a spatial-level multi-modal interaction subunit 2311 and a channel-level multi-modal interaction subunit 2312.
[0043] In the high-order collaborative interaction unit 231, the output of the spatial-level multi-modal interaction subunit 2311 of this high-order collaborative interaction unit 231 is used as the input of the channel-level multi-modal interaction subunit 2312 of this high-order collaborative interaction unit 231, and the output of the channel-level multi-modal interaction subunit 2312 is used as the input of the spatial-level multi-modal interaction subunit 2311 of the next high-order collaborative interaction unit 231. The input of the spatial-level multi-modal interaction subunit 2311 of the first high-order collaborative interaction unit 231 is the enhanced input feature pair.
[0044] The spatial-level multi-modal interaction subunit 2311 splits the input in the channel dimension and fuses the split results in the spatial dimension through a deformable attention mechanism and a residual connection to obtain the corresponding output.
[0045] Among them, the spatial-level multi-modal interaction subunit 2311 takes the input Split the compromise into feature pairs with the same shape and feature pairs . In this embodiment, the compromise is split according to the input dimension N, and the first N / 2 dimensions of the input are used as feature pairs , and the last N / 2 dimensions of the input are used as feature pairs .
[0046] The spatial multi-modal interaction subunit 2311 then fuses the feature pairs and feature pairs in the spatial dimension to obtain the output and . The calculation expression is:
[0047] ,
[0048] ,
[0049] ,
[0050] In the formula is the number of attention heads in the deformable attention mechanism, and are fully connected layers respectively, is the number of different scales corresponding to the features in the enhanced input feature pairs, is the th number of sampling points of the attention head, is the th attention weight of the attention head, is the corresponding sampling point on the th attention head updated based on the offset . In this embodiment is 5, that is, corresponding to the feature pair .
[0051] Existing spatial fusion strategies usually use Transformer for fusion. However, directly sampling Transformer for fusion of the long sequence features after multi-scale feature splicing will bring a heavy memory burden and a high time complexity, and will introduce redundant noise impurities of many invalid modes. The deformable attention mechanism adopted by the spatial multi-modal interaction subunit 2311 can effectively avoid the above problems.
[0052] The channel-level multi-modal interaction sub-unit 2312 performs channel dimension concatenation, sequence dimension splitting, three-dimensional reshaping, and channel-level attention processing on the input in sequence to obtain the corresponding output.
[0053] Among them, the channel-level multi-modal interaction sub-unit 2312 processes the input and to generate the output The calculation expression is as follows:
[0054] ,
[0055] ,
[0056] ,
[0057] ,
[0058] ,
[0059] ,
[0060] In the formula is the channel-level concatenation operation, is the splitting operation, is the reshaping operation, is the number of different scales corresponding to the features in the enhanced input feature pair, feature height, feature width, feature coordinate point corresponding value in, is the learnable parameter.
[0061] In this embodiment, the collaborative correlation between modality A and modality B is gradually captured through the iterative interaction of multiple spatial-level multi-modal interaction sub-units 2311 and channel-level multi-modal interaction sub-units 2312. Specifically, the spatial-level multi-modal interaction sub-unit 2311 integrates spatially fine-grained complementary information through the deformable attention mechanism. The channel-level multi-modal interaction sub-unit 2312 utilizes the global first-order statistics to distinguish the mutual dependence between modality A and modality B.
[0062] The target detection module 24 includes the encoder and decoder layers of Co-DETR, and generates detection boxes and the confidence of the corresponding categories as detection results based on the multi-modal features. In this embodiment, the multi-modal features are extracted through Co-DETR, and the detection boxes and the confidence of the categories are obtained through two multi-layer perceptrons respectively.
[0063] Figure 3 It is a schematic flow chart of the construction and training of a dual-stream object detection model in an embodiment of the present invention.
[0064] As Figure 3 shown, the construction and training process of the dual-stream object detection model 200 includes the following steps:
[0065] Step T1, namely the training data construction step, constructs a training data set according to multiple existing object images of different modalities.
[0066] Among them, at least two images of different modalities of the existing object are collected at the same shooting position in each scenario as the existing object images. All the existing object images corresponding to the same shooting position and the same existing object are paired pairwise to construct image pairs. The object categories and position information of the corresponding existing objects are respectively labeled for the two existing object images in the image pair, and the two object images are registered to eliminate misalignment.
[0067] In this embodiment, the shooting scenarios of the existing objects include but are not limited to complex scenarios such as occlusion, low light, fog, and rain. The object categories and position information of the existing objects are labeled by the Labelme annotation software, and the labeled data is converted into the COCO data set format type.
[0068] Step T2, namely the model construction step, constructs the dual-stream object detection model 200.
[0069] In this embodiment, step T1 is executed first and then step T2. In other embodiments, step T2 can be executed first and then step T1 according to needs, or step T1 and step T2 can be executed simultaneously.
[0070] Step T3, namely the model training step, trains the dual-stream object detection model 200 according to the training data set to obtain the trained dual-stream object detection model 200.
[0071] Among them, the loss function for training the dual-stream object detection model 200 includes classification loss and object localization loss.
[0072] The calculation expression of the classification loss is:
[0073] ,
[0074] In the formula is a regulation factor, is the predicted label generated by the dual-stream object detection model according to the image, is the true classification label corresponding to this image, is the calculation result of the classification loss.
[0075] The calculation expression of the object localization loss is:
[0076] ,
[0077] wherein is the predicted bounding box, is the predicted bounding box corresponding to the labeled bounding box, is the minimum bounding box wrapping the predicted bounding box and the labeled bounding box, is the predicted bounding box and the labeled bounding box the intersection over union ratio between them, is the calculation result of the target localization loss.
[0078] The following describes the process of the object detection method for cross - domain information fusion using the object detection method and apparatus 100 in conjunction with the accompanying drawings.
[0079] Figure 4 is a schematic flowchart of the object detection method for cross - domain information fusion in an embodiment of the present invention.
[0080] As Figure 4 shown, the object detection method for cross - domain information fusion includes the following steps:
[0081] Step S1: Use the registration unit 11 to register the modality A image and the modality B image to obtain the registered modality A image and modality B image.
[0082] Step S2: Use the detection unit 12 to input the registered modality A image and the registered modality B image into the two - stream object detection model to obtain the detection result.
[0083] Functions and effects of the embodiment
[0084] According to the object detection method and apparatus for cross - domain information fusion involved in this embodiment, on the one hand, the spatial - level multi - modal interaction sub - unit uses the deformable attention mechanism to adaptively capture the complementary feature information of the two modalities at the spatial level; on the other hand, the channel - level multi - modal interaction sub - unit uses the channel - level attention block to model the channel interaction between the two different modalities, thereby realizing the complementary information fusion and redundant information filtering between the multi - modalities. In short, this method can achieve accurate detection of various objects in complex scenarios.
[0085] Those skilled in the art should understand that the present invention is not limited by the above - mentioned embodiments. What is described in the above - mentioned embodiments and the specification only illustrates the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A target detection method using cross-domain information fusion, for obtaining a detection result of a target according to a modality A image and a modality B image corresponding to the target, characterized in that: The following steps are involved: Step S1, registering the modality A image and the modality B image to obtain a registered modality A image and modality B image; Step S2, inputting the registered modality A image and the registered modality B image into a dual-stream object detection model to obtain the detection result, Wherein, the dual-stream target detection model includes: A feature extraction module performs feature extraction on the registered modality A image to obtain a high-level semantic feature A, and performs feature extraction on the registered modality B image to obtain a high-level semantic feature B; A multi-scale enhancement module performs multi-scale enhancement on the feature pair consisting of the high-level semantic feature A and the high-level semantic feature B to obtain an enhanced input feature pair; A fusion module performs spatial-level multimodal interaction and channel-level multimodal interaction on the enhanced input feature pair to obtain multimodal features; The target detection module generates a detection frame and a corresponding category confidence as the detection result according to the multimodal feature. The construction and training process of the dual-stream target detection model includes the following steps: A training data construction step, constructing a training data set based on existing target images of multiple different modalities; Model building step: building a two-stream target detection model; The model training step is to train the dual-stream object detection model according to the training data set to obtain a trained dual-stream object detection model. In the model training step, the loss function of training the dual-stream target detection model includes classification loss and target positioning loss. The calculation expression of the classification loss is: QFL(σ)=-|y-σ| β ((1-y)log(1-σ)+ylog(σ)), Where |y-σ| β is the adjustment factor, σ is the predicted label generated by the two-stream target detection model based on the image, y is the true classification label corresponding to the image, QFL(σ) is the classification loss calculation result, The calculation expression of the target positioning loss is: Where |A| is the prediction box, |B| is the annotation box corresponding to the prediction box |A|, |C| is the minimum bounding box that wraps the prediction box |A| and the annotation box |B|, IOU is the intersection-over-union ratio between the prediction box |A| and the annotation box |B|, and L box Calculate the result for target localization loss.
2. The target detection method based on cross-domain information fusion according to claim 1 is characterized in that: in, The feature extraction module includes a ViT backbone network, The calculation expressions of the high-level semantic feature A and the high-level semantic feature B are: X modA =ViT(img modA ), X modB =ViT(img modB ), Where X modA is the high-level semantic feature A, img modA is the modality A image, X modB is the high-level semantic feature B, img modB is the mode B image.
3. The target detection method of cross-domain information fusion according to claim 1 is characterized in that: in, In the multi-scale enhancement module, the feature pairs are upsampled by multiple different multiples in width and height to obtain corresponding feature pairs of multiple different scales, and all the feature pairs of different scales are expanded into sequences in width and height and spliced to obtain the enhanced input feature pairs.
4. The target detection method of cross-domain information fusion according to claim 1 is characterized in that: in, The fusion module includes a plurality of sequentially connected high-order collaborative interaction units, The high-order collaborative interaction unit includes a space-level multimodal interaction subunit and a channel-level multimodal interaction subunit. The spatial multimodal interaction subunit splits the input in the channel dimension, and fuses the split results in the spatial dimension through a deformable attention mechanism and residual connection to obtain the corresponding output. The channel-level multimodal interaction subunit sequentially performs channel dimension splicing, sequence dimension splitting, three-dimensional reshaping and channel-level attention processing on the input to obtain the corresponding output. In the high-order collaborative interaction unit, the output of the spatial-level multimodal interaction subunit of the high-order collaborative interaction unit is used as the input of the channel-level multimodal interaction subunit of the high-order collaborative interaction unit, and the output of the channel-level multimodal interaction subunit is used as the input of the spatial-level multimodal interaction subunit of the next high-order collaborative interaction unit. The input of the spatial-level multimodal interaction subunit of the first high-order collaborative interaction unit is the enhanced input feature pair.
5. The target detection method of cross-domain information fusion according to claim 4 is characterized in that: in, The spatial multimodal interaction subunit takes the input (X modA ,X modB ) is split into feature pairs of the same shape and feature pairs The spatial-level multimodal interaction subunit is used to interact with the feature pair and the characteristic pair Fusion in the spatial dimension yields the output X modA and X modB The calculation expression is: Where M is the number of attention heads in the deformable attention mechanism, W m and W′ m are the fully connected layers, L is the number of different scales corresponding to the features in the enhanced input feature pair, and C N is the number of sampling points of the mth attention head, A m is the attention weight of the mth attention head, x(p q +Δp m ) corresponds to x k The sampling point of the mth attention head is based on the offset Δp m Updated sampling points.
6. The target detection method of cross-domain information fusion according to claim 4 is characterized in that: in, The channel-level multimodal interaction subunit is used for input X modA and X modB The calculation expression for generating output O is: X modAB =concat(X modA ,X modB ), O=concat(O 0 ,THE 1 ,THE 2 ,THE 3 ,THE 4 ), Where concat is the channel-level concatenation operation, split is the splitting operation, reshape is the reshaping operation, L is the number of different scales corresponding to the features in the enhanced input feature pair, and H is i Features Height, W i Features The width of Features The value corresponding to the coordinate point (x, y) in the middle. are learnable parameters.
7. The target detection method of cross-domain information fusion according to claim 1 is characterized in that: in, In the training data construction step, at least two images of different modalities are collected from the existing target at the same shooting position in each scene as the existing target image, Pair all the existing target images corresponding to the same shooting position and the same existing target in pairs to construct image pairs, The two existing target images in the image pair are respectively labeled with target categories and position information of the corresponding existing targets, and the two target images are registered to eliminate misalignment.
8. A target detection device for cross-domain information fusion, used in the target detection method for cross-domain information fusion according to any one of claims 1 to 7, characterized in that: include: A registration unit, used for registering the modality A image and the modality B image to obtain a registered modality A image and modality B image; A detection unit includes a dual-stream target detection model, which is used to input the registered modality A image and the registered modality B image into the dual-stream target detection model to obtain the detection result. Wherein, the dual-stream target detection model includes: A feature extraction module performs feature extraction on the registered modality A image to obtain a high-level semantic feature A, and performs feature extraction on the registered modality B image to obtain a high-level semantic feature B; A multi-scale enhancement module performs multi-scale enhancement on the feature pair consisting of the high-level semantic feature A and the high-level semantic feature B to obtain an enhanced input feature pair; A fusion module performs spatial-level multimodal interaction and channel-level multimodal interaction on the enhanced input feature pair to obtain multimodal features; The target detection module includes a detection head, which generates a detection frame and a confidence level of a corresponding category as the detection result according to the multimodal features.
Citation Information
Patent Citations
Multi-modal target detection method and device based on feature enhancement and collaborative interaction
CN117423007A