Multi-modal target detection method and device, terminal and storage medium
The introduction of a UNet network as the visual backbone in a multi-modal detection model addresses the limitations of supervised and unsupervised methods by reducing labeled data reliance and enhancing feature representation, improving target detection performance in complex environments.
Patent Information
- Application Number
- CN202510807532.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing object detection methods rely on labeled data, lack of diversity in visual backbones and high training costs, which affects the applicability and performance of the model in data scarce scenarios.
The UNet network is introduced as a visual backbone network, combined with CLIP encoder and VAE encoder, self-supervised learning is performed through multimodal object detection model, and semantic information in image tasks is used to generate text, reduce dependence on labeled data, and improve the diversity and generalization ability of feature representations.
It reduces training costs, improves the detection performance of the model in complex scenarios and the applicability of data in scenarios of scarcity, and provides a more diverse and generalized feature representation method.
Smart Images

Figure CN120318503A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-modal object detection method, device, terminal and storage medium, belonging to the technical field of object detection. Background Art
[0002] Existing disaster rescue equipment tends to be intelligent and unmanned, achieving "replacing humans with equipment". Unmanned equipment operation needs to replace "human eyes" for target recognition and positioning. One problem that current intelligent rescue equipment must solve is: in the complex post-disaster environment (irregular target shapes, occlusions, etc.), how to quickly and accurately locate the target and how to identify unknown targets are all goals that the object detection model must achieve. The task of object detection is to predict a set of bounding boxes and classification labels for a given image. Modern detectors select a supervised or unsupervised visual backbone network trained to extract image features, and then design a one-stage anchor-free, two-stage anchor-based or end-to-end manner based on Transformer (a deep learning model architecture for natural language processing and other sequence-to-sequence tasks) to train the detection head to achieve the recognition of the target of interest.
[0003] The existing technology mainly relies on supervised and unsupervised visual backbone networks in object detection tasks, such as ResNet101 (a version of a deep convolutional neural network with 101 layers) and MAE (a self-supervised pre-training framework for visual representation learning). Although these methods have achieved remarkable results in object detection tasks, there are still some technical limitations.
[0004] Among them, supervised learning methods (such as ResNet101) require a large amount of accurately labeled training data to optimize the model. However, obtaining high-quality labeled data is costly and time-consuming. Especially in object detection tasks, labeling bounding boxes and category information requires a large amount of manual intervention. This dependence on labeled data limits the application of the model in data-scarce scenarios and also increases the training cost.
[0005] Unsupervised learning methods (such as MAE) although reduce the dependence on labeled data, their representation ability is usually limited by the design of the pre-training task. For example, MAE learns image features through a masked reconstruction task, but this task may not be able to fully capture the fine-grained semantic information related to the object detection task. Therefore, the performance of unsupervised learning methods in object detection tasks is often not as good as that of supervised learning methods.
[0006] Existing methods mainly focus on discriminative models and reconstruction models in object detection tasks, while there is less exploration of generative models. Generative models (such as diffusion models) have powerful data generation capabilities and implicit knowledge learning capabilities, and can learn rich visual features from large-scale unlabeled data. However, the existing technology fails to fully utilize the potential of generative models, resulting in limited diversity of visual backbones and unable to provide more generalizable feature representations for object detection tasks.
[0007] In summary, the existing object detection methods usually have problems such as relying on labeled data, insufficient diversity of visual backbones, and high training costs, which affect the performance and practicality of object detection. Summary of the Invention
[0008] The purpose of the present invention is to overcome the deficiencies in the existing technology and provide a multi-modal object detection method, device, terminal, and storage medium. For the first time, the UNet network is introduced as a visual backbone network in object detection tasks, breaking through the limitations of traditional supervised and unsupervised methods and providing a new feature representation method for object detection tasks. Through the implicit knowledge learning mechanism of the UNet network, the present invention combines the semantic information in the text-to-image generation task with the object detection task, significantly improving the performance of the model in complex scenarios. By utilizing the self-supervised learning ability of the multi-modal object detection model, the present invention reduces the dependence on a large amount of labeled data, lowers the training cost, and improves the applicability of the model in data-scarce scenarios.
[0009] To solve the above technical problems, the present invention is implemented by the following technical solutions:
[0010] In the first aspect, the present invention provides a multi-modal object detection method, including the following steps:
[0011] Obtain a target image;
[0012] Input the target image into a pre-trained multi-modal object detection model, and output the localization and classification prediction results to complete the detection; the multi-modal object detection model includes a CLIP encoder, a VAE encoder, a UNet network, and a detection head;
[0013] Among them, the detection process of the multi-modal object detection model for the target image specifically includes:
[0014] Input the target image into the CLIP encoder to generate a text prompt embedding;
[0015] Input the target image into the VAE encoder to generate a latent feature map;
[0016] Input the latent feature map and the text prompt embedding into the UNet network to obtain a multi-scale diffusion feature map;
[0017] Calculate the cross-attention calculation result graph of the text prompt embedding and the latent feature map in the UNet network;
[0018] Perform channel dimension splicing on the cross-attention calculation result graph and the multi-scale diffusion feature map to form an enhanced feature map;
[0019] Input the enhanced feature map into the detection head to output the localization and classification prediction results.
[0020] Further, the training of the multi-modal object detection model specifically includes:
[0021] Obtain an image set and a multi-modal object detection model;
[0022] Randomly sample a group of random images from the image set;
[0023] Input the random images into the CLIP encoder to generate text prompt embeddings;
[0024] Input the random images into the VAE encoder to generate latent feature maps;
[0025] Input the latent feature maps and text prompt embeddings into the UNet network to obtain multi-scale diffusion feature maps;
[0026] Calculate the cross-attention calculation result graph of the text prompt embedding and the latent feature map in the UNet network;
[0027] Perform channel dimension splicing on the cross-attention calculation result graph and the multi-scale diffusion feature map to form an enhanced feature map;
[0028] Input the enhanced feature map into the detection head to output the localization and classification prediction results;
[0029] According to the output localization and classification prediction results, calculate the total loss function, and update the parameters of the multi-modal object detection model according to gradient backpropagation to complete the training of the multi-modal object detection model.
[0030] Further, the step of inputting the target image into the CLIP encoder to generate text prompt embeddings specifically includes:
[0031] Input the target image into the CLIP encoder to obtain the corresponding image features;
[0032] Pass the corresponding image features through two linear layers to obtain the mapped text prompt embeddings.
[0033] Further, the step of performing channel dimension splicing on the cross-attention calculation result graph and the multi-scale diffusion feature map to form an enhanced feature map specifically includes:
[0034] The second and third layers of the cross-attention calculation result map and the third and fourth layers of the multi-scale diffusion feature map are concatenated along the channel dimension to form an enhanced feature map.
[0035] Further, the enhanced feature map is input into the detection head to output the localization and classification prediction results, specifically including:
[0036] The enhanced feature map is input into the FPN module of the detection head for multi-scale feature fusion to obtain a fused feature map;
[0037] The fused feature map is input into the classification module and the localization module of the detection head to output the localization and classification prediction results.
[0038] Further, the total loss function is calculated by the following formula:
[0039] ;
[0040] In the formula: is the total loss function; is the region proposal network sub-module of the detection head; is the region of interest sub-module of the detection head; is the classification loss of rpn; is the localization loss of rpn; is the classification loss of roi; is the localization loss of roi;
[0041] Further, the classification loss of the roi is calculated by the following formula:
[0042] ;
[0043] In the formula: N is the total number of images in a group; n is the image serial number in a group of images; C is the total number of categories in the object detection dataset; c is the category serial number; is the classification prediction result of the image with image serial number n and category serial number c in a group of images; is the annotation information of;
[0044] The classification loss of the rpn is calculated by the following formula:
[0045] ;
[0046] In the formula: is the number of foreground and background categories in the object detection dataset;
[0047] The localization loss of the rpn is calculated by the following formula:
[0048] ;
[0049] ;
[0050] where: M is the total number of regions to be processed, and m is the serial number of the total number of regions among the total number of regions to be processed; is the target location prediction result of the image with the serial number m of the total number of regions among the total number of regions to be processed; is the annotation information of the smooth error loss;
[0051] The location loss of the said roi is obtained by the following formula:
[0052] ;
[0053] ;
[0054] where: is the total number of regions processed by the roi, is the serial number of the total number of regions among the total number of regions processed by the roi; is for the image with the serial number of the total number of regions among the total number of regions processed by the roi being the target location prediction result; is the annotation information of
[0055] In a second aspect, the present invention provides a multi-modal target detection device, including:
[0056] Input module: used to obtain the target image;
[0057] Detection module: used to input the target image into a pre-trained multi-modal target detection model, output the location and classification prediction results, and complete the detection;
[0058] wherein, the multi-modal target detection model includes a CLIP encoder, a VAE encoder, a UNet network, and a detection head; the detection process of the detection module for the target image specifically includes:
[0059] Scene text prompt generation module: used to input the target image into the CLIP encoder to generate text prompt embeddings;
[0060] Image encoding module: used to input the target image into the VAE encoder to generate a latent feature map;
[0061] Diffusion Feature Extraction Module: It is used to input the latent feature map and text prompt into the UNet network to obtain a multi-scale diffusion feature map;
[0062] Cross-Attention Map Fusion Module: It is used to calculate the cross-attention calculation result map of the text prompt embedding and the latent feature map in the UNet network; and form an enhanced feature map by performing channel dimension splicing according to the cross-attention calculation result map and the multi-scale diffusion feature map;
[0063] Object Detection Module: It is used to input the enhanced feature map into the detection head and output the localization and classification prediction results.
[0064] Thirdly, the present invention provides a terminal, including a processor and a storage medium;
[0065] The storage medium is used to store instructions;
[0066] The processor is used to operate according to the instructions to execute the steps of the method according to any one of the first aspect.
[0067] Fourthly, a computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the steps of the method according to any one of the first aspect.
[0068] Compared with the prior art, the beneficial effects achieved by the present invention:
[0069] In this multi-modal object detection method, the UNet network is introduced as a visual backbone network in the object detection task for the first time, breaking through the limitations of traditional supervised and unsupervised methods; traditional methods are limited by the dependence on labeled data or the design of pre-training tasks, while the present invention uses the rich implicit knowledge learned by the UNet network to provide a new, more diverse and generalization-capable feature representation method for the object detection task, opening up a new technical route for the object detection field; the present invention utilizes the self-supervised learning ability of the multi-modal object detection model, reduces the dependence on a large amount of labeled data, reduces the training cost, and at the same time improves the applicability of the model in data-scarce scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is a flowchart of a multi-modal object detection method provided by an embodiment of the present invention;
[0071] Figure 2 is a specific process schematic diagram of a multi-modal object detection method provided by an embodiment of the present invention;
[0072] Figure 3 is a structural schematic diagram of a multi-modal object detection device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.
[0074] The term "and / or" is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the preceding and following associated objects.
[0075] Embodiment 1:
[0076] As Figures 1 to 2 shown, the present invention provides a multi-modal object detection method, including the following steps:
[0077] Obtain a target image;
[0078] Input the target image into a pre-trained multi-modal object detection model, and output the localization and classification prediction results to complete the detection; the multi-modal object detection model includes a CLIP encoder, a VAE encoder, a UNet network, and a detection head;
[0079] Among them, the detection process of the multi-modal object detection model for the target image specifically includes:
[0080] Input the target image into the CLIP encoder (contrastive language-image pre-training encoder) to generate a text prompt embedding, specifically including:
[0081] Input the target image into the CLIP encoder to obtain the corresponding image features;
[0082] Pass the corresponding image features through two linear layers to obtain the mapped text prompt embedding;
[0083] Input the target image into the VAE encoder (variational autoencoder) to generate a latent feature map;
[0084] Input the latent feature map and the text prompt embedding into the UNet (a convolutional neural network architecture designed for image segmentation) network to obtain a multi-scale diffusion feature map; among them, the UNet network is a UNet based on a text-to-image diffusion model;
[0085] Calculate the cross-attention calculation result map of the text prompt embedding and the latent feature map in the UNet network;
[0086] The enhanced feature map is formed by concatenating the cross-attention calculation result map and the multi-scale diffusion feature map along the channel dimension, specifically including:
[0087] Concatenate the second and third layers of the cross-attention calculation result map and the third and fourth layers of the multi-scale diffusion feature map along the channel dimension to form the enhanced feature map;
[0088] Input the enhanced feature map into the detection head to output the localization and classification prediction results, specifically including:
[0089] Input the enhanced feature map into the FPN (Feature Pyramid Network) module of the detection head for multi-scale feature fusion to obtain the fused feature map;
[0090] Input the fused feature map into the classification module and the localization module of the detection head to output the localization and classification prediction results.
[0091] The present invention solves the problems of existing object detection methods relying on labeled data and insufficient diversity of visual backbones by introducing a generative paradigm visual backbone network and combining semantic text prompts with a location attention enhancement mechanism, thereby improving the generalization ability and performance of object detection; specifically, Figure 2 The VQGAN encoder (discrete encoder) in [[ ]] is a manifestation form of the VAE encoder.
[0092] Specifically, such as Figure 2As shown, the region-aware feature is that SAM (a pre-trained model for image segmentation) extracts the segmentation mask features of an image through a vision Transformer architecture, which can accurately capture the contours and spatial positions of arbitrary objects and is suitable for zero-shot segmentation tasks; the fully connected layer is a basic component of a neural network. The multi-layer perceptron performs high-dimensional mapping on the input features by stacking linear layers and non-linear activation functions and is often used in classification heads or feature fusion. It has a large number of parameters but high flexibility and can learn complex non-linear relationships; the latent feature is that in a latent diffusion model, an image is compressed into a low-dimensional latent variable by a VAE or VQGAN encoder, and the diffusion process (adding noise and denoising) is carried out in the latent space, significantly reducing the computational cost while retaining the generation quality. After decoding the latent feature, a high-fidelity image is output; the cross-attention layer is that in the text-to-image generation task, the UNet network introduces a cross-attention mechanism, uses the text embedding as the key-value and the image features as the query, dynamically generates spatial attention weights, and realizes fine-grained control of the text semantics over the local regions of the image; the region proposal network is the core component of Faster R-CNN. It generates anchor points on the feature map through a sliding window, predicts the object probability and bounding box offset of each anchor point, filters out high-quality candidate regions, and provides preliminary object localization for two-stage detectors; the region of interest pooling is for the candidate boxes of different sizes output by the region proposal network of Faster R-CNN. It normalizes them into feature blocks of a fixed size (such as 7×7) through bilinear interpolation or max pooling, retains the spatial information while adapting to the input requirements of the subsequent fully connected layer, and is a key operation for feature alignment in object detection.
[0093] Based on the text-to-image diffusion model framework, the present invention innovatively proposes a method of using a generative paradigm vision backbone network for object detection tasks; this method makes full use of the rich implicit knowledge learned by the diffusion model during the text-to-image training process, has excellent visual perception ability, and can provide a brand-new generative paradigm-based vision backbone network for object detection tasks.
[0094] Specifically, the present invention combines the semantic information in the text-to-image task with the object detection task through the implicit knowledge learning mechanism of the diffusion model, significantly improving the performance of the model in complex scenarios; the present invention utilizes the self-supervised learning ability of a multi-modal object detection model (generative model), reduces the dependence on a large amount of labeled data, reduces the training cost, and at the same time improves the applicability of the model in data-scarce scenarios; through the above innovations, the present invention provides a more diverse, generalization-capable and cost-effective solution for object detection tasks.
[0095] The present invention first introduces the UNet network (a diffusion model UNet based on text-to-image) as a visual backbone network in the object detection task, breaking through the limitations of traditional supervised (such as ResNet) and unsupervised (such as MAE) methods; traditional methods are limited by the dependence on labeled data or the design of pre-training tasks, while the present invention utilizes the rich implicit knowledge learned by the UNet network (deep understanding of object structure, texture, and scene relationships) to provide a new, more diverse, and generalization-capable feature representation method for the object detection task. This opens up a new technical route in the field of object detection.
[0096] An embodiment, the training of the multi-modal object detection model specifically includes:
[0097] Obtain an image set and a multi-modal object detection model;
[0098] Randomly sample a set of random images from the image set;
[0099] Input the random images into the CLIP encoder to generate text prompt embeddings;
[0100] Input the random images into the VAE encoder to generate latent feature maps;
[0101] Input the latent feature maps and text prompt embeddings into the UNet network to obtain multi-scale diffusion feature maps;
[0102] Calculate the cross-attention calculation result map of the text prompt embeddings and the latent feature maps in the UNet network;
[0103] Perform channel dimension splicing on the cross-attention calculation result map and the multi-scale diffusion feature maps to form enhanced feature maps;
[0104] Input the enhanced feature maps into the detection head to output localization and classification prediction results;
[0105] According to the output localization and classification prediction results, calculate the total loss function, and update the parameters of the multi-modal object detection model according to gradient backpropagation to complete the training of the multi-modal object detection model.
[0106] Specifically, the training of the multi-modal object detection model specifically includes:
[0107] Randomly sample a set of random images from the image set ;
[0108] Input the random images into the CLIP image encoder to obtain corresponding image features with a size of ; then, Through two linear layers, the mapped scene text prompt is obtained. ;
[0109] The specific expression is as follows:
[0110] ; ;
[0111] Among them, represents the total number of pictures in the current group of images; ch represents the number of channels; H represents the height; W represents the width; represents the multi-layer perceptron;
[0112] Input the image into the VAE encoder to obtain the latent feature header ;
[0113] Subsequently, and are simultaneously input into the denoising UNet network of the text-to-image diffusion model to obtain five-layer multi-scale diffusion feature maps ; The size of the feature map of each layer is respectively , , , ;
[0114] The specific expression is as follows:
[0115] ; ;
[0116] Among them, is the number of layers of the multi-scale diffusion feature map in the UNet network; is the multi-scale diffusion feature map;
[0117] Calculate the cross-attention calculation result map of the text prompt embedding and the latent feature map in the UNet network ; The size of the feature map of each layer is respectively , , , ; Select i = 2, 3 layers and = 3, 4 layers, and perform a channel-level merging operation on and , and replace the = 3, 4 features in the original multi-scale diffusion feature map with the result to obtain the enhanced feature map , and its size is respectively , , , , ;
[0118] The specific expressions are as follows:
[0119] ;
[0120] , ;
[0121] Among them, is the number of layers of the cross-attention calculation result map; k is the number of layers of the enhanced feature map.
[0122] Input the diffusion feature into the FPN module for multi-scale feature fusion to obtain the fused feature map ; Specifically, use 1×1 convolution to adjust the number of channels of each fused feature map so that their number of channels is the same as 256; starting from the feature map with the smallest size, each layer of the feature map will be fused with the feature map of the previous layer through upsampling (using bilinear interpolation), and the high-level features of the top layer will be retained to better utilize semantic information; where p is the number of layers of the fused feature map;
[0123] Input the fused feature map into the classification module and localization module of the detection head to output the localization and classification prediction results;
[0124] According to the output localization and classification prediction results, calculate the total loss function, and update the parameters of the multi-modal object detection model according to the gradient backpropagation. After one round of training, re-conduct the training work or end the training of the multi-modal object detection model.
[0125] In one embodiment, the total loss function is calculated by the following formula:
[0126] ;
[0127] In the formula: is the total loss function; is the region proposal network sub-module of the detection head; is the region of interest sub-module of the detection head; is the classification loss of rpn; is the localization loss of rpn; is the classification loss of roi; is the localization loss of roi;
[0128] The classification loss of the roi is calculated by the following formula:
[0129] ;
[0130] Where: N is the total number of images in a group; n is the image serial number within a group of images; C is the total number of categories in the target detection dataset; c is the category serial number; is the classification prediction result of the image with image serial number n and category serial number c in a group of images; is the annotation information of
[0131] The classification loss of the rpn is obtained by the following formula:
[0132] ;
[0133] Where: is the number of foreground and background categories in the target detection dataset;
[0134] The localization loss of the rpn is obtained by the following formula:
[0135] ;
[0136] ;
[0137] Where: M is the total number of regions processed, and m is the serial number of the region among the total number of regions processed; is the target localization prediction result of the image with serial number m among the total number of regions processed; is the annotation information of is the smooth error loss;
[0138] The localization loss of the roi is obtained by the following formula:
[0139] ;
[0140] ;
[0141] Where: is the total number of regions processed by the roi, is the serial number of the region among the total number of regions processed by the roi; is the target localization prediction result of the image with serial number is the annotation information of
[0142] An embodiment. The network proposed in this application is trained and tested in the sub-dataset Day sunny of the autonomous driving dataset Diverse-dataset. The AP metric is used to evaluate each category, and the mAP metric is used for overall evaluation. The comparison results with mainstream visual backbone network methods are shown in the following table:
[0143] Table 1. Evaluation results of the Diverse-dataset dataset;
[0144]
[0145] In the table, Ours refers to the present invention;
[0146] The method of the text-to-image diffusion model proposed in the present invention as a new visual backbone network significantly outperforms classic visual backbone networks such as Resnet50 (a version of a deep convolutional neural network with 50 layers), Resnet101, and CLIP. It also achieves comparable performance to the current state-of-the-art method Convnext (a modern convolutional neural network architecture), demonstrating that the visual backbone network based on the generative paradigm (i.e., the text-to-image diffusion model UNet) can be used for object detection tasks.
[0147] Through the implicit knowledge learning mechanism of the text-to-image diffusion model UNet, the present invention creatively combines the semantic information in the text generation image task (embedded by text prompts generated by CLIP) with the object detection task (localization and classification). The core lies in using the cross-attention mechanism inside the text-to-image diffusion model UNet to calculate the spatial association map between the text semantics and the image latent feature map. This attention map accurately indicates the spatial position of the concepts in the text description in the image. By splicing and fusing this attention map with the multi-scale visual features extracted by the diffusion model itself in the channel dimension, the model's ability to understand complex scenes (such as irregular object shapes, occlusions, background interferences, etc.) is significantly enhanced, and the model's performance in complex scenes is improved.
[0148] Embodiment Two:
[0149] As Figure 3 shown, the present invention provides a multi-modal object detection device, including:
[0150] Input module: used to obtain the target image;
[0151] Detection module: used to input the target image into the pre-trained multi-modal object detection model, output the localization and classification prediction results, and complete the detection;
[0152] Among them, the multi-modal object detection model includes a CLIP encoder, a VAE encoder, a UNet network, and a detection head; the detection process of the detection module for the target image specifically includes:
[0153] Scene text prompt generation module: used to input the target image into the CLIP encoder to generate text prompt embeddings;
[0154] Image encoding module: used to input the target image into the VAE encoder to generate a latent feature map;
[0155] Diffusion feature extraction module: used to input the latent feature map and text prompt embeddings into the UNet network to obtain a multi-scale diffusion feature map;
[0156] Cross-attention map fusion module: used to calculate the cross-attention calculation result map of the text prompt embeddings and the latent feature map in the UNet network; perform channel dimension splicing on the cross-attention calculation result map and the multi-scale diffusion feature map to form an enhanced feature map;
[0157] Object detection module: used to input the enhanced feature map into the detection head to output the localization and classification prediction results.
[0158] Specifically, the image encoding module: compresses the original image into low-dimensional latent features, which are the input of the diffusion feature extraction module. After inputting the image (size 448×800 pixels), it performs spatial compression through the VAE encoder to generate a latent feature map (size 64×112×4); the VAE encoder is based on the pre-trained weights of Stable Diffusion (text-to-image diffusion model), and the parameters are frozen to avoid updating during training.
[0159] Scene text prompt generation module: generates text embeddings that are semantically aligned with the input image to enhance the feature expression ability of the diffusion model; after inputting the image (size 448×800 pixels), it extracts the global feature vector (dimension 768) through the CLIP image encoder, and the feature vector is reduced in dimension to the same dimension as the text embedding of the diffusion model (dimension 768) through two layers of MLP (multi-layer perceptron) to obtain text prompt embeddings; using the image-text alignment feature of CLIP, the visual features are transformed into semantic text conditions to enhance the high-level semantic perception ability of the diffusion model.
[0160] Diffusion feature extraction module: extracts multi-scale visual features from the UNet network of the diffusion model; used to input the latent feature map and text prompt embeddings into the UNet network to obtain a multi-scale diffusion feature map; the UNet network is pre-trained based on the LAION-4B dataset (a large-scale publicly available dataset containing approximately 4 billion image-text pairs), and the parameters are frozen to retain the generation ability.
[0161] Cross - attention map fusion module: Captures the positional association between text and image through a cross - modal attention mechanism, enhancing the positional sensitivity of features; used to calculate the cross - attention calculation result map (size 32×56×768) of the text prompt embedding and the latent feature map in the UNet network; forms an enhanced feature map (size 32×56×1024) by performing channel - dimension splicing on the cross - attention calculation result map and the multi - scale diffusion feature map; introduces low - level position information to improve the localization accuracy of object detection.
[0162] Object detection module: Completes object classification and bounding box regression based on the enhanced feature map; the enhanced feature map is input into the Faster R - CNN (a deep - learning model for object detection) detection head to generate the final detection result; during training, only the parameters of the detection head are optimized, and other modules are frozen to reduce the computational cost.
[0163] Example Three:
[0164] An embodiment of the present invention also provides a terminal, including a processor and a storage medium;
[0165] The storage medium is used to store instructions;
[0166] The processor is used to operate according to the instructions to execute the steps of the method described in Example One.
[0167] Example Four:
[0168] An embodiment of the present invention also provides a computer - readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method described in Example One.
[0169] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk storage, CD - ROM, optical storage, etc.) containing computer - usable program code.
[0170] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0171] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0173] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A multi-modal object detection method, characterized in that, It includes the following steps: Obtain the target image; Input the target image into a pre-trained multi-modal object detection model to output the localization and classification prediction results, completing the detection; the multi-modal object detection model includes a CLIP encoder, a VAE encoder, a UNet network, and a detection head; Among them, the detection process of the multi-modal object detection model for the target image specifically includes: Input the target image into the CLIP encoder to generate a text prompt embedding; Input the target image into the VAE encoder to generate a latent feature map; Input the latent feature map and the text prompt embedding into the UNet network to obtain a multi-scale diffusion feature map; Calculate the cross-attention calculation result map of the text prompt embedding and the latent feature map in the UNet network; Perform channel dimension splicing according to the cross-attention calculation result map and the multi-scale diffusion feature map to form an enhanced feature map; Input the enhanced feature map into the detection head to output the localization and classification prediction results.
2. The multimodal object detection method according to claim 1, wherein The training of the multi-modal object detection model specifically includes: Obtain an image set and a multi-modal object detection model; Randomly sample a group of random images from the image set; Input the random images into the CLIP encoder to generate a text prompt embedding; Input the random images into the VAE encoder to generate a latent feature map; Input the latent feature map and the text prompt embedding into the UNet network to obtain a multi-scale diffusion feature map; Calculate the cross-attention calculation result map of the text prompt embedding and the latent feature map in the UNet network; Perform channel dimension splicing according to the cross-attention calculation result map and the multi-scale diffusion feature map to form an enhanced feature map; Input the enhanced feature map into the detection head to output the localization and classification prediction results; According to the output localization and classification prediction results, calculate the total loss function, and update the parameters of the multi-modal object detection model according to the gradient backpropagation to complete the training of the multi-modal object detection model.
3. The multimodal object detection method according to claim 1, wherein The step of inputting the target image into the CLIP encoder to generate a text prompt embedding specifically includes: Input the target image into the CLIP encoder to obtain the corresponding image features; Pass the corresponding image features through two linear layers to obtain the mapped text prompt embedding.
4. The multimodal object detection method according to claim 1, wherein The step of performing channel dimension splicing according to the cross-attention calculation result map and the multi-scale diffusion feature map to form an enhanced feature map specifically includes: Splice the second and third layers of the cross-attention calculation result map and the third and fourth layers of the multi-scale diffusion feature map along the channel dimension to form an enhanced feature map.
5. The multimodal object detection method according to claim 1, wherein The step of inputting the enhanced feature map into the detection head to output the localization and classification prediction results specifically includes: Input the enhanced feature map into the FPN module of the detection head for multi-scale feature fusion to obtain the fused feature map; Input the fused feature map into the classification module and the localization module of the detection head to output the localization and classification prediction results.
6. The multimodal object detection method according to claim 2, wherein The total loss function is calculated by the following formula: ; In the formula: is the total loss function; is the region generation network sub-module of the detection head; is the region of interest sub-module of the detection head; is the classification loss of the rpn; is the localization loss of the rpn; is the classification loss of the roi; is the localization loss of the roi.
7. The multimodal object detection method according to claim 6, wherein The classification loss of the roi is calculated by the following formula: ; Where: N is the total number of images in a group; n is the image serial number within a group of images; C is the total number of categories in the target detection dataset; c is the category serial number; is the classification prediction result of the image with image serial number n and category serial number c in a group of images; is the annotation information of; The classification loss of the rpn is calculated by the following formula: ; Wherein: is the number of foreground and background categories in the target detection dataset; The localization loss of the rpn is calculated by the following formula: ; ; Where: M is the total number of regions to be processed, and m is the serial number of the total number of regions in the total number of regions to be processed; is the target location prediction result of the image with the serial number m of the total number of regions in the total number of regions to be processed; is the annotation information of; is the smoothing error loss; The localization loss of the roi is calculated by the following formula: ; ; Wherein: is the total number of regions processed by the ROI, is the serial number of the total number of regions among the total number of regions processed by the ROI; is the target location prediction result of the image with the serial number among the total number of regions processed by the ROI; is the annotation information of.
8. A multimodal object detection device, characterized in that, Including: Input module: used to obtain the target image; Detection module: used to input the target image into a pre-trained multi-modal object detection model, output the localization and classification prediction results, and complete the detection; Among them, the multi-modal object detection model includes a CLIP encoder, a VAE encoder, a UNet network, and a detection head; the specific detection process of the detection module for the target image includes: Scene text prompt generation module: used to input the target image into the CLIP encoder to generate text prompt embeddings; Image encoding module: used to input the target image into the VAE encoder to generate a latent feature map; Diffusion feature extraction module: used to input the latent feature map and text prompt embeddings into the UNet network to obtain a multi-scale diffusion feature map; Cross-attention map fusion module: used to calculate the cross-attention calculation result map of the text prompt embeddings and the latent feature map in the UNet network; perform channel dimension splicing on the cross-attention calculation result map and the multi-scale diffusion feature map to form an enhanced feature map; Object detection module: used to input the enhanced feature map into the detection head to output the localization and classification prediction results.
9. A terminal, characterized in that, Including a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Remote sensing target detection method based on diffusion model
CN119152285A
Monocular 3D target detection method and device
CN119206196A
Contextual visual-based SAR target detection method and apparatus, and storage medium
US20230184927A1
Cited By
Automatic detection method and equipment for blood cells in blood smear image
CN122089709A
Method and apparatus for automatic detection of blood cells in blood smear images
CN122089709B