Open target detection method and device, equipment and storage medium
By constructing a target network and combining it with a feature deformation adapter and a cross-modal loss function, the problem of semantic differences between local ROI features and global image features is solved, and the detection accuracy and generalization ability of the open object detection model are improved.
Patent Information
- Application Number
- CN202510878468.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-26
AI Technical Summary
In existing open vocabulary object detection methods, the spatial deformation problem during local ROI feature extraction leads to significant semantic differences between local region features and global image features extracted by the CLIP model, which in turn causes a decrease in the accuracy of local open object detection.
Construct a target network, including a region proposal network, a feature deformation adapter, an image encoder, and a text encoder. Use the feature deformation adapter to align features of local regions of interest. Combine the local instance-level distillation loss and the cross-modal feature alignment loss to update the network parameters and improve the feature alignment capability.
It effectively reduces the feature semantic differences caused by local ROI feature deformation, enhances the deformation feature alignment capability, and improves the detection generalization capability of the open object detection model for unknown categories.
Smart Images

Figure CN120707834A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of target detection technology, and in particular to an open target detection method, apparatus, device and storage medium. Background Art
[0002] The core challenge of open vocabulary object detection lies in identifying new objects outside the training set. Existing methods usually rely on the cross-modal image-text alignment capabilities of visual language pre-training models such as the CLIP model. However, CLIP's image-text pair training mechanism has inherent limitations: it extracts features from the entire image through global semantic encoding and lacks the ability to perceive refined features of local fine-grained regions of the image. In the open vocabulary object detection task, the detection model needs to focus on the fine-grained features of the local ROI (region of interest) of the image. However, the spatial deformation problem during the extraction of local ROI features will lead to significant semantic differences between the local region features and the global image features pre-trained by the CLIP model, which in turn leads to the problem of reduced accuracy in local open object detection. Summary of the Invention
[0003] The present application provides an open target detection method, apparatus, device and storage medium for improving the spatial deformation problem during local ROI feature extraction, which results in significant semantic differences between local region features and global image features extracted by the CLIP model, thereby causing a technical problem of decreased accuracy in local open target detection.
[0004] In view of this, the first aspect of the present application provides an open target detection method, comprising:
[0005] Constructing a target network, wherein the target network includes a region proposal network, a feature deformation adapter, an image encoder, and a text encoder;
[0006] Extracting local regions of interest from training images through a region proposal network, and aligning the features of the local regions of interest through a feature deformation adapter to obtain adapted local region features;
[0007] Inputting the adapted local region features into an image encoder for processing to generate image embedding features, and calculating a local instance-level distillation loss using the image embedding features and the adapted local region features;
[0008] Inputting the text description corresponding to the training image into a text encoder for processing to generate text embedding features, and calculating a cross-modal feature alignment loss using the image embedding features and the text embedding features;
[0009] Combining the local instance-level distillation loss and the cross-modal feature alignment loss to update the network parameters of the target network to obtain a trained open object detection model;
[0010] Object detection is performed on the image to be detected using the open object detection model.
[0011] Optionally, there are multiple feature deformation adapters, each of which includes a feature adapter and a deformation adapter. The feature adapter consists of a pooling layer, an activation layer, and a fully connected layer, and all feature adapters share the pooling layer and the activation layer.
[0012] The feature adapter is used to perform feature mapping on the local region of interest and perform residual connection with the local region of interest to obtain local region features;
[0013] The deformation adapter is used to select a corresponding deformation adaptation weight according to a discrete interval where the aspect ratio of the local area feature is located to weight the local area feature to obtain an adapted local area feature.
[0014] Optionally, the method further includes:
[0015] Performing global feature extraction on the training image by the image encoder to obtain a first global feature;
[0016] Performing average pooling and feature aggregation on the adapted local area features corresponding to the training image to obtain a second global feature;
[0017] Calculating the cosine similarity between the first global feature and the second global feature to obtain a global feature training loss;
[0018] Accordingly, the network parameters of the target network are updated by combining the local instance-level distillation loss and the cross-modal feature alignment loss to obtain a trained open object detection model, including:
[0019] The local instance-level distillation loss, the cross-modal feature alignment loss, and the global feature training loss are combined to update the network parameters of the target network to obtain a trained open object detection model.
[0020] Optionally, extracting a local region of interest from a training image using a region proposal network includes:
[0021] Processing the training image through a region proposal network to generate candidate boxes with confidence scores;
[0022] From the candidate frames, candidate frames whose intersection-over-union ratio with the real frames of the base category is lower than the intersection-over-union ratio threshold and whose confidence scores are higher than the confidence threshold are screened to obtain new category candidate frames; the area corresponding to the new category candidate frames is the local region of interest.
[0023] Optionally, the calculation formula for the local instance-level distillation loss is:
[0024]
[0025] Where N is the number of candidate frames of the new class target, M is the number of real frames of the basic class, is the local region feature of the i-th local region of interest after deformation adaptation; The text embedding feature for the j-th text description; The text embedding features generated for the region of interest of the base category, Embed features for text and the adapted local area features The cosine similarity between .
[0026] Optionally, calculating the cross-modal feature alignment loss using the image embedding feature and the text embedding feature includes:
[0027] The text embedding features of the text description and the corresponding image embedding features are used as positive sample pairs, and the text embedding features of the text description and the non-corresponding image embedding features are used as negative sample pairs. The cosine similarity between the text embedding features and the image embedding features is calculated to obtain the cross-modal feature alignment loss.
[0028] Optionally, the calculation formula for the cross-modal feature alignment loss is:
[0029]
[0030] Where, is the set of known basic categories, L is the number of known basic categories; A collection of text descriptions; The text embedding feature for the j-th text description; Basic category Corresponding text embedding features; Text embedding features generated for the region of interest of the base category; The text description generated for the local region of interest i of unknown category; is the weight parameter.
[0031] A second aspect of the present application provides an open object detection device, comprising:
[0032] A network construction unit, configured to construct a target network, wherein the target network includes a region proposal network, a feature deformation adapter, an image encoder, and a text encoder;
[0033] A feature adaptation unit is used to extract local regions of interest from training images using a region proposal network, and perform feature alignment on the local regions of interest using a feature deformation adapter to obtain adapted local region features;
[0034] a first loss calculation unit, configured to input the adapted local region features into an image encoder for processing to generate image embedding features, and calculate a local instance-level distillation loss using the image embedding features and the adapted local region features;
[0035] a second loss calculation unit, configured to input the text description corresponding to the training image into a text encoder for processing, generate text embedding features, and calculate a cross-modal feature alignment loss using the image embedding features and the text embedding features;
[0036] a parameter updating unit, configured to update the network parameters of the target network by combining the local instance-level distillation loss and the cross-modal feature alignment loss to obtain a trained open object detection model;
[0037] The target detection unit is used to perform target detection on the image to be detected using the open target detection model.
[0038] A third aspect of the present application provides an electronic device, the device comprising a processor and a memory;
[0039] The memory is used to store program code and transmit the program code to the processor;
[0040] The processor is configured to execute any one of the open object detection methods described in the first aspect according to instructions in the program code.
[0041] In a fourth aspect, the present application provides a computer-readable storage medium, which is used to store program code. When the program code is executed by a processor, it implements the open target detection method described in any one of the first aspects.
[0042] It can be seen from the above technical solutions that this application has the following advantages:
[0043] The open target detection method provided in this application reduces the feature semantic differences caused by local area feature deformation and enhances the deformation feature alignment capability by designing a feature deformation adapter; and proposes a granular hierarchical distillation framework to make full use of the multi-level semantic features of the image, build a cross-level knowledge transfer link, and effectively improve the open target detection model's detection generalization capability for unknown categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 A flowchart of an open target detection method provided in an embodiment of the present application;
[0046] Figure 2 A schematic structural diagram of an open target detection device provided in an embodiment of the present application;
[0047] Figure 3 A schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0049] For easier understanding, please refer to Figure 1 , an embodiment of the present application provides an open target detection method, comprising:
[0050] Step 110: Build a target network.
[0051] The target network in the embodiments of this application includes a region proposal network, a feature deformation adapter, and a CLIP model. The CLIP model includes an image encoder and a text encoder. The image encoder can be a classic convolutional network (such as ResNet-50) or a visual transformer (ViT), and the text encoder can use a Transformer network. The CLIP model in this application can use a pretrained CLIP model.
[0052] Among them, there are multiple feature deformation adapters, which include feature adapters and deformation adapters. The feature adapters are composed of pooling layers, activation layers, and fully connected layers. All feature adapters share the pooling layer and activation layer. The feature adapter is used to perform feature mapping on the local region of interest and perform residual connection with the local region of interest to obtain local region features. The deformation adapter is used to weight the local region features according to the discrete interval where the aspect ratio of the local region features is located to obtain the adapted local region features.
[0053] Step 120: extracting local regions of interest from the training image using a region proposal network, and aligning the local regions of interest using a feature deformation adapter to obtain adapted local region features.
[0054] Open vocabulary object detection requires only basic category training data and a pre-trained CLIP model to train an object detector capable of detecting new categories. Its core process consists of two parts: object localization and region classification. Given that basic training data covers only a limited number of categories and lacks prior information about new categories, simple distillation based solely on the pre-trained CLIP model is insufficient to effectively capture the rich semantic features of new categories.
[0055] To address the above issues, this application uses a region proposal network to process the original training image, generate candidate boxes with confidence scores, and filter out those whose IOU (intersection-over-union) with the ground-truth boxes of the basic categories is lower than the IOU threshold. (Indicates that it has low overlap with the basic category and may belong to a new category), and the confidence score is higher than the confidence threshold The candidate box (background area can be further filtered) is identified as a potential new category candidate box, and the area corresponding to the new category candidate box is the local area of interest. It can be 0.2, It can be 0.7.
[0056] It should be noted that the base category true box is the true bounding box of the target annotated in the training image for the known basic category, which contains the location coordinates of the target and the corresponding category label; the base category true box is essentially the manual annotation result of the basic category, and is the benchmark for target detection model training and new category screening. Its acquisition depends on high-quality dataset annotation.
[0057] Since the candidate boxes of the generated local regions of interest have different shapes, if the local region features are simply obtained through ROIAlign, due to the problem of feature deformation, it will not be able to be effectively aligned with the image embedding features extracted by the CLIP model. This embodiment of the application generates K feature adapters for each candidate box of the local region of interest. , each feature adapter is composed of a pooling layer, an activation layer R and a fully connected layer FC. For the i-th local region of interest r i , through residual connection and identity mapping weighted processing, the corresponding local area feature F can be obtained i , taking the kth local region feature of the i-th local region of interest as an example:
[0058]
[0059] Where, is the weight parameter, Represents the local region of interest r i The output features after pooling processing, Represents the activation layer pair The output features after processing, Represents a fully connected layer pair Output features after processing.
[0060] For the i-th local region of interest, K local region features can be output Among them, all local area features are calculated using shared pooling layers and activation layers, and differentiated local area features are generated through different fully connected layers.
[0061] Based on the K local area features output above, it is necessary to assign deformation adaptation weights to each local area feature , so that each deformation adapter only needs to process local area features within a specific range. Specifically, according to the aspect ratio of the local area features The corresponding K discrete intervals are divided. When K=10, such as the discrete interval ,according to D K Range is weighted, select a specific deformation adapter, such as the current In the discrete interval , then the deformation adaptation weight Then in The position is set to 1, and the other weight positions are set to 0, which is equivalent to selecting only the deformation adapter that processes the range to align the local area features. Finally, the deformation adapter recommended for the i-th local area of interest is × ,This scheme achieves targeted adaptation to local area features of different shapes through ,structured adapter design and aspect ratio interval division, ,effectively solving the cross-modal alignment failure problem caused by ,feature deformation.
[0062] Step 130: Input the adapted local region features into an image encoder for processing to generate image embedding features, and calculate the local instance-level distillation loss using the image embedding features and the adapted local region features.
[0063] The local region features weighted by the targeted instance deformation adapter are input into the image encoder for processing to generate the corresponding image embedding features. The local instance-level distillation loss is calculated using the image embedding features and the adapted local region features.
[0064] During training, the image embedding features of the local regions of interest (ROIs) of all training images read in the current batch are generated through the image encoder to obtain an image embedding feature queue. For any current ROI, the adapted local region features of the current ROI are trained on instance-level distillation through the image embedding feature queue, effectively classifying potential new class targets. Specifically, by leveraging the cross-modal image-text alignment capability of the pre-trained CLIP model, the cosine similarity between the image embedding feature queue extracted by the image encoder and each adapted local region feature is calculated. If the target in the current ROI is a base category, its similarity with the corresponding base category text is significantly higher. If the target in the current ROI is an unknown category, the CLIP model's powerful feature extraction capability can be used to extract features from the ROI, and the text category obtained by the CLIP model is used as the text description label of the potential new target. This method not only retains the classification accuracy of the base category, but also leverages the open vocabulary representation capability of the CLIP model, providing an interpretable label generation path for semantic modeling of unknown categories, effectively improving the CLIP model's generalization ability for new targets.
[0065] Specifically, for a local region of interest within a basic category, a predefined text description (e.g., category name) is predefined for the basic category (e.g., "car" or "airplane") based on predefined text labels. This predefined text description is processed through a text encoder to generate text embedding features. The cosine similarity between the image embedding features of the local region of interest and all predefined text embedding features is then calculated, and the text category corresponding to the maximum similarity is taken as the predicted result. If the similarity exceeds a similarity threshold, it is determined to be a basic category, and supervised training is performed using the real labels.
[0066] For local regions of interest of unknown categories: Based on open vocabulary semantic matching, assuming that the local region of interest is a "drone" (unknown category), the image embedding features of the local region of interest are first extracted through the image encoder; secondly, the similarity between the image embedding features and the text embedding features of the basic categories (such as "car" and "airplane") is calculated, and the similarity is lower than the similarity threshold; then, the cosine similarity is calculated in the text library of the CLIP model to match the closest text embedding feature, such as "small aircraft", which is used as the pseudo-label of the target in the local region of interest; the CLIP model learns the characteristics of "drone" through this pseudo-label, and the accuracy of the text label can be further fine-tuned based on more local regions of interest.
[0067] Step 140: Input the text description corresponding to the training image into a text encoder for processing to generate text embedding features, and calculate the cross-modal feature alignment loss using the image embedding features and the text embedding features;
[0068] The text descriptions corresponding to the training images are input into a text encoder for processing. The text descriptions are projected into the same dimensional space as the image embedding features to generate text embedding features. The text embedding features of the training images are paired with the corresponding image embedding features as positive pairs, and the text embedding features of the training images are paired with the non-corresponding image embedding features as negative pairs. Cosine similarity is used to calculate the distance between the text embedding features and the image embedding features, resulting in a cross-modal feature alignment loss. The training goal is to maximize the cosine similarity between the image and text embeddings for each positive pair, and minimize the cosine similarity between the image and text embeddings for each negative pair.
[0069] Step 150: combining the local instance-level distillation loss and the cross-modal feature alignment loss to update the network parameters of the target network to obtain a trained open object detection model;
[0070] In one embodiment, the local instance-level distillation loss and the cross-modal feature alignment loss can be weightedly summed to obtain a total loss, and the network parameters of the target network can be updated using the total loss until the target network converges (such as reaching the maximum number of training times or the training error is lower than a preset error threshold, etc.), thereby obtaining a trained open object detection model.
[0071] In another embodiment, in order to obtain global image context features of multi-category targets and complex environments, in addition to feature distillation through local regions of interest, contrast distillation is also performed using global image features. An image encoder is used to extract global features from the entire training image to obtain a first global feature; a trainable detector is used to average pool several adapted local area features and aggregate the pooled features to obtain a second global feature; then, feature-level distillation is performed on the first global feature extracted by the image encoder and the second global feature extracted by the detector, and the global feature training loss is calculated by designing cosine similarity to achieve cross-model alignment of global context information.
[0072] Finally, the local instance-level distillation loss, cross-modal feature alignment loss and global feature training loss are weightedly summed to obtain the total loss. The network parameters of the target network are updated with this total loss to obtain the trained open object detection model.
[0073] The total loss Loss calculation formula in the embodiment of this application is:
[0074]
[0075] Where, , , To balance the hyperparameters of each loss term, they are usually tuned using a validation set. is the local instance-level distillation loss, is the global feature training loss, is the cross-modal feature alignment loss.
[0076] The calculation formula of local instance-level distillation loss is:
[0077]
[0078]
[0079] Where, is the number of potential candidate boxes for new target categories, The set of candidate boxes generated by the region proposal network, is the set of ground-truth bounding boxes for the base category, M is the number of ground-truth bounding boxes for the base category, Candidate box With bounding box The intersection-over-intersection ratio, Candidate box The confidence score of is the intersection-over-union threshold, used to filter candidate boxes with low overlap with the basic category; is the confidence threshold, used to filter the background area; is the local region feature of the i-th local region of interest after deformation adaptation; The deformation adaptation weight assigned to the kth deformation adapter for the i-th local region of interest; is the output feature of the kth feature adapter; is the i-th local region of interest; FC k is the fully connected layer of the k-th feature adapter; Text embedding features generated for the region of interest of the base category; The text embedding feature for the j-th text description; Embed features for text and the adapted local area features The cosine similarity between .
[0080] The calculation formula for global feature training loss is:
[0081]
[0082] Where, is the first global feature of the training image I; The features after average pooling of the adapted local region features of the i-th local region of interest in the training image I; is the feature after average pooling corresponding to the training image I The second global feature after aggregation.
[0083] The calculation formula of cross-modal feature alignment loss is:
[0084]
[0085] Where, is the set of known basic categories, L is the number of known basic categories; A text description set (including basic categories and generated new category descriptions); The text embedding feature for the j-th text description; Basic category Corresponding text embedding features; The text description generated for the local region of interest i of unknown category; is the weight parameter.
[0086] Step 160: Perform object detection on the image to be detected using an open object detection model.
[0087] The image to be detected is input into the open object detection model, and the local region of interest is extracted through the region proposal network in the open object detection model. The local region of interest is aligned with the feature deformation adapter to obtain the adapted local region features; the adapted local region features are processed by the image encoder to obtain the image embedding features; the similarity between the image embedding features and the text embedding features in the text library is calculated, and the text category corresponding to the text embedding feature with the highest similarity is used as the target category in the local region of interest to obtain the target detection result.
[0088] This application uses a two-layer mechanism of local instance distillation combined with global feature alignment, combined with the cross-modal semantic representation capability of the CLIP model, to provide a label generation path based on semantic similarity for unknown targets while retaining the recognition accuracy of basic classes. It effectively solves the zero-sample generalization problem in the classification of new targets and is particularly suitable for target detection tasks in multi-category and complex scenarios.
[0089] This application designs an instance deformation adapter to reduce the feature semantic differences caused by local ROI feature deformation and enhance the deformation feature alignment capability; in addition, it proposes a granular hierarchical distillation framework to fully utilize the multi-level semantic features of the image, build a cross-level knowledge transfer link, and effectively improve the detection generalization capability of the open object detection model for unknown categories.
[0090] Please refer to Figure 2 , the embodiment of the present application further provides an open target detection device, comprising:
[0091] A network construction unit 210 is used to construct a target network, which includes a region proposal network, a feature deformation adapter, an image encoder, and a text encoder;
[0092] A feature adaptation unit 220 is configured to extract a local region of interest from a training image using a region proposal network, perform feature alignment on the local region of interest using a feature deformation adapter, and obtain adapted local region features;
[0093] A first loss calculation unit 230 is configured to input the adapted local region features into an image encoder for processing to generate image embedding features, and calculate a local instance-level distillation loss using the image embedding features and the adapted local region features;
[0094] A second loss calculation unit 240 is configured to input the text description corresponding to the training image into a text encoder for processing, generate text embedding features, and calculate a cross-modal feature alignment loss using the image embedding features and the text embedding features;
[0095] A parameter updating unit 250 is used to update the network parameters of the target network by combining the local instance-level distillation loss and the cross-modal feature alignment loss to obtain a trained open object detection model;
[0096] The object detection unit 260 is configured to perform object detection on the image to be detected using an open object detection model.
[0097] As a further improvement, the number of feature deformation adapters is multiple, and the feature deformation adapter includes a feature adapter and a deformation adapter. The feature adapter consists of a pooling layer, an activation layer, and a fully connected layer. All feature adapters share the pooling layer and the activation layer.
[0098] The feature adapter is used to perform feature mapping on the local region of interest and perform residual connection with the local region of interest to obtain local region features;
[0099] The deformation adapter is used to select the corresponding deformation adaptation weight according to the discrete interval where the aspect ratio of the local area feature is located to weight the local area feature to obtain the adapted local area feature.
[0100] As a further improvement, the apparatus further includes a third loss calculation unit, configured to perform global feature extraction on the training image through an image encoder to obtain a first global feature;
[0101] Perform average pooling and feature aggregation on the adapted local area features corresponding to the training image to obtain the second global feature;
[0102] Calculate the cosine similarity between the first global feature and the second global feature to obtain the global feature training loss;
[0103] Correspondingly, the parameter update unit is specifically used to update the network parameters of the target network by combining the local instance-level distillation loss, the cross-modal feature alignment loss and the global feature training loss to obtain a trained open object detection model.
[0104] This application designs an instance deformation adapter to reduce the feature semantic differences caused by local ROI feature deformation and enhance the deformation feature alignment capability; in addition, it proposes a granular hierarchical distillation framework to fully utilize the multi-level semantic features of the image, build a cross-level knowledge transfer link, and effectively improve the detection generalization capability of the open object detection model for unknown categories.
[0105] Please refer to Figure 3 , an embodiment of the present application further provides an electronic device, the device including a processor 310 and a memory 320;
[0106] The memory 320 is used to store program codes and transmit the program codes to the processor 310;
[0107] The processor 310 is configured to execute the open object detection method in the aforementioned method embodiment according to instructions in the program code.
[0108] An embodiment of the present application further provides a computer-readable storage medium, which is used to store program code. When the program code is executed by a processor, the open target detection method in the aforementioned method embodiment is implemented.
[0109] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0110] In the specification of this application and the above-mentioned drawings, the terms "first," "second," "third," "fourth," etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.
[0111] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.
[0112] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0113] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0114] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0115] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for executing all or part of the steps of the method described in each embodiment of the present application through a computer device (which can be a personal computer, server, or network device, etc.). The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (full name: Read-Only Memory, English abbreviation: ROM), random access memory (full name: Random Access Memory, English abbreviation: RAM), disk or optical disk, and other media that can store program code.
[0116] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An open target detection method, characterized in that: include: Constructing a target network, wherein the target network includes a region proposal network, a feature deformation adapter, an image encoder, and a text encoder; Extracting local regions of interest from training images through a region proposal network, and aligning the features of the local regions of interest through a feature deformation adapter to obtain adapted local region features; Inputting the adapted local region features into an image encoder for processing to generate image embedding features, and calculating a local instance-level distillation loss using the image embedding features and the adapted local region features; Inputting the text description corresponding to the training image into a text encoder for processing to generate text embedding features, and calculating a cross-modal feature alignment loss using the image embedding features and the text embedding features; Combining the local instance-level distillation loss and the cross-modal feature alignment loss to update the network parameters of the target network to obtain a trained open object detection model; Object detection is performed on the image to be detected using the open object detection model.
2. The open target detection method according to claim 1, characterized in that: There are multiple feature deformation adapters, each of which includes a feature adapter and a deformation adapter. The feature adapter consists of a pooling layer, an activation layer, and a fully connected layer, and all feature adapters share the pooling layer and the activation layer. The feature adapter is used to perform feature mapping on the local region of interest and perform residual connection with the local region of interest to obtain local region features; The deformation adapter is used to select a corresponding deformation adaptation weight according to a discrete interval where the aspect ratio of the local area feature is located to weight the local area feature to obtain an adapted local area feature.
3. The open target detection method according to claim 1, characterized in that: The method further comprises: Performing global feature extraction on the training image by the image encoder to obtain a first global feature; Performing average pooling and feature aggregation on the adapted local area features corresponding to the training image to obtain a second global feature; Calculating the cosine similarity between the first global feature and the second global feature to obtain a global feature training loss; Accordingly, the network parameters of the target network are updated by combining the local instance-level distillation loss and the cross-modal feature alignment loss to obtain a trained open object detection model, including: The local instance-level distillation loss, the cross-modal feature alignment loss, and the global feature training loss are combined to update the network parameters of the target network to obtain a trained open object detection model.
4. The open target detection method according to claim 1, wherein: The method of extracting a local region of interest from a training image through a region proposal network includes: Processing the training image through a region proposal network to generate candidate boxes with confidence scores; From the candidate frames, candidate frames whose intersection-over-union ratio with the real frames of the base category is lower than the intersection-over-union ratio threshold and whose confidence scores are higher than the confidence threshold are screened to obtain new category candidate frames; the area corresponding to the new category candidate frames is the local region of interest.
5. The open target detection method according to claim 4, characterized in that: The calculation formula of the local instance-level distillation loss is: Where N is the number of candidate frames of the new class target, M is the number of real frames of the basic class, is the local region feature of the i-th local region of interest after deformation adaptation; The text embedding feature for the j-th text description; The text embedding features generated for the region of interest of the base category, Embed features for text and the adapted local area features The cosine similarity between .
6. The open target detection method according to claim 1, characterized in that: The calculating the cross-modal feature alignment loss using the image embedding feature and the text embedding feature includes: The text embedding features of the text description and the corresponding image embedding features are used as positive sample pairs, and the text embedding features of the text description and the non-corresponding image embedding features are used as negative sample pairs. The cosine similarity between the text embedding features and the image embedding features is calculated to obtain the cross-modal feature alignment loss.
7. The open target detection method according to claim 1, characterized in that: The calculation formula of the cross-modal feature alignment loss is: Where, is the set of known basic categories, L is the number of known basic categories; Describes the collection of texts; The text embedding feature for the j-th text description; Basic category Corresponding text embedding features; Text embedding features generated for the region of interest of the base category; The text description generated for the local region of interest i of unknown category; is the weight parameter.
8. An open target detection device, characterized in that: include: A network construction unit, configured to construct a target network, wherein the target network includes a region proposal network, a feature deformation adapter, an image encoder, and a text encoder; A feature adaptation unit is used to extract local regions of interest from training images using a region proposal network, and perform feature alignment on the local regions of interest using a feature deformation adapter to obtain adapted local region features; a first loss calculation unit, configured to input the adapted local region features into an image encoder for processing to generate image embedding features, and calculate a local instance-level distillation loss using the image embedding features and the adapted local region features; a second loss calculation unit, configured to input the text description corresponding to the training image into a text encoder for processing, generate text embedding features, and calculate a cross-modal feature alignment loss using the image embedding features and the text embedding features; a parameter updating unit, configured to update the network parameters of the target network by combining the local instance-level distillation loss and the cross-modal feature alignment loss to obtain a trained open object detection model; The target detection unit is used to perform target detection on the image to be detected using the open target detection model.
9. An electronic device, characterized in that: The device includes a processor and a memory; The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the open object detection method according to any one of claims 1 to 7 according to instructions in the program code.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and when the program code is executed by a processor, the open object detection method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Semantic knowledge guided vehicle re-identification method
CN118230321A
Open vocabulary target detection method and system based on multi-target classification, medium and program product
CN119919634A
Medical image generation method and device based on bimodal fusion, equipment and medium
CN120047578A
Water surface target detection method and system based on visual language large model
CN120107690A
Image processing method, electronic device and storage medium
US20220253631A1