Target detection model training method, device and electronic equipment

Through the combination of feature extraction and aggregation network, the efficiency of multimodal object detection model training is improved, and the problem of low training efficiency under multiple visual cue images is solved.

CN119693775BActive Publication Date: 2025-05-02HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510191809.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-02
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

In multimodal object detection model training, the training efficiency is low through feature extraction and aggregation of sample images and visual cue images, especially when there are many visual cue images.

Method used

A training method for object detection model is proposed, which extracts visual cue features through feature extraction networks, and uses feature aggregation network to aggregate reference features and visual cue features to generate aggregated visual cue features for training models.

Benefits of technology

Through this method, the efficiency of object detection model training is improved, and feature information can be used more effectively, especially when processing multiple visual prompt images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693775B_ABST
    Figure CN119693775B_ABST
Patent Text Reader

Abstract

The present application provides a method, device and electronic device for training a target detection model. In the present application, when a sample image and multiple visual cue images corresponding to the sample image are obtained, for each category, a feature extraction network is used to perform feature extraction on multiple visual cue images corresponding to the category to obtain multiple visual cue features corresponding to the category, and a feature aggregation network is used to perform feature aggregation on reference features corresponding to the category and multiple visual cue features corresponding to the category to obtain aggregated visual cue features corresponding to the category, and a model to be trained is trained according to the aggregated visual cue features corresponding to each category to obtain a target detection model, and a reference feature corresponding to each category is configured in the feature aggregation network so that for each category, multiple cue features corresponding to the category can be aggregated into the reference feature to obtain an aggregated visual cue feature representing the category, thereby improving the efficiency of model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of target detection technology, and in particular to a target detection model training method, device and electronic equipment. Background Art

[0002] In the multimodal target detection model training scenario, the target detection model is often trained through sample images and visual cue images corresponding to the sample images. The current multimodal target detection model training method obtains sample visual features of sample images and visual cue features of visual cue images through feature extraction, wherein a sample image may include multiple visual cue images corresponding to the multiple visual cue images, and multiple visual cue features may be extracted from the multiple visual cue images. The efficiency of training the training model through the sample visual features of the sample image and each visual cue feature is low. Summary of the invention

[0003] In view of this, the present application provides a target detection model training method, device and electronic device to improve the efficiency of model training.

[0004] The technical solutions provided by this application are as follows:

[0005] According to an embodiment of the first aspect of the present application, a method for training a target detection model is provided, wherein the model to be trained includes a feature extraction network and a feature aggregation network, and the method includes:

[0006] Acquire a plurality of visual cue images corresponding to the sample image, wherein for each visual cue image, the visual cue image includes at least one category of reference objects;

[0007] For each category, a feature extraction network is used to extract features from multiple visual cue images corresponding to the category to obtain multiple visual cue features corresponding to the category;

[0008] For each category, a feature aggregation network is used to perform feature aggregation on a reference feature corresponding to the category and a plurality of visual cue features corresponding to the category, so as to obtain an aggregated visual cue feature corresponding to the category; wherein the feature aggregation network is configured with a reference feature corresponding to each category;

[0009] The model to be trained is trained according to the aggregated visual cue features corresponding to each category to obtain a trained target detection model; wherein the target detection model is used to detect the image to be tested to obtain the target category of the object in the image to be tested.

[0010] Optionally, the extracting features of the multiple visual cue images corresponding to the category by a feature extraction network to obtain multiple visual cue features corresponding to the category includes:

[0011] For each visual cue image, extract features of the visual cue image through the feature extraction network to obtain an overall visual feature corresponding to the visual cue image;

[0012] For each category, the sub-visual features of the reference object corresponding to the category are extracted from the overall visual features corresponding to each visual image to obtain the visual prompt features corresponding to the category.

[0013] Optionally, the feature aggregation network includes at least one feature interaction layer and at least one feature aggregation layer; the feature aggregation network performs feature aggregation on the reference feature corresponding to the category and multiple visual cue features corresponding to the category to obtain the aggregated visual cue feature corresponding to the category, including: for each feature interaction layer, obtaining the first input feature of the feature interaction layer, the first input feature includes the reference features of all categories; for each category, performing similarity calculation on the reference feature of the category and the reference feature of each category through the feature interaction layer, determining the weight corresponding to the reference feature of each category according to the similarity between the reference feature of the category and the reference feature of each category, and performing weighted summation on the reference feature of the category and the reference features of other categories according to the weight to obtain the interaction feature of the category;

[0014] For each feature aggregation layer, a second input feature of the feature aggregation layer is obtained, wherein the second input feature includes interaction features of all categories and multiple visual cue features corresponding to each category; for each category, a similarity calculation is performed on the interaction features of the category and the multiple visual cue features corresponding to the category through the feature aggregation layer, and a weight corresponding to each visual cue feature corresponding to the reference feature of the category is determined according to the similarity between the reference feature of the category and the multiple visual cue features corresponding to the category, and a weighted sum is performed on the reference feature of the category and the multiple visual cue features corresponding to the category according to the weight to obtain a fusion feature of the category.

[0015] Optionally, the feature aggregation network includes multiple network layers, the 1st to Mth network layers among the multiple network layers are feature interaction layers, and starting from the M+1th network layer, they are alternately feature aggregation layers and feature interaction layers.

[0016] Optionally, the model to be trained further includes a backbone network, a feature enhancement network and a target decoder, and the model to be trained is trained according to the aggregated visual cue features corresponding to each category to obtain a trained target detection model, including:

[0017] Inputting the sample image into the backbone network to obtain sample image features, wherein the sample image includes a target object of at least one category;

[0018] The sample image features and the aggregated visual cue features corresponding to each category are input into a feature enhancement network, and similarity calculation is performed on the sample image features and the aggregated visual cue features corresponding to each category through the feature enhancement network, and the weights of the sample image features and the aggregated visual cue features corresponding to each category are determined according to the similarity between the sample image features and the aggregated visual cue features corresponding to each category, and a weighted sum is performed according to the weights of the sample image features and the aggregated visual cue features corresponding to each category to obtain the sample visual enhancement feature; and for each category, the aggregated visual enhancement feature corresponding to the category is obtained according to the weight of the aggregated visual cue features corresponding to the category and the weight of the sample image features;

[0019] Inputting the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category into a target decoder, and determining the similarity between the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category by the target decoder, so as to obtain a visual cue prediction category of the target object in the sample image according to the similarity;

[0020] The model to be trained is trained based on the visual cue prediction category of the target object to obtain a trained target detection model.

[0021] Optionally, the model to be trained further includes a text encoder, and the method further includes:

[0022] Outputting the text prompt to a text encoder to obtain text prompt features corresponding to each category, wherein the process of outputting the text prompt to the text encoder and the process of extracting features of multiple visual prompt images corresponding to each category through a feature extraction network are performed alternately;

[0023] Input the text prompt features corresponding to each category and the sample image features output by the backbone network into the feature enhancement network, calculate the similarity between the sample image features and the text prompt features corresponding to each category through the feature enhancement network, determine the weights of the sample image features and the text prompt features corresponding to each category according to the similarity between the sample image features and the text prompt features corresponding to each category, and perform weighted summation according to the weights of the sample image features and the text prompt features corresponding to each category to obtain the sample text enhancement features; and for each category, obtain the text prompt enhancement features corresponding to the category according to the weight of the text prompt features corresponding to the category and the weight of the sample image features;

[0024] The sample text enhancement feature and the text prompt enhancement feature corresponding to each category are input into a target decoder, and the target decoder determines the similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category, so as to obtain the text prompt prediction category of the target object in the sample image according to the similarity;

[0025] The model to be trained is trained based on the text prompt prediction category and the visual prompt prediction category to obtain a trained object detection model.

[0026] Optionally, the model to be trained is obtained by the following method:

[0027] The initial network model is obtained by training, and the initial network model includes a backbone network, a feature enhancement network, a target decoder and a text encoder; the feature aggregation network is generated based on the text encoder, and the feature aggregation network and the feature extraction network are added to the initial network model to obtain the model to be trained; the text encoder includes multiple feature interaction layers, and the feature aggregation network is obtained by inserting at least one feature aggregation layer between the multiple feature interaction layers included in the text encoder;

[0028] The training process of the initial network model includes:

[0029] Output the text prompt to the text encoder to obtain the text prompt features corresponding to each category;

[0030] Inputting the text prompt features and sample image features corresponding to each category into the feature enhancement network; the sample image features are obtained by extracting features of the sample images through the backbone network;

[0031] The sample image features and the text prompt features corresponding to each category are enhanced through the feature enhancement network to obtain the sample text enhancement features and the text prompt enhancement features corresponding to each category;

[0032] The similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category is determined by the target decoder, and the text prompt prediction category of the target object included in the sample image is determined according to the similarity; the initial network model is obtained by training according to the text prompt prediction category.

[0033] According to an embodiment of the second aspect of the present application, a target detection model training device is provided, wherein the model to be trained includes a feature extraction network and a feature aggregation network, and the device includes:

[0034] An acquisition unit, configured to acquire a plurality of visual cue images corresponding to the sample image, wherein for each visual cue image, the visual cue image includes at least one category of reference objects;

[0035] An extraction unit is used for extracting features from a plurality of visual cue images corresponding to each category through a feature extraction network to obtain a plurality of visual cue features corresponding to the category;

[0036] an aggregation unit, for performing feature aggregation on a reference feature corresponding to the category and a plurality of visual cue features corresponding to the category through a feature aggregation network for each category, so as to obtain an aggregated visual cue feature corresponding to the category; wherein the feature aggregation network is configured with a reference feature corresponding to each category;

[0037] A training unit is used to train the model to be trained according to the aggregated visual cue features corresponding to each category to obtain a trained target detection model; wherein the target detection model is used to detect the image to be tested to obtain the target category of the object in the image to be tested.

[0038] Optionally, the extraction unit is specifically used for:

[0039] For each visual cue image, extract features of the visual cue image through the feature extraction network to obtain an overall visual feature corresponding to the visual cue image;

[0040] For each category, the sub-visual features of the reference object corresponding to the category are intercepted from the overall visual features corresponding to each visual image to obtain the visual prompt features corresponding to the category;

[0041] And / or, the polymeric unit is specifically used for:

[0042] For each feature interaction layer, a first input feature of the feature interaction layer is obtained, wherein the first input feature includes reference features of all categories; for each category, similarity calculation is performed on the reference feature of the category and the reference feature of each category through the feature interaction layer, a weight corresponding to the reference feature of each category is determined according to the similarity between the reference feature of the category and the reference feature of each category, and a weighted sum is performed on the reference feature of the category and the reference features of other categories according to the weight to obtain the interaction feature of the category;

[0043] For each feature aggregation layer, a second input feature of the feature aggregation layer is obtained, wherein the second input feature includes interaction features of all categories and a plurality of visual cue features corresponding to each category; for each category, similarity calculation is performed on the interaction features of the category and the plurality of visual cue features corresponding to the category through the feature aggregation layer, and a weight corresponding to each visual cue feature corresponding to the reference feature of the category is determined according to the similarity between the reference feature of the category and the plurality of visual cue features corresponding to the category, and a weighted sum is performed on the reference feature of the category and the plurality of visual cue features corresponding to the category according to the weight, so as to obtain a fusion feature of the category;

[0044] And / or, the feature aggregation network includes multiple network layers, the first to M network layers of the multiple network layers are feature interaction layers, and starting from the M+1th network layer, they are alternately feature aggregation layers and feature interaction layers;

[0045] And / or, the model to be trained further includes a backbone network, a feature enhancement network and a target decoder, and the training unit is specifically used for:

[0046] Inputting the sample image into the backbone network to obtain sample image features, wherein the sample image includes a target object of at least one category;

[0047] The sample image features and the aggregated visual cue features corresponding to each category are input into a feature enhancement network, and similarity calculation is performed on the sample image features and the aggregated visual cue features corresponding to each category through the feature enhancement network, and the weights of the sample image features and the aggregated visual cue features corresponding to each category are determined according to the similarity between the sample image features and the aggregated visual cue features corresponding to each category, and a weighted sum is performed according to the weights of the sample image features and the aggregated visual cue features corresponding to each category to obtain the sample visual enhancement feature; and for each category, the aggregated visual enhancement feature corresponding to the category is obtained according to the weight of the aggregated visual cue features corresponding to the category and the weight of the sample image features;

[0048] Inputting the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category into a target decoder, and determining the similarity between the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category by the target decoder, so as to obtain a visual cue prediction category of the target object in the sample image according to the similarity;

[0049] Training the model to be trained based on the visual cue prediction category of the target object to obtain a trained target detection model;

[0050] And / or, the model to be trained further includes a text encoder, and the training unit is further used for:

[0051] Outputting the text prompt to a text encoder to obtain text prompt features corresponding to each category, wherein the process of outputting the text prompt to the text encoder and the process of extracting features of multiple visual prompt images corresponding to each category through a feature extraction network are performed alternately;

[0052] Input the text prompt features corresponding to each category and the sample image features output by the backbone network into the feature enhancement network, calculate the similarity between the sample image features and the text prompt features corresponding to each category through the feature enhancement network, determine the weights of the sample image features and the text prompt features corresponding to each category according to the similarity between the sample image features and the text prompt features corresponding to each category, and perform weighted summation according to the weights of the sample image features and the text prompt features corresponding to each category to obtain the sample text enhancement features; and for each category, obtain the text prompt enhancement features corresponding to the category according to the weight of the text prompt features corresponding to the category and the weight of the sample image features;

[0053] The sample text enhancement feature and the text prompt enhancement feature corresponding to each category are input into a target decoder, and the target decoder determines the similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category, so as to obtain the text prompt prediction category of the target object in the sample image according to the similarity;

[0054] Training the model to be trained based on the text prompt prediction category and the visual prompt prediction category to obtain a trained object detection model;

[0055] And / or, the model to be trained is obtained by the following method:

[0056] The initial network model is obtained by training, and the initial network model includes a backbone network, a feature enhancement network, a target decoder and a text encoder; the feature aggregation network is generated based on the text encoder, and the feature aggregation network and the feature extraction network are added to the initial network model to obtain the model to be trained; the text encoder includes multiple feature interaction layers, and the feature aggregation network is obtained by inserting at least one feature aggregation layer between the multiple feature interaction layers included in the text encoder;

[0057] The training process of the initial network model includes:

[0058] Output the text prompt to the text encoder to obtain the text prompt features corresponding to each category;

[0059] Inputting the text prompt features and sample image features corresponding to each category into the feature enhancement network; the sample image features are obtained by extracting features of the sample images through the backbone network;

[0060] The sample image features and the text prompt features corresponding to each category are enhanced through the feature enhancement network to obtain the sample text enhancement features and the text prompt enhancement features corresponding to each category;

[0061] The similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category is determined by the target decoder, and the text prompt prediction category of the target object included in the sample image is determined according to the similarity; the initial network model is obtained by training according to the text prompt prediction category.

[0062] According to an embodiment of the third aspect of the present application, an electronic device is provided, the electronic device comprising:

[0063] A processor and a computer-readable storage medium, wherein the computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the processor executes the method described in the first aspect.

[0064] It can be seen from the above technical scheme that when obtaining a sample image and multiple visual cue images corresponding to the sample image, the present application, for each category, performs feature extraction on the multiple visual cue images corresponding to the category through a feature extraction network to obtain multiple visual cue features corresponding to the category, and for each category, performs feature aggregation on the reference feature corresponding to the category and the multiple visual cue features corresponding to the category through a feature aggregation network to obtain the aggregated visual cue features corresponding to the category, and trains the model to be trained according to the aggregated visual cue features corresponding to each category to obtain a trained target detection model, and sets a feature aggregation network and configures the reference features corresponding to each category for the feature aggregation network, so that for each category, the multiple cue features corresponding to the category can be aggregated into the reference feature to obtain the aggregated visual cue features representing the category, thereby greatly improving the efficiency of model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0066] Figure 1 A flowchart of the target detection model training method provided in the embodiment of the present application;

[0067] Figure 2 A schematic diagram of a feature aggregation network structure provided in an embodiment of the present application;

[0068] Figure 3 A schematic diagram of a target detection model training framework provided in an embodiment of the present application;

[0069] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0070] Figure 5A structural diagram of a target detection model training device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0071] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings.

[0072] Please refer to Figure 1 , Figure 1 A flowchart of the target detection model training method provided in an embodiment of the present application.

[0073] It should be noted that the target detection model provided in the present application refers to a model that can detect the position and / or category of the target object in the image to be detected.

[0074] Before training begins, you can pre-set the category of objects to be detected. The category can be one or more. For example, if the category is set to "cat" and "dog", the trained object detection model can recognize the cats and dogs in the image to be tested and mark their corresponding positions.

[0075] like Figure 1 As shown, the method may include the following steps:

[0076] Step 101: Acquire multiple visual cue images corresponding to a sample image.

[0077] In this embodiment, the sample image includes at least one category of target objects, and the visual cue image refers to an image including visual cue information related to the sample image, that is, the visual cue image includes at least one category of reference objects.

[0078] For example, if the sample image includes a dog, that is, the target object is "dog", then the visual cue image may be a picture of dogs of various types or colors. The visual cue image may also include the sample image itself, and this application does not impose any restrictions on this.

[0079] In this embodiment, the target object in the sample image may be one or more than one target object, and the preset categories to which the multiple target objects belong may be one or more than one category.

[0080] Correspondingly, when the sample image includes target objects of multiple categories, the multiple visual cue images corresponding to the sample image may also be that each category of target object corresponds to multiple visual cue images, that is, each visual cue image includes at least one reference object of at least one category.

[0081] For example, the preset categories are cats and dogs, and the sample image includes two cats and a dog. Both "cats" and "dogs" can be taken as target objects. At this time, the multiple visual cue images corresponding to the sample image may include multiple visual cue images corresponding to the target objects of each preset category, and the visual cue images obtained for the target object "dog" include at least one dog. In addition, it may also include multiple dogs, or objects of other categories such as cats. The present application does not impose any restrictions on this.

[0082] At this point, the description of step 101 ends, and step 102 is executed next.

[0083] Step 102: for each category, a feature extraction network is used to perform feature extraction on a plurality of visual cue images corresponding to the category to obtain a plurality of visual cue features corresponding to the category.

[0084] In this embodiment, feature extraction may be performed on the multiple visual cue images obtained in step 101 through a feature extraction network.

[0085] It should be noted that when performing feature extraction on visual cue images, since different visual cue images may include reference objects of different preset categories, and the same visual cue image may also include one or more reference objects of different preset categories, for each visual cue image, it is necessary to perform feature extraction on each preset category of reference objects through a feature extraction network.

[0086] As an embodiment, a method for extracting features from a plurality of visual cue images corresponding to the category through a feature extraction network to obtain a plurality of visual cue features corresponding to the category may include:

[0087] For each visual cue image, extract features of the visual cue image through a feature extraction network to obtain the overall visual features corresponding to the visual cue image;

[0088] For each category, the sub-visual features of the reference object corresponding to the category are extracted from the overall visual features corresponding to each visual image to obtain the visual prompt features corresponding to the category.

[0089] For example, if a visual cue image includes a cat and a dog, you can first perform feature extraction on the entire visual cue image to obtain the overall visual feature, and then further extract the sub-visual features belonging to the cat in the overall visual feature as the visual cue feature corresponding to the preset type "cat", and then extract the sub-visual features belonging to the dog in the overall visual feature as the visual cue feature corresponding to the preset type "dog".

[0090] After performing the above feature extraction on the reference object of each category in each visual cue image, a plurality of visual cue features corresponding to each category are obtained.

[0091] It should be noted that, by first performing overall feature extraction on the visual cue image and then intercepting the extracted overall visual features according to the reference objects corresponding to each category, the contextual information around the reference object is retained, that is, the influence of the surrounding images of the reference object on the reference object is retained, which can enhance the generalization of feature extraction and avoid overfitting.

[0092] At this point, the description of step 102 ends, and step 103 is executed next.

[0093] Step 103: for each category, a feature aggregation network is used to perform feature aggregation on a reference feature corresponding to the category and a plurality of visual cue features corresponding to the category to obtain an aggregated visual cue feature corresponding to the category.

[0094] In this embodiment, the feature aggregation network is configured with reference features corresponding to each category, and the reference features are initially initialized features. The feature aggregation network is used to aggregate multiple visual cue features corresponding to each preset category and the reference features corresponding to the preset category to obtain an aggregated visual cue feature corresponding to the reference category.

[0095] Specifically, the feature aggregation network includes at least one feature interaction layer and at least one feature aggregation layer; the feature aggregation network performs feature aggregation on the reference feature corresponding to the category and the multiple visual cue features corresponding to the category, and the specific method of obtaining the aggregated visual cue features corresponding to the category may include:

[0096] For each feature interaction layer, a first input feature of the feature interaction layer is obtained, and the first input feature includes reference features of all categories; for each category, similarity calculation is performed on the reference feature of the category and the reference feature of each category through the feature interaction layer, and the weight corresponding to the reference feature of each category is determined according to the similarity between the reference feature of the category and the reference feature of each category, and the reference feature of the category and the reference features of other categories are weighted summed according to the weight to obtain the interaction feature of the category;

[0097] For each feature aggregation layer, a second input feature of the feature aggregation layer is obtained, and the second input feature includes interaction features of all categories and multiple visual cue features corresponding to each category; for each category, similarity is calculated for the interaction features of the category and the multiple visual cue features corresponding to the category through the feature aggregation layer, and the weight corresponding to each visual cue feature corresponding to the reference feature of the category is determined according to the similarity between the reference feature of the category and the multiple visual cue features corresponding to the category, and the reference feature of the category and the multiple visual cue features corresponding to the category are weightedly summed according to the weight to obtain a fusion feature of the category.

[0098] In this embodiment, the input of each feature interaction layer (referred to as the first input feature) can be a reference feature corresponding to all categories, or an interaction feature corresponding to all categories output by other feature interaction layers, or a fusion feature corresponding to all categories output by any feature aggregation layer.

[0099] Specifically, if the feature interaction layer is the first network layer in the feature aggregation network, its first input feature is the reference feature corresponding to all categories. At this time, for each category, the feature interaction layer can be used to calculate the similarity between the reference feature of the category and the reference feature of each category, and the weight corresponding to the reference feature of each category is determined according to the similarity between the reference feature of the category and the reference feature of each category. The reference feature of the category and the reference features of other categories are weightedly summed according to the weight to obtain the interaction feature of the category.

[0100] If the feature interaction layer is a network layer after other feature interaction layers, then its first input feature is the output of the feature interaction layer, that is, the interaction features corresponding to all categories. At this time, the interaction features corresponding to each category can be used as the reference features corresponding to each category, and the above-mentioned steps of performing feature interaction on the reference features corresponding to each category are executed, which will not be repeated here.

[0101] If the feature interaction layer is a network layer after any feature aggregation layer, its first input feature is the output of the feature aggregation layer, that is, the fused features corresponding to all categories. At this time, the fused features corresponding to each category can be used as the reference features corresponding to each category, and the above-mentioned steps of performing feature interaction on the reference features corresponding to each category are executed, which will not be repeated here.

[0102] Similarly, the interaction features corresponding to each category obtained through the feature interaction layer can be used as input features of the next feature interaction layer, or the interaction features of all categories can be used as input features of the next feature aggregation layer, or the interaction features of all categories can be used as aggregated visual cue features.

[0103] In this embodiment, the input of each feature aggregation layer (recorded as the second input feature) can be the interactive features corresponding to all categories, or can be the fusion features corresponding to all categories output by other feature aggregation layers.

[0104] Specifically, if the feature aggregation layer is the network layer after the feature interaction layer, its second input feature is the interaction feature corresponding to all categories. At this time, for each category, the feature aggregation layer can be used to calculate the similarity between the interaction feature of the category and the multiple visual cue features corresponding to the category. According to the similarity between the reference feature of the category and the multiple visual cue features corresponding to the category, the weight of each visual cue feature corresponding to the reference feature of the category is determined, and the reference feature of the category and the multiple visual cue features corresponding to the category are weightedly summed according to the weight to obtain the fusion feature of the category.

[0105] If the feature interaction layer is a network layer after other feature aggregation layers, then its second input feature is the output of the feature aggregation layer, that is, the fused features corresponding to all categories. At this time, the fused features corresponding to each category can be used as the interaction features corresponding to each category, and the above-mentioned steps of feature fusion of the interaction features corresponding to each category are performed, which will not be repeated here.

[0106] Similarly, the fused features corresponding to each category obtained through the feature aggregation layer can be used as the input features of the next feature interaction layer, or the fused features of all categories can be used as the input features of the next feature aggregation layer, or the fused features of all categories can be used as aggregated visual cue features.

[0107] As an embodiment, the feature aggregation network includes multiple network layers, the 1st to Mth network layers among the multiple network layers are feature interaction layers, and starting from the M+1th network layer, they alternately become feature aggregation layers and feature interaction layers.

[0108] As an embodiment, the feature aggregation network can reuse the structure and parameters of a trained text encoder, that is, it is obtained by inserting at least one feature aggregation layer between multiple feature interaction layers included in the text encoder.

[0109] For example, please refer to Figure 2 , Figure 2 A schematic diagram of a feature aggregation network structure provided in an embodiment of the present application.

[0110] like Figure 2 As shown, the text encoder includes 12 feature interaction layers, and a feature aggregation layer is added between every two feature interaction layers from the 6th feature interaction layer to the 12th feature interaction layer;

[0111] Among them, the feature exchange layer can adopt the Self-Attention layer structure, and the feature aggregation layer can adopt the Cross-attention layer structure.

[0112] Specifically, if there are n preset categories, n reference features can be initialized first, corresponding to the n preset categories, respectively, and denoted as [MASK1], [MASK2], ... [MASKn]. These reference features first pass through 6 feature interaction layers, and a feature interaction is performed in each feature interaction layer to improve the distinctiveness of the reference features corresponding to each preset category. Starting from the 7th layer, before entering the feature interaction layer, it first enters the feature aggregation layer.

[0113] In the feature aggregation layer, for each reference feature corresponding to a preset category, the reference feature corresponding to the preset category obtained through feature interaction is interacted and fused with multiple visual cue features corresponding to the preset category, so that the aggregated reference feature can better represent the preset category.

[0114] After 12 feature interaction layers and 6 feature aggregation layers (multiple interactions and fusions), an aggregate visual cue feature corresponding to each preset category is obtained (corresponding to Figure 2 in which the feature vector of visual cue category 1 to the feature vector of visual cue category n).

[0115] Among them, the structure and weights of the 12-layer feature interaction all come from the trained text encoder. Compared with training a dedicated image encoder from scratch, this can significantly accelerate the network convergence speed and improve the final performance of the network.

[0116] This concludes Figure 2 Description.

[0117] In this embodiment, the trained text encoder will be described in detail below and will not be described again here.

[0118] At this point, the description of step 103 ends, and step 104 is executed next.

[0119] Step 104 , training the model to be trained according to the aggregated visual cue features corresponding to each category to obtain a trained object detection model.

[0120] In this embodiment, the target detection model is used to detect the image to be detected to obtain the target category of the object in the image to be detected.

[0121] As an embodiment, the model to be trained also includes a backbone network, a feature enhancement network and a target decoder. The model to be trained is trained according to the aggregated visual cue features corresponding to each category. The specific method of obtaining the trained target detection model includes:

[0122] Inputting a sample image into the backbone network to obtain sample image features, wherein the sample image includes a target object of at least one category;

[0123] The sample image features and the aggregated visual cue features corresponding to each category are input into the feature enhancement network, and the similarity between the sample image features and the aggregated visual cue features corresponding to each category is calculated through the feature enhancement network, and the weights of the sample image features and the aggregated visual cue features corresponding to each category are determined according to the similarity between the sample image features and the aggregated visual cue features corresponding to each category, and the sample image features and the aggregated visual cue features corresponding to each category are weighted summed to obtain the sample visual enhancement features; and for each category, the aggregated visual enhancement features corresponding to the category are obtained according to the weight of the aggregated visual cue features corresponding to the category and the weight of the sample image features;

[0124] The sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category are input into the target decoder, and the similarity between the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category is determined by the target decoder, so as to obtain the visual cue prediction category of the target object in the sample image according to the similarity;

[0125] The model to be trained is trained based on the visual cue prediction category of the target object to obtain a trained object detection model.

[0126] In this embodiment, the sample image features of the sample image are obtained by performing feature extraction on the sample image through a backbone network. The sample image features of the sample image can be obtained by performing feature extraction on the sample image through any backbone network, such as through a convolutional neural network CNN, a deep convolutional neural network, etc. This application does not impose any restrictions on this.

[0127] In this embodiment, the predicted position and predicted category of the target object included in the sample image can be determined by using the aggregated visual cue features corresponding to each reference category obtained in step 103 and the sample image features of the sample image.

[0128] In this embodiment, the aggregated visual cue features corresponding to each reference category and the sample image features of the sample image can be enhanced through a feature enhancement network so that features with high similarity are more similar, and features with low similarity are more different.

[0129] After the feature enhancement is completed, the obtained aggregated visual enhancement features corresponding to each reference category and the sample visual enhancement features of the sample image may be input into the target decoder.

[0130] The target decoder can determine the similarity relationship between the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each reference category, take the reference category corresponding to the aggregated visual enhancement feature with the highest similarity to the sample visual enhancement feature as the predicted category of the target object in the sample image, and take the area corresponding to the sample visual enhancement feature as the area where the target object is located.

[0131] When there are multiple target objects, the similarity relationship between multiple local features of the sample visual feature and the aggregated visual cue features corresponding to each reference category can be determined. For each local feature, the reference category corresponding to the aggregated visual cue feature with the highest similarity to the local feature is used as the predicted category of the target object corresponding to the local feature in the sample image, and the area corresponding to the local feature in the sample image is used as the area where the target object is located.

[0132] In this embodiment, the predicted category obtained through visual cues is recorded as the visual cues predicted category.

[0133] As an embodiment, a method for obtaining a model to be trained includes:

[0134] The initial network model is obtained by training, and the initial network model includes a backbone network, a feature enhancement network, a target decoder and a text encoder; a feature aggregation network is generated based on the text encoder, and a feature aggregation network and a feature extraction network are added to the initial network model to obtain a model to be trained; the text encoder includes multiple feature interaction layers, and the feature aggregation network is obtained by inserting at least one feature aggregation layer between the multiple feature interaction layers included in the text encoder;

[0135] Among them, the training process of the initial network model includes:

[0136] Output the text prompt to the text encoder to obtain the text prompt features corresponding to each category;

[0137] The text prompt features and sample image features corresponding to each category are input into the feature enhancement network; the sample image features are obtained by extracting features from the sample images through the backbone network;

[0138] The sample image features and the text prompt features corresponding to each category are enhanced through the feature enhancement network to obtain the sample text enhancement features and the text prompt enhancement features corresponding to each category;

[0139] The similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category is determined through the target decoder, and the text prompt prediction category of the target object included in the sample image is determined according to the similarity; the initial network model is obtained by training according to the text prompt prediction category.

[0140] In this embodiment, the predicted category obtained through text hint is recorded as the text hint predicted category.

[0141] In this embodiment, the model to be trained can be obtained by further adding a feature aggregation network and a feature extraction network to the initial network model obtained by training the sample image and the multiple text prompt information corresponding to the sample image. For details, please refer to Figure 3 The box diagram on the left.

[0142] like Figure 3 As shown, the sample image is input into the backbone network for feature extraction to determine the sample image features of the sample image.

[0143] The text encoder is further used to extract features of multiple text prompt information corresponding to the sample image to obtain text prompt features corresponding to each preset type.

[0144] The sample image and the text prompt features corresponding to each preset type are input into the feature enhancement network, so as to calculate the similarity between the sample image features and the text prompt features corresponding to each category through the feature enhancement network, determine the weights of the sample image features and the text prompt features corresponding to each category according to the similarity between the sample image features and the text prompt features corresponding to each category, and perform weighted summation based on the weights of the sample image features and the text prompt features corresponding to each category to obtain the sample text enhancement feature; and for each category, obtain the text prompt enhancement feature corresponding to the category according to the weight of the text prompt feature corresponding to the category and the weight of the sample image features.

[0145] The sample text enhancement features and the text prompt enhancement features corresponding to each type are input into the target decoder, and the decoder determines the similarity between the sample text enhancement features and the text prompt enhancement features corresponding to each type, and determines the predicted category and predicted position of each target object.

[0146] The initial network model is obtained by training the predicted category and predicted position, and the feature aggregation network and feature extraction network are further added to obtain the model to be trained.

[0147] This concludes Figure 3 Description of the box on the left.

[0148] In this embodiment, in the process of training the initial network model, the text encoder has actually been trained. The subsequent feature aggregation network reuses the structure and parameters of the text encoder, so there is no need to start training from scratch. The trained resources of the text encoder can be used, which greatly improves the model convergence efficiency.

[0149] As an embodiment, the model to be trained also includes a text encoder, and the method further includes:

[0150] Outputting the text prompt to the text encoder to obtain the text prompt features corresponding to each category, and the process of outputting the text prompt to the text encoder and the process of extracting features of multiple visual prompt images corresponding to each category through the feature extraction network are performed alternately;

[0151] Input the text prompt features corresponding to each category and the sample image features output by the backbone network into the feature enhancement network, calculate the similarity between the sample image features and the text prompt features corresponding to each category through the feature enhancement network, determine the weights of the sample image features and the text prompt features corresponding to each category according to the similarity between the sample image features and the text prompt features corresponding to each category, and perform weighted summation according to the weights of the sample image features and the text prompt features corresponding to each category to obtain the sample text enhancement features; and for each category, obtain the text prompt enhancement features corresponding to the category according to the weight of the text prompt features corresponding to the category and the weight of the sample image features;

[0152] The sample text enhancement feature and the text prompt enhancement feature corresponding to each category are input into the target decoder, and the similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category is determined by the target decoder to obtain the text prompt prediction category of the target object in the sample image according to the similarity;

[0153] The model to be trained is trained based on the textual prompt prediction category and the visual prompt prediction category to obtain a trained object detection model.

[0154] In this embodiment, the training process using visual prompt images can also be alternately performed using text prompts. Figure 3 The description process of the left frame is the same and will not be repeated here.

[0155] Please refer to Figure 3 , Figure 3 Schematic diagram of the target detection model training framework provided in an embodiment of the present application.

[0156] like Figure 3 As shown, the figure on the right is a schematic diagram of training the target detection model to be trained.

[0157] In this embodiment, training through text prompts and training through visual prompts can be performed alternately, that is, during the training process, the input of one iteration is a pure text prompt, and the input of another iteration is a pure visual prompt, and this alternation is repeated back and forth, so that the trained network has the ability to be based on pure text prompts, pure visual prompts, and visual and text fusion prompts.

[0158] In this embodiment, the specific training process has been described above and will not be repeated here.

[0159] This concludes Figure 3 Description.

[0160] This concludes Figure 1 Description of the flow chart of the object detection model training method.

[0161] In the present application, when obtaining a sample image and multiple visual cue images corresponding to the sample image, for each category, feature extraction is performed on the multiple visual cue images corresponding to the category through a feature extraction network to obtain multiple visual cue features corresponding to the category, and feature aggregation is performed on the reference feature corresponding to the category and the multiple visual cue features corresponding to the category through a feature aggregation network to obtain the aggregated visual cue features corresponding to the category, and a to-be-trained model is trained according to the aggregated visual cue features corresponding to each category to obtain a target detection model, and the reference features corresponding to each category are configured in the feature aggregation network so that for each category, the multiple cue features corresponding to the category can be aggregated into the reference feature to obtain the aggregated visual cue features used to represent the category, thereby improving the efficiency of model training.

[0162] In addition, since the feature aggregation network in the present application reuses a pre-trained text encoder with generalization ability, it can make more full use of the trained resources of the text encoder during the training process, greatly improving the convergence rate of model training.

[0163] At the same time, in terms of training strategy, this application first performs pure text prompt pre-training to obtain the network to be trained and the trained text encoder, and then fine-tunes the network to be trained for pure text and pure visual prompts. Regarding the visual prompt feature aggregation process, this application proposes to add some new feature aggregation layers to the existing text encoder structure. By reusing the structure and network weights of the text encoder, the network convergence speed can be accelerated and the final effect of the network can be improved.

[0164] Please refer to Figure 4 , Figure 4 It is a schematic structural diagram of an electronic device proposed in an embodiment of the present application. At the hardware level, the electronic device includes a processor, an internal bus, a network interface, and a computer-readable storage medium, and of course may also include hardware required for other services. The processor reads the corresponding computer program from the computer-readable storage medium and then runs it to form a terminal interaction device at the logical level. Of course, in addition to the software implementation method, the present application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0165] Please refer to Figure 5 , Figure 5 : is a structural diagram of a target detection model training device proposed in an embodiment of the present application, and the model to be trained includes a feature extraction network and a feature aggregation network. Figure 5 As shown, the device may include an acquisition unit 501, an extraction unit 502, an aggregation unit 503, and a training unit 504. Specifically, the device includes:

[0166] An acquisition unit 501 is used to acquire a plurality of visual cue images corresponding to the sample image, wherein for each visual cue image, the visual cue image includes at least one category of reference objects;

[0167] An extraction unit 502 is used for extracting features from a plurality of visual cue images corresponding to each category through a feature extraction network to obtain a plurality of visual cue features corresponding to the category;

[0168] Aggregation unit 503, for each category, performing feature aggregation on a reference feature corresponding to the category and a plurality of visual cue features corresponding to the category through a feature aggregation network to obtain an aggregated visual cue feature corresponding to the category; wherein the feature aggregation network is configured with a reference feature corresponding to each category;

[0169] The training unit 504 is used to train the model to be trained according to the aggregated visual cue features corresponding to each category to obtain a trained target detection model; wherein the target detection model is used to detect the image to be tested to obtain the target category of the object in the image to be tested.

[0170] Optionally, the extraction unit 502 is specifically configured to:

[0171] For each visual cue image, extract features of the visual cue image through a feature extraction network to obtain the overall visual features corresponding to the visual cue image;

[0172] For each category, the sub-visual features of the reference object corresponding to the category are intercepted from the overall visual features corresponding to each visual image to obtain the visual prompt features corresponding to the category;

[0173] And / or, the aggregation unit 503 is specifically used for:

[0174] For each feature interaction layer, a first input feature of the feature interaction layer is obtained, and the first input feature includes reference features of all categories; for each category, similarity calculation is performed on the reference feature of the category and the reference feature of each category through the feature interaction layer, and the weight corresponding to the reference feature of each category is determined according to the similarity between the reference feature of the category and the reference feature of each category, and the reference feature of the category and the reference features of other categories are weighted summed according to the weight to obtain the interaction feature of the category;

[0175] For each feature aggregation layer, a second input feature of the feature aggregation layer is obtained, where the second input feature includes interaction features of all categories and multiple visual cue features corresponding to each category; for each category, similarity is calculated between the interaction features of the category and the multiple visual cue features corresponding to the category through the feature aggregation layer, and a weight corresponding to each visual cue feature corresponding to the reference feature of the category is determined according to the similarity between the reference feature of the category and the multiple visual cue features corresponding to the category, and a weighted sum is performed on the reference feature of the category and the multiple visual cue features corresponding to the category according to the weight, so as to obtain a fusion feature of the category;

[0176] And / or, the feature aggregation network includes multiple network layers, the first to M network layers of the multiple network layers are feature interaction layers, and starting from the M+1th network layer, they are alternately feature aggregation layers and feature interaction layers;

[0177] And / or, the model to be trained further includes a backbone network, a feature enhancement network and a target decoder, and the training unit 504 is specifically used for:

[0178] Inputting a sample image into the backbone network to obtain sample image features, wherein the sample image includes a target object of at least one category;

[0179] The sample image features and the aggregated visual cue features corresponding to each category are input into the feature enhancement network, and the similarity between the sample image features and the aggregated visual cue features corresponding to each category is calculated through the feature enhancement network, and the weights of the sample image features and the aggregated visual cue features corresponding to each category are determined according to the similarity between the sample image features and the aggregated visual cue features corresponding to each category, and the sample image features and the aggregated visual cue features corresponding to each category are weighted summed to obtain the sample visual enhancement features; and for each category, the aggregated visual enhancement features corresponding to the category are obtained according to the weight of the aggregated visual cue features corresponding to the category and the weight of the sample image features;

[0180] The sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category are input into the target decoder, and the similarity between the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category is determined by the target decoder, so as to obtain the visual cue prediction category of the target object in the sample image according to the similarity;

[0181] The model to be trained is trained based on the visual cue prediction category of the target object to obtain a trained target detection model;

[0182] And / or, the model to be trained further includes a text encoder, and the training unit 504 is further used for:

[0183] Outputting the text prompt to the text encoder to obtain the text prompt features corresponding to each category, and the process of outputting the text prompt to the text encoder and the process of extracting features of multiple visual prompt images corresponding to each category through the feature extraction network are performed alternately;

[0184] Input the text prompt features corresponding to each category and the sample image features output by the backbone network into the feature enhancement network, calculate the similarity between the sample image features and the text prompt features corresponding to each category through the feature enhancement network, determine the weights of the sample image features and the text prompt features corresponding to each category according to the similarity between the sample image features and the text prompt features corresponding to each category, and perform weighted summation according to the weights of the sample image features and the text prompt features corresponding to each category to obtain the sample text enhancement features; and for each category, obtain the text prompt enhancement features corresponding to the category according to the weight of the text prompt features corresponding to the category and the weight of the sample image features;

[0185] The sample text enhancement feature and the text prompt enhancement feature corresponding to each category are input into the target decoder, and the similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category is determined by the target decoder to obtain the text prompt prediction category of the target object in the sample image according to the similarity;

[0186] The model to be trained is trained based on the textual prompt prediction category and the visual prompt prediction category to obtain a trained object detection model;

[0187] And / or, the model to be trained is obtained by:

[0188] The initial network model is obtained by training, and the initial network model includes a backbone network, a feature enhancement network, a target decoder and a text encoder; a feature aggregation network is generated based on the text encoder, and a feature aggregation network and a feature extraction network are added to the initial network model to obtain a model to be trained; the text encoder includes multiple feature interaction layers, and the feature aggregation network is obtained by inserting at least one feature aggregation layer between the multiple feature interaction layers included in the text encoder;

[0189] Among them, the training process of the initial network model includes:

[0190] Output the text prompt to the text encoder to obtain the text prompt features corresponding to each category;

[0191] The text prompt features and sample image features corresponding to each category are input into the feature enhancement network; the sample image features are obtained by extracting features from the sample images through the backbone network;

[0192] The sample image features and the text prompt features corresponding to each category are enhanced through the feature enhancement network to obtain the sample text enhancement features and the text prompt enhancement features corresponding to each category;

[0193] The similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category is determined through the target decoder, and the text prompt prediction category of the target object included in the sample image is determined according to the similarity; the initial network model is obtained by training according to the text prompt prediction category.

[0194] So far, completed Figure 5 Description of the object detection model training setup.

[0195] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed, the method disclosed in the above example of the present application can be implemented.

[0196] Exemplarily, the computer-readable storage medium may be any electronic, magnetic, optical or other physical storage device that may contain or store information, such as executable instructions, data, etc. For example, the computer-readable storage medium may be: RAM (Radom Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disk, DVD, etc.), or similar storage medium, or a combination thereof.

[0197] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of protection of the present application.

Claims

1. A target detection model training method, characterized in that: The model to be trained includes a feature extraction network and a feature aggregation network. The method includes: Acquire a plurality of visual cue images corresponding to the sample image, wherein for each visual cue image, the visual cue image includes at least one category of reference objects; For each category, a feature extraction network is used to extract features from multiple visual cue images corresponding to the category to obtain multiple visual cue features corresponding to the category; For each category, a feature aggregation network is used to perform feature aggregation on a reference feature corresponding to the category and a plurality of visual cue features corresponding to the category, so as to obtain an aggregated visual cue feature corresponding to the category; wherein the feature aggregation network is configured with a reference feature corresponding to each category; The model to be trained is trained according to the aggregated visual cue features corresponding to each category to obtain a trained target detection model; wherein the target detection model is used to detect the image to be tested to obtain the target category of the object in the image to be tested.

2. The method according to claim 1, characterized in that The feature extraction network is used to extract features from multiple visual cue images corresponding to the category to obtain multiple visual cue features corresponding to the category, including: For each visual cue image, extract features of the visual cue image through the feature extraction network to obtain an overall visual feature corresponding to the visual cue image; For each category, the sub-visual features of the reference object corresponding to the category are extracted from the overall visual features corresponding to each visual image to obtain the visual prompt features corresponding to the category.

3. The method according to claim 1, characterized in that The feature aggregation network includes at least one feature interaction layer and at least one feature aggregation layer; the feature aggregation network performs feature aggregation on the reference feature corresponding to the category and the multiple visual cue features corresponding to the category to obtain the aggregated visual cue features corresponding to the category, including: For each feature interaction layer, a first input feature of the feature interaction layer is obtained, wherein the first input feature includes reference features of all categories; for each category, similarity calculation is performed on the reference feature of the category and the reference feature of each category through the feature interaction layer, a weight corresponding to the reference feature of each category is determined according to the similarity between the reference feature of the category and the reference feature of each category, and a weighted sum is performed on the reference feature of the category and the reference features of other categories according to the weight to obtain the interaction feature of the category; For each feature aggregation layer, a second input feature of the feature aggregation layer is obtained, wherein the second input feature includes interaction features of all categories and multiple visual cue features corresponding to each category; for each category, a similarity calculation is performed on the interaction features of the category and the multiple visual cue features corresponding to the category through the feature aggregation layer, and a weight corresponding to each visual cue feature corresponding to the reference feature of the category is determined according to the similarity between the reference feature of the category and the multiple visual cue features corresponding to the category, and a weighted sum is performed on the reference feature of the category and the multiple visual cue features corresponding to the category according to the weight to obtain a fusion feature of the category.

4. The method according to claim 3, characterized in that The feature aggregation network includes multiple network layers, wherein the first to M network layers of the multiple network layers are feature interaction layers, and starting from the M+1th network layer, they are alternately feature aggregation layers and feature interaction layers.

5. The method according to claim 1, characterized in that The model to be trained also includes a backbone network, a feature enhancement network and a target decoder. The model to be trained is trained according to the aggregated visual cue features corresponding to each category to obtain a trained target detection model, including: Inputting the sample image into the backbone network to obtain sample image features, wherein the sample image includes a target object of at least one category; The sample image features and the aggregated visual cue features corresponding to each category are input into a feature enhancement network, and similarity calculation is performed on the sample image features and the aggregated visual cue features corresponding to each category through the feature enhancement network, and the weights of the sample image features and the aggregated visual cue features corresponding to each category are determined according to the similarity between the sample image features and the aggregated visual cue features corresponding to each category, and a weighted sum is performed according to the weights of the sample image features and the aggregated visual cue features corresponding to each category to obtain the sample visual enhancement feature; and for each category, the aggregated visual enhancement feature corresponding to the category is obtained according to the weight of the aggregated visual cue features corresponding to the category and the weight of the sample image features; Inputting the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category into a target decoder, and determining the similarity between the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category by the target decoder, so as to obtain a visual cue prediction category of the target object in the sample image according to the similarity; The model to be trained is trained based on the visual cue prediction category of the target object to obtain a trained target detection model.

6. The method according to claim 1, characterized in that The model to be trained also includes a text encoder, and the method further includes: Outputting the text prompt to a text encoder to obtain text prompt features corresponding to each category, wherein the process of outputting the text prompt to the text encoder and the process of extracting features of multiple visual prompt images corresponding to each category through a feature extraction network are performed alternately; Input the text prompt features corresponding to each category and the sample image features output by the backbone network into the feature enhancement network, calculate the similarity between the sample image features and the text prompt features corresponding to each category through the feature enhancement network, determine the weights of the sample image features and the text prompt features corresponding to each category according to the similarity between the sample image features and the text prompt features corresponding to each category, and perform weighted summation according to the weights of the sample image features and the text prompt features corresponding to each category to obtain the sample text enhancement features; and for each category, obtain the text prompt enhancement features corresponding to the category according to the weight of the text prompt features corresponding to the category and the weight of the sample image features; The sample text enhancement feature and the text prompt enhancement feature corresponding to each category are input into a target decoder, and the target decoder determines the similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category, so as to obtain the text prompt prediction category of the target object in the sample image according to the similarity; The model to be trained is trained based on the text prompt prediction category and the visual prompt prediction category to obtain a trained object detection model.

7. The method according to claim 1, characterized in that The model to be trained is obtained by the following method: The initial network model is obtained by training, and the initial network model includes a backbone network, a feature enhancement network, a target decoder and a text encoder; the feature aggregation network is generated based on the text encoder, and the feature aggregation network and the feature extraction network are added to the initial network model to obtain the model to be trained; the text encoder includes multiple feature interaction layers, and the feature aggregation network is obtained by inserting at least one feature aggregation layer between the multiple feature interaction layers included in the text encoder; The training process of the initial network model includes: Output the text prompt to the text encoder to obtain the text prompt features corresponding to each category; Inputting the text prompt features and sample image features corresponding to each category into the feature enhancement network; the sample image features are obtained by extracting features of the sample images through the backbone network; The sample image features and the text prompt features corresponding to each category are enhanced through the feature enhancement network to obtain the sample text enhancement features and the text prompt enhancement features corresponding to each category; The similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category is determined by the target decoder, and the text prompt prediction category of the target object included in the sample image is determined according to the similarity; the initial network model is obtained by training according to the text prompt prediction category.

8. A target detection model training device, characterized in that: The model to be trained includes a feature extraction network and a feature aggregation network, and the device includes: An acquisition unit, configured to acquire a plurality of visual cue images corresponding to the sample image, wherein for each visual cue image, the visual cue image includes at least one category of reference objects; An extraction unit is used for extracting features from a plurality of visual cue images corresponding to each category through a feature extraction network to obtain a plurality of visual cue features corresponding to the category; an aggregation unit, for performing feature aggregation on a reference feature corresponding to the category and a plurality of visual cue features corresponding to the category through a feature aggregation network for each category, so as to obtain an aggregated visual cue feature corresponding to the category; wherein the feature aggregation network is configured with a reference feature corresponding to each category; A training unit is used to train the model to be trained according to the aggregated visual cue features corresponding to each category to obtain a trained target detection model; wherein the target detection model is used to detect the image to be tested to obtain the target category of the object in the image to be tested.

9. The device according to claim 8, characterized in that The extraction unit is specifically used for: For each visual cue image, extract features of the visual cue image through the feature extraction network to obtain an overall visual feature corresponding to the visual cue image; For each category, the sub-visual features of the reference object corresponding to the category are intercepted from the overall visual features corresponding to each visual image to obtain the visual prompt features corresponding to the category; And / or, the polymeric unit is specifically used for: For each feature interaction layer, obtaining a first input feature of the feature interaction layer, wherein the first input feature includes reference features of all categories; For each category, the similarity between the reference feature of the category and the reference feature of each category is calculated through the feature interaction layer, the weight corresponding to the reference feature of each category is determined according to the similarity between the reference feature of the category and the reference feature of each category, and the reference feature of the category and the reference features of other categories are weighted summed according to the weight to obtain the interaction feature of the category; For each feature aggregation layer, a second input feature of the feature aggregation layer is obtained, wherein the second input feature includes interaction features of all categories and a plurality of visual cue features corresponding to each category; for each category, similarity calculation is performed on the interaction features of the category and the plurality of visual cue features corresponding to the category through the feature aggregation layer, and a weight corresponding to each visual cue feature corresponding to the reference feature of the category is determined according to the similarity between the reference feature of the category and the plurality of visual cue features corresponding to the category, and a weighted sum is performed on the reference feature of the category and the plurality of visual cue features corresponding to the category according to the weight, so as to obtain a fusion feature of the category; And / or, the feature aggregation network includes multiple network layers, the first to M network layers of the multiple network layers are feature interaction layers, and starting from the M+1th network layer, they are alternately feature aggregation layers and feature interaction layers; And / or, the model to be trained further includes a backbone network, a feature enhancement network and a target decoder, and the training unit is specifically used for: Inputting the sample image into the backbone network to obtain sample image features, wherein the sample image includes a target object of at least one category; The sample image features and the aggregated visual cue features corresponding to each category are input into a feature enhancement network, and similarity between the sample image features and the aggregated visual cue features corresponding to each category is calculated through the feature enhancement network, and the weights of the sample image features and the aggregated visual cue features corresponding to each category are determined according to the similarity between the sample image features and the aggregated visual cue features corresponding to each category, and a weighted sum is performed according to the weights of the sample image features and the aggregated visual cue features corresponding to each category to obtain the sample visual enhancement feature; And for each category, according to the weight of the aggregated visual cue feature corresponding to the category and the weight of the sample image feature, the aggregated visual enhancement feature corresponding to the category is obtained; Inputting the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category into a target decoder, and determining the similarity between the sample visual enhancement feature and the aggregated visual enhancement feature corresponding to each category by the target decoder, so as to obtain a visual cue prediction category of the target object in the sample image according to the similarity; Training the model to be trained based on the visual cue prediction category of the target object to obtain a trained target detection model; And / or, the model to be trained further includes a text encoder, and the training unit is further used for: Outputting the text prompt to a text encoder to obtain text prompt features corresponding to each category, wherein the process of outputting the text prompt to the text encoder and the process of extracting features of multiple visual prompt images corresponding to each category through a feature extraction network are performed alternately; Input the text prompt features corresponding to each category and the sample image features output by the backbone network into the feature enhancement network, calculate the similarity between the sample image features and the text prompt features corresponding to each category through the feature enhancement network, determine the weights of the sample image features and the text prompt features corresponding to each category according to the similarity between the sample image features and the text prompt features corresponding to each category, and perform weighted summation according to the weights of the sample image features and the text prompt features corresponding to each category to obtain the sample text enhancement features; And for each category, according to the weight of the text prompt feature corresponding to the category and the weight of the sample image feature, the text prompt enhancement feature corresponding to the category is obtained; The sample text enhancement feature and the text prompt enhancement feature corresponding to each category are input into a target decoder, and the target decoder determines the similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category, so as to obtain the text prompt prediction category of the target object in the sample image according to the similarity; Training the model to be trained based on the text prompt prediction category and the visual prompt prediction category to obtain a trained object detection model; And / or, the model to be trained is obtained by the following method: The initial network model is obtained by training, and the initial network model includes a backbone network, a feature enhancement network, a target decoder and a text encoder; the feature aggregation network is generated based on the text encoder, and the feature aggregation network and the feature extraction network are added to the initial network model to obtain the model to be trained; the text encoder includes multiple feature interaction layers, and the feature aggregation network is obtained by inserting at least one feature aggregation layer between the multiple feature interaction layers included in the text encoder; The training process of the initial network model includes: Output the text prompt to the text encoder to obtain the text prompt features corresponding to each category; Inputting the text prompt features and sample image features corresponding to each category into the feature enhancement network; the sample image features are obtained by extracting features of the sample images through the backbone network; The sample image features and the text prompt features corresponding to each category are enhanced through the feature enhancement network to obtain the sample text enhancement features and the text prompt enhancement features corresponding to each category; The similarity between the sample text enhancement feature and the text prompt enhancement feature corresponding to each category is determined by the target decoder, and the text prompt prediction category of the target object included in the sample image is determined according to the similarity; the initial network model is obtained by training according to the text prompt prediction category.

10. An electronic device, characterized in that: The electronic device includes: A processor and a computer-readable storage medium, wherein computer program instructions are stored in the computer-readable storage medium, and when the computer program instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image classification model training method and device

    CN116259060A

  • Prompt learning method, image classification method and related device

    CN118155214A