Open vocabulary target detection method and system based on multi-target classification, medium and program product

By building an open vocabulary object detection model for multi-object classification, using context information and multi-object classification module, the problem of insufficient detection performance of new categories in the existing technology is solved, and more accurate detection and recognition of new categories of objects is achieved.

CN119919634AActive Publication Date: 2025-05-02XIAMEN UNIV

Patent Information

Application Number
CN202411980224.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The existing open vocabulary object detection methods are prone to introduce noise when acquiring new category knowledge and lack of attention to context, resulting in the model's overfitting of base categories and limitations on the detection performance of new categories.

Method used

Using an open vocabulary object detection method based on multi-object classification, a detection model including a backbone feature extraction network, a region suggestion network, a candidate box expansion module, a distillation module and a multi-object classification module are used to detect new categories of objects that do not appear in the training set, and use new categories of information to classify them through the multi-object classification module.

Benefits of technology

It realizes more accurate open vocabulary object detection, which can effectively detect new categories of objects that have not appeared in the training set, and improves the model's understanding and recognition performance of new categories of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919634A_ABST
    Figure CN119919634A_ABST
Patent Text Reader

Abstract

The invention discloses an open vocabulary target detection method and system based on multi-target classification, a medium and a program product, and belongs to the technical field of computer vision. Comprising the following steps: performing feature extraction on an image by using a trunk feature extraction network to obtain a feature map; the region suggestion network generates a group of candidate boxes in the feature map, and applies a candidate box expansion module to all the candidate boxes to obtain an expansion box; learning knowledge from a CLIP image encoder by using a distillation module to obtain distillation loss; inputting the features extracted by the candidate box and the extension box and the text features into a multi-target classification module to obtain classification loss; training the model in combination with distillation loss and classification loss; and finally, inputting a to-be-detected picture into the trained model, and generating a predicted target area, a corresponding category name and a corresponding confidence coefficient, thereby realizing more accurate open vocabulary target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to an open vocabulary target detection method, system, medium and program product based on multi-target classification. Background Art

[0002] In recent years, with the improvement of computer performance and the development of big data, visual information data has increased rapidly, including multimedia data such as static images, dynamic images, video files, audio files, etc., which are spread at a very fast speed on various social media. As one of the most basic problems in the field of computer vision, target detection is widely used in many fields such as target tracking, behavior understanding, human-computer interaction, face recognition, etc., and has attracted extensive attention and research from many scholars in the early 20th century. Humans mainly receive external information through vision, so application technology based on visual information will be a forward-looking research point of artificial intelligence. Among them, face recognition, video surveillance, target detection, Internet image content review, biometric recognition and other technologies have become research hotspots today. These technologies are also widely used in medical care, elderly care, transportation, urban operation, security and other fields, such as: medical image diagnosis, posture estimation, station security, automatic driving, vehicle speed detection, video surveillance behavior analysis, etc.

[0003] The goal of traditional object detection is to identify and label all objects of known categories in an image. In order to detect a wider range of categories in real-world scenarios, the model must add more categories and their corresponding fine-grained data to the training set, which is costly. In addition, linearly adding new categories will exacerbate the long-tail distribution problem of categories.

[0004] Open vocabulary object detection has been proposed as a solution in recent years. Open vocabulary object detection aims to detect new categories that have never been seen during training. In contrast to zero-shot learning that strictly limits the information about new categories, open vocabulary object detection allows access to semantic information related to new categories, which comes from image-text pairs with image-level labels or pre-trained visual-language models, thus facilitating the learning of a cross-modal visual-language representation space that links image regions with word descriptions.

[0005] Existing open-vocabulary object detection mainly transfers the knowledge of new categories from pre-trained vision-language models to object detectors through knowledge distillation to gain the ability to recognize new categories.

[0006] OVR-CNN (Open Vocabulary Region-Convolutional Neural Network) first used contrastive learning to expand the vocabulary on a large number of image-text pairs, while fine-tuning on detection data with limited vocabulary. The models after this were modified to improve performance in the following aspects: knowledge distillation, region text pre-training, prompt modeling, region text alignment, and transfer-based learning. For example, DK-DETR (DistillingKnowledge DEtection TRansformer) models object relationships by aligning feature relationship matrices between student and teacher models. OADP (Object Aware Distillation Pyramid) enhances the student model by aligning image patch embeddings. HierKD (HierarchicalKnowledge Distillation) extracts multi-scale global features of images to align descriptions. GKC (GatedContextual Kernel) avoids overfitting by introducing additional category synonyms. ZeroSeg (Zero-shot Segmentation) extracts visual concepts by processing images into multi-scale views.

[0007] However, these methods introduce noise when acquiring knowledge of new categories and pay insufficient attention to context during training, which leads to overfitting of the model to the base categories and limited performance in detecting new categories. The context in these images contains the intrinsic structure of multiple semantic concepts, which may include objects of new categories, which are the targets we need to detect.

[0008] The existing background cue learning method in open vocabulary object detection in CN202410651398 uses the RPN network to directly generate candidate boxes that may contain new types of objects. The RPN network itself is only trained on the base class, so the candidate boxes generated in this way have a lot of room for improvement. Summary of the invention

[0009] The object of the present invention is to provide an open vocabulary target detection method, system, medium and program product based on multi-target classification, which can use context information to detect new categories of objects that have not appeared in the training set, thereby achieving more accurate open vocabulary target detection.

[0010] In order to achieve the above object, the solution of the present invention is:

[0011] An open vocabulary object detection method based on multi-object classification uses context information to detect new categories of objects that have not appeared in the training set, including the following steps:

[0012] Step S1, constructing an open vocabulary target detection model based on multi-target classification; the model includes a backbone feature extraction network ResNet50, a region proposal network RPN, a candidate box expansion module CBM, a distillation module KD and a multi-target classification module MSC;

[0013] Step S2, training the model, specifically includes the following steps:

[0014] Step S21, providing a training image sample set;

[0015] Step S22, randomly select an image from the training image sample set, use the backbone feature extraction network ResNet50 to extract features from the image to obtain a feature map, and input the feature map into the region proposal network RPN;

[0016] Step S23, the region proposal network RPN generates a set of candidate boxes in the feature map, and then applies the candidate box expansion module CBM to all the candidate boxes to obtain an extended box located near the candidate box for adding context information;

[0017] Step S24, using the distillation module KD to learn knowledge from the CLIP image encoder, and using the backbone feature extraction network ResNet50 of the model to extract the features F of the extended box and the candidate box E ; Use CLIP image encoder to process the candidate frame and the expanded frame to obtain its feature F C , and then apply InfoNCE Loss to calculate the two corresponding features F E and F C Distillation loss L KD ;

[0018] Step S25, using the CLIP text encoder to process the basic category names of the training image sample set to obtain the text feature F T , the feature F E and text features F T Input the candidate box into the multi-target classification module MSC to classify the candidate box and get the total classification loss L MSC ;

[0019] Step S26, combining the distillation loss L KD and the total classification loss L MSC The model is trained to obtain a trained open vocabulary target detection model for multi-target classification;

[0020] Step S3, input the image to be detected into the trained model, use the trained model to generate the predicted target area and the category name corresponding to the target area on the image, and give the confidence corresponding to this set of predictions. This confidence comes from the classification probability score of the target area by the multi-target classification module MSC.

[0021] Further, in step S22, an image is randomly selected from the training image sample set as the image I∈R used in training H×W×C , where H, W, and C represent the image height, image width, and image dimension respectively. The backbone feature extraction network ResNet50 is used to extract features from the image I to obtain a feature map;

[0022] Step S23, generate a set of candidate boxes P = {p0, p1, ..., p n}, where p i represents the i-th candidate box, which contains four values, which are used to record the relative position of the i-th candidate box in the picture, and n is the number of candidate boxes; and in order to increase the context information, the candidate box expansion module CBM is applied to all candidate boxes P to obtain the expanded box B = {B0, B1, ..., B n}, expand box B i Adopt and p i The same way of expression, B i represents the i-th extended box, which also contains four values. The values ​​are used to record the relative position of the i-th extended box in the picture. n is the number of extended boxes. At the same time, the regression module is applied to optimize the candidate box P and obtain the regression loss L reg .

[0023] Furthermore, in step S23, firstly, the candidate box p is taken i The eight nearby regions are used to expand each candidate box and obtain P ie ={p i ,r i1 ,…,r i8}, where p i ∈P, where p i represents the original i-th candidate box selected previously, and r i1 ,…,r i8 Represents the candidate box p i Eight nearby areas;

[0024] Then filter P ie , to remove the area beyond the image boundary and obtain P ir ={p i ,r i1 ,…,r ik}, where k represents the number of remaining regions, and k≤8;

[0025] In order to find the existence of new types of objects, according to p i The similarity of P ir Sort by small to large, and shield the area with low objectivity score to get j≤k;

[0026] Finally, from Choose three with p i The boxes with the lowest similarity and compare them with p i Combined to form an expansion frame. Further, the process of the last step in step S23 is expressed as:

[0027] B i ={p i , p im c , p i(m+1) c , p i(m+2) c}

[0028] Among them, m represents the selection of a similarity region that is most likely to contain a new class of objects, and p im c , p i(m+1) c , p i(m+2) c The three surrounding areas are finally selected. The final expansion box B is obtained based on these three areas and the original area. i ;

[0029] Expressed in formula, it is:

[0030] p i c =Mask(Sort(Φ(p i ·r i )));

[0031] Where Φ(·) represents cosine similarity, Sort represents the process of calculating similarity and sorting, and Mask represents the selection of surrounding areas based on the similarity sorting results. i ∈P ir And r i ≠p i .

[0032] Further, in step S24, the backbone feature extraction network ResNet50 of the model extracts the features of the extended box and the features of the candidate box F E : F E ={F E 1,F E 2,…,F En ,F E (n+1) ,…,F E 2n}, where the first n features are the features of the expanded box, and the last n features are the features of the candidate box;

[0033] Then use the CLIP image encoder to process the candidate box and the expanded box to obtain its feature F C : F C ={F C 1,F C 2,…,F C n ,F C (n+1) ,…,F C 2n}, where the first n features are the features of the expanded box, and the last n features are the features of the candidate box;

[0034] Then apply InfoNCEloss to calculate the distillation loss L between the two corresponding features. KD , by optimizing the distillation loss-L KD , so that the feature F generated by this model E The feature F generated by the CLIP image encoder C This is close to achieving the learning of knowledge in the CLIP image encoder.

[0035] Further, in step S25, the category name is processed using the CLIP text encoder to obtain the text feature F T ={F T 1,F T 2,…,F T q}, where the meaning of q changes with the state of the model. When the model is training, since only the basic category information of the training image sample set can be used for training, q represents the number of basic categories in the training image sample set. In the test process, since the model detects objects of both basic categories and new categories at the same time, q represents the number of all categories in the test image sample set. Then, the multi-target classification module MSC is used to classify the feature F E and text features F T Classification is performed to obtain the classification loss L MSC ;

[0036] First, the candidate frames are classified: the feature F extracted by the backbone feature extraction network ResNet50 of the present invention is E ={F E 1,F E 2,…,F E n+1 ,F En+2 ,…,F E 2n}Into the multi-target classification module MSC, for each pair of extended boxes B and their corresponding candidate boxes p, the real category G = {g1, g2, ..., g n}, g1, g2, …, g n Corresponding to the basic target categories in the training image sample set;

[0037] Then, by calculating F E i and F T gi , the cosine similarity of (gi∈G) generates the classification score, where F E i and F T gi F E and F T For each candidate box, calculate its classification loss L according to a single category. cls , for each expanded box, the classification loss is calculated, which is divided into L B and L N Two parts, among which L B Represents the loss of the basic category object, L N is the loss of potential new class objects, defined as:

[0038]

[0039] Where Φ(·) represents cosine similarity, Softmax represents normalization of the value, and log represents taking the logarithm;

[0040] Select the top k categories with the highest classification scores G′={g1′,g2′,…,g k ′} as the aligned category, the loss is defined as:

[0041]

[0042] Where Φ(·) represents cosine similarity, softmax represents normalization of the value, and log represents taking the logarithm;

[0043] Total classification loss L MSC Defined as:

[0044] L MSC =L cls +L B +λ4L N

[0045] Among them, λ4 represents the weight of the loss of potential new category objects.

[0046] The present invention also provides an open vocabulary target detection method system based on multi-target classification, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the method.

[0047] The present invention also provides a computer program product, comprising a computer program / instruction, which implements the method when executed by a processor.

[0048] The present invention also provides a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method described above is implemented.

[0049] After adopting the above technical solution, the present invention mines potential new class targets from contextual information, and designs detailed steps to screen candidate boxes, thereby obtaining richer new class target information and being able to detect new categories of objects that have not appeared in the training set; at the same time, a multi-target classification method is adopted to utilize new class information, which also improves the model's understanding of new class information. The present invention takes the OV-COCO and OV-LVIS datasets, trains the model according to the above steps, and uses the test set for testing. The present invention uses a validation set to evaluate the performance on the OV-COCO and OV-LVIS datasets. Evaluation indicators Represents the average precision (mAP) of the new category under the IoU threshold of 0.5; the evaluation indicator AP r , represents the mean average precision (mAP) of rare categories (new categories). The recognition performance of this method for new categories on the OV-COCO dataset reaches 38.5% (ResNet50-FPN), and the recognition performance of new categories on the OV-LVIS dataset reaches 24.1% (detection task) and 23.4% (segmentation task), which is higher than other methods, proving that the present invention has a better effect in identifying new categories of targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a schematic diagram of the network structure of the present invention. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.

[0052] like Figure 1 The present invention provides an open vocabulary target detection method based on multi-target classification, and the detection method of the present invention comprises the following steps:

[0053] Step S1, constructing an open vocabulary object detection model based on multi-target classification (referred to as CANet), which includes a backbone feature extraction network (ResNet50), a region proposal network (RegionProposalNetwork, RPN), a candidate box expansion module (Context Broaden Module, CBM), a distillation module (Knowledge Distillation, KD) and a multi-target classification module (Multi-Semantic Classifier, MSC);

[0054] The backbone feature extraction network ResNet50 is used to extract features from any image in the training image sample set to obtain a feature map; the region proposal network RPN is used to generate a set of candidate boxes for the feature map extracted by the backbone feature extraction network ResNet50; the candidate box expansion module CBM is used to expand the candidate box, and can also rank the expanded part according to the similarity between the candidate box and the expanded part, and retain the best candidate box; the distillation module KD is used to close the image features generated by the backbone feature extraction network ResNet50 and the CLIP (Contrastive Language-Image Pre-Training) image encoder of the model; the multi-target classification module MSC is used to classify the candidate box and the expanded box at the same time;

[0055] Step S2, training the model: randomly select an image from the training image sample set, use the backbone feature extraction network ResNet50 to extract features from the image to obtain a feature map, and generate a set of candidate boxes from the feature map through the region proposal network RPN; the candidate box expansion module CBM expands all candidate boxes to obtain expanded boxes located near the candidate boxes for adding context information; the distillation module KD is used to learn knowledge from the CLIP image encoder to extract the features F of the expanded boxes and the candidate boxes E , use the CLIP image encoder to process the candidate box and the expanded box to obtain its feature F C , and then apply InfoNCE Loss to calculate the two corresponding features F E and F C Distillation loss L KD ; Use CLIP text encoder to process the basic category names of the training image sample set to obtain text features F T , the feature F E and text features F T Input into the multi-target classification module MSC for classification, and obtain the classification loss L MSC ; Combined distillation loss L KD and classification loss LMSC The model is trained to obtain a trained open vocabulary target detection model for multi-target classification;

[0056] Step S3, input the image to be detected into the trained model, use the trained model to generate the predicted target area and the category name corresponding to the target area on the image, and give the confidence corresponding to this set of predictions. This confidence comes from the classification probability score of the target area by the multi-target classification module MSC.

[0057] In step S2, the model training method specifically includes the following steps:

[0058] Step S21, given a data set with image-level labels, the set is divided into a training image sample set and a test image sample set;

[0059] Step S22: randomly select an image from the training image sample set as the image I∈R used in training. H ×W×C , where H, W, and C represent the image height, image width, and image dimension respectively. The backbone feature extraction network ResNet50 is used to extract features from the image I to obtain a feature map;

[0060] Step S23, generate a set of candidate boxes P = {p0, p1, ..., p n}, where p i represents the i-th candidate box, which contains four values, which are used to record the relative position of the i-th candidate box in the image, and n is the number of candidate boxes. In addition, in order to increase the context information, the candidate box expansion module CBM is applied to all candidate boxes P to obtain the expanded box B = {B0, B1, ..., B n}, expand box B i Adopt and p i The same way of expression, B i represents the i-th extended box, and also contains four values, which are used to record the relative position of the i-th extended box in the picture, and n is the number of extended boxes. At the same time, the present invention applies a regression module to optimize the candidate box P and obtain the regression loss L reg , that is, by optimizing the regression loss, the candidate box P extracted by the model can be closer to the actual location of the object.

[0061] Specifically, first take the candidate box p i The eight nearby regions are used to expand each candidate box and obtain P ie ={p i ,r i1 ,…,r i8}, where p i ∈P, where pi represents the original i-th candidate box selected previously, and r i1 ,…,r i8 Represents the candidate box p i Eight nearby areas;

[0062] Then filter P ie , to remove the area beyond the image boundary and obtain P ir ={p i ,r i1 ,…,r ik}, where k represents the number of remaining regions, and k≤8;

[0063] Furthermore, the present invention aims to find the existence of new classes of objects. i The similarity of P ir Sort by small to large, and shield the area with low objectivity score to get j≤k;

[0064] Finally, the present invention Choose three with p i The boxes with the lowest similarity and compare them with p i Combine to form an expansion box.

[0065] This process can be expressed as:

[0066] B i ={p i , p im c , p i(m+1) c , p i(m+2) c}

[0067] Among them, m represents the selection of a similarity region that is most likely to contain a new class of objects, and p im c , p i(m+1) c , p i(m+2) c The three surrounding areas are finally selected. The final expansion box B is obtained based on these three areas and the original area. i .

[0068] It can be expressed in formula as:

[0069] p i c =Mask(Sort(Φ(p i ·r i )));

[0070] Where Φ(·) represents cosine similarity, Sort represents the process of calculating similarity and sorting, and Mask represents the selection of surrounding areas based on the similarity sorting results. i ∈P ir And r i ≠p i .

[0071] Step S24, the present invention applies the distillation module KD to learn knowledge from the CLIP image encoder:

[0072] The backbone feature extraction network ResNet50 of the model will extract the features of the extended box and the features of the candidate box F E : F E ={F E 1,F E 2,…,F E n ,F E (n+1) ,…,F E 2n}, where the first n features are the features of the extended frame, and the last n features are the features of the candidate frame. At the same time, the present invention also uses the CLIP image encoder to process the candidate frame and the extended frame to obtain their features F C : F C= {F C 1,F C 2,…,F C n ,F C (n+1) ,…,F C 2n}, where the first n features are the features of the extended box, and the last n features are the features of the candidate box. Here, the present invention uses InfoNCEloss (information contrast loss) to calculate the distillation loss L between the two corresponding features. KD , by optimizing this distillation loss L KD , so that the features generated by the present invention are close to the features generated by the CLIP image encoder, thereby realizing the learning of the knowledge in the CLIP image encoder.

[0073] Step S25, use CLIP text encoder to process the category name to obtain text feature F T ={F T 1,F T 2,…,F T q}, where the meaning of q changes with the state of the model. When the model is training, since only the basic category information of the training image sample set can be used for training, q represents the number of basic categories in the training image sample set. In the test process, since the model detects objects of both basic categories and new categories at the same time, q represents the number of all categories in the test image sample set. Then, the multi-target classification module MSC is used to classify the feature F E and text features F T Classification is performed to obtain the classification loss L MSC ;

[0074] First, classify the candidate boxes:

[0075] The feature F extracted by the backbone feature extraction network ResNet50 of the present invention E ={F E 1,F E 2,…,F E n+1 ,F E n+2 ,…,F E 2n}Into the multi-target classification module MSC, for each pair of extended boxes B and their corresponding candidate boxes p, the real category G = {g1, g2, ..., g n}, g1, g2, …, g n Corresponding to the basic target category in the training image sample set.

[0076] Then, by calculating F E i and F T gi , the cosine similarity of (gi∈G) generates the classification score, where F E i and F T gi F E and F T For each candidate box, calculate its classification loss L according to a single category. cls , for each expanded box, the classification loss is calculated, which can be divided into L B and L N Two parts, among which L B Represents the loss of the basic category object, L N is the loss of potential new class objects, defined as:

[0077]

[0078] Where Φ(·) represents cosine similarity, Softmax represents normalization of the values, and log represents taking the logarithm.

[0079] Select the top k categories with the highest classification scores G′={g1′,g2′,…,g k ′} as the aligned category, the loss is defined as:

[0080]

[0081] Where Φ(·) represents cosine similarity, softmax represents normalization of the values, and log represents taking the logarithm.

[0082] Total classification loss L MSC Defined as:

[0083] L MSC =L cls +L B +λ4L N

[0084] Among them, λ4 represents the weight of the loss of potential new category objects.

[0085] Step S26, combining the distillation loss L KD and the total classification loss L MSC The model is trained to obtain a trained open vocabulary object detection model for multi-object classification.

[0086] Extensive experiments on the OV-COCO and OV-LVIS datasets show that the proposed CANet achieves significant and consistent performance improvements compared to other competitive methods.

[0087] The present invention is developed on the Ubuntu platform, and the developed deep learning framework is based on Pytorch. The main language used in the present invention is Python.

[0088] Take the OV-COCO and OV-LVIS datasets, train the model according to the above steps and test it using the test image sample set.

[0089] Table 1 and Table 2 are the test results of the present invention and other methods on two datasets. It can be found that compared with other methods, the present invention has the best effect. CANet (Ours) is the result of the present invention. We use the validation set to evaluate the performance of our method on the OV-COCO and OV-LVIS datasets. Evaluation indicators Represents the average precision (mAP) of the new category under the IoU threshold of 0.5; the evaluation indicator AP r, represents the mean average precision (mAP) of rare categories (new categories). The recognition performance of this method for new categories on the OV-COCO dataset reaches 38.5% (ResNet50-FPN), and the recognition performance of new categories on the OV-LVIS dataset reaches 24.1% (detection task) and 23.4% (segmentation task), which is higher than other methods, proving that the present invention has a better effect in identifying new categories of targets.

[0090]

[0091] Table 1 Comparison of model performance with the latest technical methods on OV-COCO

[0092]

[0093] Table 2 Comparison of model performance with state-of-the-art methods on OV-LVIS

[0094] The present invention also provides an open vocabulary target detection method system based on multi-target classification, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the method.

[0095] The computer program product comprises a computer program / instruction, and the method described is implemented when the computer program / instruction is executed by a processor.

[0096] The present invention also provides a computer-readable storage medium, in which the above-mentioned computer program is stored, and when the computer program is executed by a processor, the method described above is implemented.

Claims

1. An open vocabulary object detection method based on multi-object classification, characterized in that: Using context information, we can detect new categories of objects that have not appeared in the training set, including the following steps: Step S1, constructing an open vocabulary target detection model based on multi-target classification; the model includes a backbone feature extraction network ResNet50, a region proposal network RPN, a candidate box expansion module CBM, a distillation module KD and a multi-target classification module MSC; Step S2, training the model, specifically includes the following steps: Step S21, providing a training image sample set; Step S22, randomly select an image from the training image sample set, use the backbone feature extraction network ResNet50 to extract features from the image to obtain a feature map, and input the feature map into the region proposal network RPN; Step S23, the region proposal network RPN generates a set of candidate boxes in the feature map, and then applies the candidate box expansion module CBM to all the candidate boxes to obtain an extended box located near the candidate box for adding context information; Step S24, using the distillation module KD to learn knowledge from the CLIP image encoder, and using the backbone feature extraction network ResNet50 of the model to extract the features F of the extended box and the candidate box E ; Use CLIP image encoder to process the candidate frame and the expanded frame to obtain its feature F C , and then apply InfoNCE Loss to calculate the two corresponding features F E and F C Distillation loss L KD ; Step S25, using the CLIP text encoder to process the basic category names of the training image sample set to obtain the text feature F T , the feature F E and text features F T Input the candidate box into the multi-target classification module MSC to classify the candidate box and get the total classification loss L MSC ; Step S26, combining the distillation loss L KD and the total classification loss L MSC The model is trained to obtain a trained open vocabulary target detection model for multi-target classification; Step S3, input the image to be detected into the trained model, use the trained model to generate the predicted target area and the category name corresponding to the target area on the image, and give the confidence corresponding to this set of predictions. This confidence comes from the classification probability score of the target area by the multi-target classification module MSC.

2. The open vocabulary target detection method based on multi-target classification according to claim 1, characterized in that: Step S22: randomly select an image from the training image sample set as the image I∈R used in training. H×W×C , where H, W, and C represent the image height, image width, and image dimension respectively. The backbone feature extraction network ResNet50 is used to extract features from the image I to obtain a feature map; Step S23, generate a set of candidate boxes P = {p0, p1, ..., p n }, where p i represents the i-th candidate box, which contains four values, which are used to record the relative position of the i-th candidate box in the picture, and n is the number of candidate boxes; and in order to increase the context information, the candidate box expansion module CBM is applied to all candidate boxes P to obtain the expanded box B = {B0, B1, ..., B n }, expand box B i Adopt and p i The same expression, B i represents the i-th extended box, which also contains four values. The values ​​are used to record the relative position of the i-th extended box in the picture. n is the number of extended boxes. At the same time, the regression module is applied to optimize the candidate box P and obtain the regression loss L reg .

3. The open vocabulary target detection method based on multi-target classification according to claim 2, characterized in that: In step S23, first take the candidate box p i The eight nearby regions are used to expand each candidate box and obtain P ie ={p i ,r i1 ,…,r i8 }, where p i ∈P, where p i represents the original i-th candidate box selected previously, and r i1 ,…,r i8 Represents the candidate box p i Eight nearby areas; Then filter P ie , to remove the area beyond the image boundary and obtain P ir ={p i ,r i1 ,…,r ik }, where k represents the number of remaining regions, and k≤8; In order to find the existence of new types of objects, according to p i The similarity of P ir Sort by small to large, and shield the area with low objectivity score to get P i c ={p i ,r i1 ,…,r ij }, j≤k; Finally, from P i c Choose three with p i The boxes with the lowest similarity and compare them with p i Combine to form an expansion box.

4. The open vocabulary target detection method based on multi-target classification according to claim 3, characterized in that: The process of the last step in step S23 is expressed as: B i ={p i ,p im c ,p i(m+1) c ,p i(m+2) c } Among them, m represents the selection of a similarity region that is most likely to contain a new class of objects, and p im c , p i(m+1) c , p i(m+2) c The three surrounding areas are finally selected. The final expansion box B is obtained based on these three areas and the original area. i ; Expressed in formula, it is: p i c =Mask(Sort(Φ(p i ·r i ))); Where Φ(·) represents cosine similarity, Sort represents the process of calculating similarity and sorting, and Mask represents the selection of surrounding areas based on the similarity sorting results. i ∈P ir And r i ≠p i .

5. The open vocabulary target detection method based on multi-target classification according to claim 1, characterized in that: In step S24, the backbone feature extraction network ResNet50 of the model extracts the features of the extended box and the features of the candidate box F E :F E ={F E 1,F E 2,…,F E n ,F E (n+1) ,…,F E 2n }, where the first n features are the features of the expanded box, and the last n features are the features of the candidate box; Then use the CLIP image encoder to process the candidate box and the expanded box to obtain its feature F C :F C ={F C 1,F C 2,…,F C n ,F C (n+1) ,…,F C 2n }, where the first n features are the features of the expanded box, and the last n features are the features of the candidate box; Then apply InfoNCEloss to calculate the distillation loss L between the two corresponding features. KD , by optimizing the distillation loss L KD , so that the feature F generated by this model E The feature F generated by the CLIP image encoder C This is close to achieving the learning of knowledge in the CLIP image encoder.

6. The open vocabulary target detection method based on multi-target classification according to claim 5, characterized in that: In step S25, the category name is processed using the CLIP text encoder to obtain the text feature F T ={F T 1,F T 2,…,F T q }, where the meaning of q changes with the state of the model. When the model is training, since only the basic category information of the training image sample set can be used for training, q represents the number of basic categories in the training image sample set. During the test process, since the model detects objects of both basic categories and new categories at the same time, q represents the number of all categories in the test image sample set. Then, the feature F is classified by the multi-target classification module MSC. E and text features F T Classification is performed to obtain the classification loss L MSC ; First, the candidate frames are classified: the feature F extracted by the backbone feature extraction network ResNet50 of the present invention is E ={F E 1,F E 2,…,F E n+1 ,F E n+2 ,…,F E 2n }Into the multi-target classification module MSC, for each pair of extended boxes B and their corresponding candidate boxes p, the real category G = {g1, g2, ..., g n }, g1, g2, …, g n Corresponding to the basic target categories in the training image sample set; Then, by calculating F E i and F T gi , the cosine similarity of (gi∈G) generates the classification score, where F E i and F T gi F E and F T For each candidate box, calculate its classification loss L according to a single category. cls , for each expanded box, the classification loss is calculated, which is divided into L B and L N Two parts, among which L B Represents the loss of the basic category object, L N is the loss of potential new class objects, defined as: Where Φ(·) represents cosine similarity, Softmax represents normalization of the value, and log represents taking the logarithm; Select the top k categories with the highest classification scores G′={g1′,g2′,…,g k ′} as the aligned category, the loss is defined as: Where Φ(·) represents cosine similarity, softmax represents normalization of the value, and log represents taking the logarithm; Total classification loss L MSC Defined as: L MSC =L cls +L B +λ4L N Among them, λ4 represents the weight of the loss of potential new category objects.

7. An open vocabulary target detection method system based on multi-target classification, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the method according to any one of claims 1 to 6.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Zero sample semantic segmentation method, system and equipment based on knowledge distillation and medium

    CN115761235A

  • Social media content-oriented multi-target group classification method

    CN117094835A

  • Background prompt learning method in open vocabulary target detection

    CN118628712A

  • Multi-target tracking method and device for open vocabulary scene and medium

    CN118840393A

  • Method, system, device and medium for zero-shot semantic segmentation based on knowledge distillation

    GB2625638A

Cited By

  • Open target detection method and device, equipment and storage medium

    CN120707834A

  • An open target detection method, apparatus, device, and storage medium

    CN120707834B

  • Open vocabulary man-machine interaction detection method based on calibration diffusion model

    CN120949945A

  • An open-vocabulary human-robot interaction detection method based on a calibrated diffusion model

    CN120949945B