Zero sample target detection algorithm based on open environment

By combining ViT model and GANs network, auxiliary positioning modules and teacher-student networks are introduced, which solves the generalization ability and unknown category identification problems of models in open environments, and achieves efficient and accurate zero-sample object detection.

CN120259726APending Publication Date: 2025-07-04UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510248593.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing open environment object detection model has shortcomings in generalization ability and the accuracy of unknown categories, especially when facing unseen objects, the detection accuracy and accuracy are not high, and the long-tail effect leads to poor recognition performance.

Method used

The ViT model is used to combine generative adversarial networks (GANs) and auxiliary positioning modules to extract image features through self-attention mechanisms, and use the teacher-student network to update parameters to enhance the model's ability to identify unknown categories.

Benefits of technology

It significantly improves the accuracy and generalization ability of zero-sample object detection, can maintain efficient detection in unknown categories and complex environments, and improves the robustness and applicability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259726A_ABST
    Figure CN120259726A_ABST
Patent Text Reader

Abstract

The invention discloses a zero sample target detection method in an open environment. ViT and GANs are combined to construct a double-branch architecture. The traditional method depends on manual annotation or a weak generalization model, so that the detection precision in a dynamic scene is insufficient. According to the method, an image is shunted through ResNet; an upper branch integration area suggestion network (ARPN) and an auxiliary positioning network (ARPN) are used for positioning and calculating regression loss; the penultimate second attention head of the lower branch ViT is classified, and after the other branch is clustered, teacher-student network optimization parameters based on GANs are input. According to the method, end-to-end training is completed by combining regression loss and classification loss, and the detection efficiency is remarkably improved. A special data set construction method for the open environment and a multi-scale data enhancement strategy are innovatively provided, and the adaptability of the model to a complex background is enhanced. Compared with the traditional technology, the scheme breaks through the zero sample scene detection precision bottleneck, and provides a new normal form for open environment target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a zero-sample target detection method based on an open environment, belongs to the field of target detection in deep learning, and mainly realizes zero-sample target detection. Background Art

[0002] Zero-shot object detection based on open environments is a recently emerging research field with a wide range of application scenarios, such as:

[0003] Security monitoring: In the field of security monitoring, open target detection technology can greatly improve the system's detection of abnormalities.

[0004] The ability to respond to situations. Traditional surveillance systems often rely on the recognition of specific patterns to detect threats or abnormal behaviors, such as suspicious packages of known shapes or specific types of pedestrian behavior. However, threats in the real world may take on a variety of unpredictable forms. Open target detection enables the system to detect these previously undefined abnormal objects or behaviors and alert monitoring personnel in a timely manner.

[0005] Autonomous driving: Autonomous vehicles rely on accurate object detection to navigate and avoid collisions.

[0006] The object categories included in the training set of the autonomous driving system are relatively limited, but the real world is full of unknowns and surprises. For example, new traffic signs, non-standard engineering facilities or other unknown obstacles may appear on the road. Open object detection technology can help the autonomous driving system stay alert and respond appropriately when facing unknown objects, thereby enhancing safety.

[0007] Ecological monitoring: biological ecology research. For biologists and ecologists, it is crucial to be able to identify unknown species of plants and animals during field surveys. As the biodiversity of the earth continues to change, the discovery of new species is the norm. Using open object detection technology, researchers can automatically detect potential new species in image data from field surveys, speed up the process of species identification, and promote the recording and protection of biodiversity.

[0008] Military field: The application of open target detection in the military field includes enhancing situational awareness, target identification and tracking, threat assessment and monitoring, improving automated protection capabilities, and providing tactical support. These applications can improve the intelligence level of command and control systems and enhance the response speed and combat capabilities of the military. Although the current open environment target detection model has achieved automatic recognition of zero-sample targets to a certain extent, there are still some significant shortcomings:

[0009] 1. Weak generalization ability: Most existing object detection models rely on a large amount of labeled data for training. However, in practical applications, it is often difficult and time-consuming to collect and label a large number of zero-shot object images. This results in the model being prone to recognition errors or detection failures when encountering unseen environments or backgrounds.

[0010] 2. Long-tail effect: Due to the large distribution differences of zero-shot objects in different categories in nature, the number of samples of some rare objects is particularly scarce. This long-tail effect caused by sample imbalance makes the detection performance of the model for minority categories poor, affecting the overall recognition accuracy.

[0011] 3. Low recognition accuracy for unknown categories: A core challenge in open-environment object detection is to correctly handle unknown categories that have not been trained while recognizing known categories. However, the current models still need to improve their detection accuracy and reliability when facing new or unseen animal species, which limits the reliability of the technology in practical application scenarios.

[0012] In summary, in order to improve the efficiency of zero-shot object detection, a new algorithm based on open-environment object detection is urgently needed. This algorithm can overcome the limitations of existing technologies and improve the recognition speed and accuracy of zero-shot objects. Summary of the Invention

[0013] To address the above problems, the present invention provides a zero-shot object detection algorithm based on an open environment, aiming to overcome the limitations of existing technologies and improve the speed and accuracy of sample object recognition. This algorithm integrates the latest computer vision technologies, including the ViT (Vision Transformer) model and generative adversarial networks (GANs), and introduces an innovative auxiliary positioning module.

[0014] First, the present invention is based on the ViT model, which has attracted much attention in computer vision tasks due to its superior feature extraction ability. The ViT model effectively captures important features in the image through the self-attention mechanism, thus improving the performance of the object detector. At the same time, the present invention adopts a student-teacher model based on the GANs network to achieve dynamic update of the detector parameters. The student-teacher model can effectively improve the generalization ability of the detector for different categories by simulating the knowledge transfer process in the real scenario, especially maintaining good detection effects in the case of long-tail distributions with fewer samples.

[0015] Secondly, to enhance the recognition ability of unknown-class objects, the present invention designs an auxiliary positioning module. This module is specifically used to analyze and locate unseen classes that may exist in the image and distinguish the foreground from the background. By introducing this module, the system can more accurately identify and classify new species or abnormal individuals that are not in the training dataset, thereby significantly improving the reliability and applicability of the detector in an open environment. The technical solution adopted by the present invention is as follows:

[0016] Step 1: Design of a zero-shot object detection algorithm based on the ViT model

[0017] ViT adopts a different strategy. It draws on the idea of the Transformer model in NLP (Natural Language Processing), treats the image as a series of "words" (patches), and learns the long-range dependencies between different patches through the self-attention mechanism. In addition, before processing the image, image enhancement operations such as flipping, cropping, and adding noise are performed on the image.

[0018] Step S10: Preprocessing of the input picture: Input the zero-shot object picture into the ResNet backbone network, and after extracting features, divide it into two paths.

[0019] Step S11: Feature localization and regression loss calculation: The upper path generates the initial candidate box positions through the RPN network, and then introduces the ARPN (Auxiliary Localization Network) based on the Selective Search algorithm for further optimization. Combining the output results of the two modules, provides accurate object localization information and calculates the regression loss. Step S12: Feature classification and clustering: The lower path is processed through the ViT network to obtain the output of the penultimate attention head. This output is fed into a specially designed classification head to achieve the classification of known classes, and at the same time, allocate a part of the features to the clustering head to identify and cluster the unknown class vectors.

[0020] Step S13: Teacher-student network learning and update: Input the unknown class vector clusters obtained from the clustering head into the teacher-student network based on GANs for learning. This process aims to update the detector parameters through knowledge transfer to enhance the sensitivity to unseen objects in a changing environment.

[0021] Step S14: Total loss calculation: Sum the regression loss and the classification loss to obtain the total loss of the overall detection model, which is used to guide model training and parameter adjustment.

[0022] Step 2: Design of the auxiliary positioning module

[0023] Step S20: Introduce the ARPN module: Combine the existing RPN network and introduce an auxiliary localization module (ARPN) based on the Selective Search algorithm to further enhance the attention to potential unknown category regions. Improve the localization accuracy by distinguishing between the front and back scenes.

[0024] Step Three: Teacher-Student Network Based on GANs

[0025] Step S30: Construct a teacher-student network: Design and train a teacher-student network based on the GANs architecture so that it can learn the representation differences between unknown categories from the clustered vector clusters.

[0026] Step S31: Update the detector parameters: Utilize the learning results of the teacher-student network to provide effective parameter updates for the detector, thereby significantly enhancing the robustness in the open-set environment. Brief Description of the Drawings

[0027] Figure 1 It is: The flowchart of the zero-shot object detection algorithm based on the ViT model provided by the embodiments of the present invention.

[0028] Figure 2 It is: The structure diagram of the ARPN auxiliary localization module.

[0029] Figure 3 It is: The learning flowchart of the GANs teacher-student network.

[0030] Through the above technical solutions, the present invention provides an innovative open-set object detection method that can efficiently detect and classify zero-shot objects, not only having stable performance on known categories but also showing strong adaptability in unknown category recognition. Detailed Embodiments

[0031] The following will describe the detailed embodiments of the present invention with reference to the accompanying drawings. The embodiments of the present invention provide a zero-shot object detection algorithm based on an open environment, including the following steps:

[0032] 1. Dataset construction: Collect sample images under different environments and lighting conditions to ensure diversity. Use manual annotation or semi-automated annotation tools to record the category of each object and its position in the picture, especially accurately annotate known and unknown categories. Preprocess the original image data, such as resizing, cropping, and normalizing, to adapt to the input requirements of the model.

[0033] 2. Backbone network processing: The input image first passes through the ResNet backbone network to calculate the feature representation. Then, the feature map is divided into two paths for subsequent processing:

[0034] - The upper path is processed by the RPN (Region Proposal Network) and the ARPN (Auxiliary Localization Network) respectively. These two modules work together to improve the accuracy of object localization and output the regression loss.

[0035] - The other path enters the ViT (Vision Transformer) network. The output of the penultimate attention head is fed into the designed classification head to generate the classification loss. At the same time, the clustering head is used for clustering analysis of the unknown class vectors.

[0036] 3. Teacher-student model update: The vector clusters generated by clustering are input into the teacher-student network based on GANs (Generative Adversarial Networks). By iteratively updating the detector parameters, the generalization ability of the model in unknown class recognition is enhanced.

[0037] 4. Total loss calculation: The regression loss from the RPN and ARPN is added to the classification loss obtained from the classification head to obtain the final total loss function for guiding the model training process.

[0038] 5. Model validation and optimization: The performance of the model is evaluated on an independent test set, ensuring that the test set covers different scenarios and unseen zero-shot target samples. The effects are quantified by metrics such as precision, recall, and F1-score. The model parameters are continuously adjusted according to the feedback to optimize the detection accuracy.

[0039] 6. Deployment and application: The trained model is deployed into the zero-shot target monitoring system, which can be embedded in ground camera devices or mobile platforms to achieve real-time detection functions. A user-friendly interface is designed to display the detection results and provide interactive feedback, facilitating operators to take actions based on the information.

[0040] 7. Continuous learning and update: Feedback is collected according to the on-site usage situation, and the dataset is continuously updated, especially the error detection cases, to promote the evolution of the model and improve the recognition accuracy.

[0041] The above are only the specific implementation manners of the present invention. Any feature disclosed in the specification, unless clearly described, can be replaced by other equivalent or similar-purpose alternative features; all the disclosed features and steps, except for mutually exclusive features or steps, can be combined arbitrarily.

Claims

1. A zero-shot object detection algorithm based on an open environment, comprising the following steps: (a) Obtain a sample image through an image acquisition device; (b) Input the image into a ResNet backbone network and divide it into two paths: one path is transmitted to an RPN network and an auxiliary localization network (ARPN) for object localization and output of a regression loss; the other path passes through a Vision Transformer (ViT) network; (c) Extract features from the penultimate attention head of the ViT network and input them into a designed classification head to output a classification loss; (d) Use a clustering head to cluster unknown class vectors and input the vector clusters into a GANs-based teacher-student network for learning to update the parameters of the detector; (e) Sum the regression loss and the classification loss to obtain the final model loss.

2. The zero-shot object detection algorithm according to claim 1, wherein The auxiliary localization network is used to distinguish foreground and background in the image to locate potential unknown classes.

3. The zero-shot object detection algorithm according to claim 1 or 2, characterized in that, The classification head is based on the KNN algorithm to enhance classification accuracy.

4. The zero-shot object detection algorithm according to any one of claims 1 to 3, characterized in that, The GANs-based teacher-student network is used to improve the generalization ability of the model for unknown classes.

5. The zero-shot object detection algorithm according to any one of claims 1 to 4, characterized in that, The detector is applicable to mobile devices or cloud platforms to achieve fast on-site recognition or remote monitoring of zero-shot objects.

6. The zero-shot object detection algorithm according to any one of claims 1 to 5, characterized in that, Improve the generalization ability of the model for sample images through data augmentation techniques to cope with different lighting and environmental conditions.

7. A zero-shot object detection system, comprising: (a) An image acquisition unit for capturing a target image; (b) A processing unit configured with a ResNet backbone network, an RPN network, an ARPN network, a Vision Transformer network, a classification head, a clustering head, and a GANs-based teacher-student network for performing object detection and parameter update; (c) A storage unit for saving a sample image data set and deep learning model parameters; (d) An output unit for displaying detection results and providing interactive feedback; Wherein, the processing unit can exchange data with other devices or platforms through a wireless communication module.

8. The zero-shot object detection system according to claim 7, characterized in that, The image acquisition unit has an adaptive control module for automatically adjusting exposure and focus to meet the shooting requirements under various environmental conditions.

9. The zero-shot object detection system according to claim 7 or 8, wherein The output unit is a touch screen display for receiving user instructions and displaying relevant zero-shot object information.