Cross-domain few-sample target detection method based on single-stage network
Through the transformer-based encoding-decoding network and weight sharing mechanism, the problem of low accuracy in cross-domain few-sample object detection is solved, and high-precision and robust cross-domain detection are achieved to adapt to detection tasks in different domains.
Patent Information
- Application Number
- CN202511087161.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing single-stage network has low accuracy in cross-domain small sample object detection, which cannot effectively adapt to detection tasks in different domains, and cannot adapt to detection tasks in different domains well without fine-tuning.
A cross-domain few-sample object detection method based on a single-stage network is designed, and a transformer-based encoding-decoding network is adopted to extract features through an encoder with a weight sharing mechanism, and predict through the detection head, avoiding the inaccuracy of regional suggestions and achieving high accuracy and robustness of cross-domain detection.
The accuracy and robustness of detection are improved in cross-domain scenarios, and can adapt to detection tasks in different domains without fine-tuning, enhancing the learning ability and adaptability of the model.
Smart Images

Figure CN120580422A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and more particularly to a cross-domain few-sample target detection method based on a single-stage network. Background Art
[0002] Few-shot object detection is a widely used computer vision technique in data-limited situations. Its primary goal is to effectively detect novel object classes using limited labeled data. This technique typically trains on a base class with sufficient data and then identifies novel object classes based on a small number of examples. This approach effectively alleviates the need for large-scale labeled data collection and has garnered widespread attention from researchers in deep learning and computer vision.
[0003] Most current research in few-shot object detection assumes that the base and new classes are in the same domain, often achieving good performance. However, in real-world applications, many scenarios present the inevitable need for cross-domain few-shot object detection. For example, the base classes may consist of objects from everyday real-world scenarios (such as pedestrians, animals, or vehicles), while the new classes may originate from vastly different backgrounds, such as fish in underwater scenes, aircraft in remote sensing imagery, insects in nature, or defect detection targets in specific industrial environments. Furthermore, cross-domain few-shot object detection is challenging because the source and target domains often differ significantly in terms of environment, lighting conditions, scale variations, and contextual information. Therefore, models need strong robustness and generalization capabilities to adapt to the changes in the new domain while still accurately detecting objects. Research in this area is moving towards designing new algorithms, building effective models, and employing advanced techniques such as knowledge transfer, feature reuse, and self-supervised learning to improve the performance and application breadth of cross-domain few-shot object detection. Therefore, in response to the needs and challenges of cross-domain few-sample target detection, carrying out research and development of related technologies has important theoretical significance and practical value. It can promote technological progress in the field of computer vision and meet the needs of different industries for intelligent recognition.
[0004] Based on whether or not a region proposal network is used, current methods can be roughly divided into single-stage and two-stage methods. Two-stage methods typically include feature extraction, feature interaction, region proposal, similarity calculation, and detection steps. These methods first use a region proposal network to extract regions of interest from the query image. Then, they calculate the similarity between these regions and the support class features. By leveraging region priors, the model achieves better detection accuracy while adapting to diverse support classes. Single-stage methods, which include feature extraction, feature interaction, and detection, exclude the use of region proposals. While single-stage methods avoid the potential errors of inaccurate proposals, they have not yet demonstrated superior performance compared to two-stage methods. Furthermore, when the number of classes in the test set is unknown, single-stage methods are typically designed to process only one support class per feed-forward step, which limits their flexibility in multi-class scenarios. In contrast, two-stage methods achieve more precise region localization and adaptability, while single-stage methods emphasize simplicity but face challenges in complex detection scenarios.
[0005] Current few-shot detection and cross-domain few-shot detection methods typically consist of two components: a detection backbone for feature extraction and a detection head for classification and regression. In this context, region proposal networks (RPNs) serve as a crucial component of the detection backbone in two-stage approaches. Single-stage approaches, on the other hand, map support classes to a class-agnostic space, which is similarly integrated into the backbone. Both mechanisms help mitigate overfitting on base classes. However, they do not fully address class and domain bias in the detection head. In two-stage detection frameworks, RPNs are widely considered class-agnostic but are often biased towards the observed domain. While some methods have explored k-shot fine-tuning of RPNs on cross-domain datasets, improvements have been limited. This highlights the challenges of using region-based priors. In contrast, single-stage frameworks avoid reliance on proposals generated by RPNs by leveraging image-level structure. However, current single-stage work requires fine-tuning for novel classes, and without fine-tuning, they often fail to detect objects in cross-domain datasets. This suggests that current single-stage approaches suffer from poor domain robustness and are unable to adapt well to detection tasks across different domains without fine-tuning.
[0006] Therefore, how to improve the accuracy of cross-domain few-sample target detection of single-stage networks is an urgent problem that technicians in this field need to solve. Summary of the Invention
[0007] In view of this, the present invention provides a cross-domain few-sample target detection method based on a single-stage network. Based on the single-stage network, a transformer-based encoding-decoding network is designed, and the detection head is reinterpreted and defined. A single-stage cross-domain few-sample target detector with high cross-domain few-sample detection accuracy and good performance without fine-tuning is designed. This solves the problem that the existing cross-domain few-sample target detection methods have low accuracy in cross-domain scenarios due to inaccurate region proposals, and the problem that the existing methods cannot adapt to open set detection task scenarios.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A cross-domain few-shot object detection method based on a single-stage network includes the following steps:
[0010] Step 1: Collect query images and support sets, where the support sets include support set images of various categories;
[0011] Step 2: Build a cross-domain few-shot object detection network based on a single-stage network and train it using query images and support set images. The cross-domain few-shot object detection network includes a shared weight encoder, a decoder, and a detection head. The shared weight encoder extracts query image features and support class features and performs feature interaction on the query image features and support class features to obtain the target query feature representation. The decoder combines the target query feature representation with the input target query and outputs high-level semantic features of the target. The detection head makes predictions based on the high-level semantic features of the target to obtain the target detection result.
[0012] Step 3: Input the image to be detected into the trained cross-domain few-shot object detection network and output the object detection result.
[0013] Preferably, the shared weight encoder includes a query image encoder and a support set encoder that adopt a weight sharing mechanism; the query image encoder extracts query image features from the query image, and the support set encoder extracts support class features from the support set image. The layers with the same update parameters in the query image encoder and the support set encoder use the same weights, and the query image encoder fuses the query image features and the support class features to generate a target query feature representation.
[0014] Preferably, the support set encoder includes a basic feature extraction layer and a self-attention layer; the basic feature extraction layer uses the base large model DINOv2 to extract the support target features of the support set image; the self-attention layer adjusts the support target features through self-attention, obtains support class features, and transmits them to the query image encoder.
[0015] Preferably, the query image encoder includes a basic feature extraction layer, a cross-attention layer, a feedforward neural network and a feature fusion module; the basic feature extraction layer uses the base large model DINOv2 to extract the query image features of the query image; the cross-attention layer introduces cross-attention and uses the same weights as the self-attention layer of the support set encoder to calculate the correlation between the query image features to obtain the query image feature matrix; the feature fusion module fuses the query image feature matrix and the support class features to highlight the information of the support set image features in the query image to obtain refined query image features; the feedforward neural network uses linear layers and activation function layers to perform nonlinear transformation on the refined query image features to obtain the target query feature representation.
[0016] The technical effect of the above technical solution is that the key task of the query image encoder is to extract the global features and contextual information of the query image, and to highlight the support set image feature information in the query image through the feature interaction module; the efficiency and consistency of the calculation are achieved by weight sharing, and the modules with the same name in the two encoders use the same weights, that is, the query image and the support set image are sent to the encoder and inferred in parallel. After passing through the same parameter-containing modules, such as the basic feature extraction layer, the self-attention layer and the cross-attention layer, the support set image features are refined by the support set encoder, aiming to refine the features to highlight their importance. The shared weight mechanism promotes the collaborative work of the query features and the support features through effective information interaction, thereby enhancing the learning ability and accuracy of the model in the case of few samples. This information sharing design enables the target detection algorithm to better cope with visual differences in different fields, ensuring the accuracy and robustness of detection; by introducing cross-attention, the network can pay attention to other feature vectors in the entire input sequence when processing feature vectors, thereby capturing long-distance dependencies. For query image features, the attention mechanism automatically guides the network to focus on important areas in the image, thereby ignoring some interfering background information; the feedforward neural network follows the cross-attention operation and introduces nonlinear transformations by combining linear layers and activation functions to refine the encoder's feature representation capabilities.
[0017] Preferably, the feature fusion module includes a dot product layer, a channel rearrangement layer and a region of interest extraction layer; the dot product layer performs a dot product operation on the query image feature matrix and the support class features to obtain a product result; the channel rearrangement layer performs channel rearrangement on the product result to obtain subspace features; the region of interest extraction layer extracts the region of interest in the subspace features to obtain refined query image features.
[0018] Preferably, the decoder includes a self-attention layer, a cross-attention layer and a feedforward neural network; the self-attention layer applies a self-attention mechanism to the input K target queries to obtain the target query distribution; the cross-attention layer uses cross-attention to integrate the target query feature representation and the target query distribution at different levels of information, and fuses the target query feature representation and the target query distribution to obtain the target integrated feature; the feedforward neural network converts the target integrated feature into the target high-level semantic feature that describes the target category and target position.
[0019] Preferably, the detection head includes a position regression head and a category regression head; the position regression head predicts the target position based on the target high-level semantic features, and the category regression head predicts the target category based on the target high-level semantic features; the target position and target category constitute the target detection result.
[0020] Preferably, the category regression head determines the corresponding target category by parsing the structured information of the input support set image, associating the class probability predicted by the target high-level semantic features with the category position in the structured information of the support set image; the structured information is the category sequence arranged in order of the support set image, including the target category and category position.
[0021] Preferably, the process of class regression head predicting target class is:
[0022] Step 21: The support set contains N image samples, and the target domain dataset contains C categories. If , then the support set is filled with background placeholders, and the image samples outside the C class in the support set are filled with background placeholders. The image samples in the filled support set are expressed as , pos represents the image sample;
[0023] Step 22: Apply the sigmoid activation function to the target high-level semantic features corresponding to each target query output by the decoder to predict the class probability of the category position corresponding to the support set image ;
[0024] Step 23: Define the output sequence format of the category regression head, map the single-stage network to the meta-learning task, and map the corresponding relationship between the category position and class probability in the support set image according to the output sequence format. , S n Indicates the category position of the nth image sample;
[0025] Step 24: Filter the maximum class probability according to the set threshold, determine the category position based on the corresponding relationship, and output the corresponding target category based on the category position.
[0026] Through the above technical solution, it can be seen that compared with the prior art, the present invention discloses a cross-domain few-sample target detection method based on a single-stage network, which performs cross-domain few-sample detection tasks based on a single-stage network, thereby avoiding the inaccurate region proposals of the two-stage network. First, the query image and the support target are input into the same base large model to extract features, and then they are passed through a weight-sharing encoding-decoding network for more refined feature interaction and high-level semantic feature representation. Finally, the detection head is mapped to a meta-learning task related to the corresponding position of the support set to obtain the target position and category information. This method is mainly aimed at fine-tuning scenarios, and has achieved advanced performance under the conditions of domain differences and data limitations, while still having excellent performance under cross-domain open set detection conditions.
[0027] The present invention can be applied to fields such as autonomous driving, monitoring systems, industrial automation, and medical image analysis. It can improve the accuracy and robustness of the model in target detection in different fields or environments when target samples are scarce, thereby enhancing the adaptability of the computer vision system when processing unseen target samples. First, the query image and the support target are passed through a large base model with weight sharing to obtain query image features and support target features; then the support target features are first refined through a support set encoder to refine the support class features, and then the refined support class features and query image features are sent to the query image encoder. While extracting the global features and contextual information of the query image, feature interaction is performed with the support class features to highlight the support class feature information in the query image, provide a comprehensive understanding of the input image, and generate a feature representation for target detection; then the feature representation for target detection generated by the encoder is sent to the decoder, which is responsible for combining the feature representation generated by the encoder with the target query, and finally generates the category and position of each target through the detection head.
[0028] Based on an encoder-decoder network, this method effectively enhances the model's understanding and recognition of target objects through feature extraction, interactive fusion, and information reconstruction, thereby improving detection accuracy and robustness. It also integrates support class features and query image features to enrich the network's feature expression capabilities. Furthermore, it reinterprets and redefines the detection head, enabling the target detection system to maintain efficient and accurate performance across domains. This method has strong application value for cross-domain small-sample target detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0030] Figure 1 Schematic diagram of the cross-domain few-sample target detection network structure provided by the present invention;
[0031] Figure 2 This is a schematic diagram of the feature fusion module structure provided by the present invention. DETAILED DESCRIPTION
[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0033] The embodiment of the present invention discloses a cross-domain few-sample target detection method based on a single-stage network, comprising the following steps:
[0034] S1: Collect query images and support sets, which include support set images of various categories;
[0035] S2: Construct a cross-domain few-shot target detection network based on a single-stage network, and use the query image and support set image to train the cross-domain few-shot target detection network; the cross-domain few-shot target detection network is as follows: Figure 1 As shown in the figure, it includes a shared weight encoder, a decoder, and a detection head; the shared weight encoder extracts query image features and support class features, and performs feature interaction on the query image features and support class features to obtain the target query feature representation; the decoder combines the target query feature representation with the input K target queries and outputs the target high-level semantic features; the detection head makes predictions based on the target high-level semantic features to obtain the target detection results;
[0036] S3: Input the image to be detected into the trained cross-domain few-shot object detection network and output the object detection result.
[0037] Further, such as Figure 1 As shown, the shared weight encoder includes a query image encoder and a support set encoder that adopt a weight sharing mechanism; the query image encoder extracts query image features from the query image, and the support set encoder extracts support class features from the support set image. The same modules in the query image encoder and the support set encoder use the same weights, and the query image encoder fuses the query image features and the support class features to generate a target detection feature representation.
[0038] Furthermore, the support set encoder includes a basic feature extraction layer and a self-attention layer; the basic feature extraction layer uses the base large model DINOv2 to extract the support target features of the support set image; the self-attention layer adjusts the support target features through self-attention to obtain support class features and transmits them to the query image encoder.
[0039] The technical effect of the above technical solution is that the parameters of the large base model are kept frozen during the entire processing process as a feature extractor. The main purpose is to ensure that the model will not interfere with the accuracy of feature extraction due to noise introduced by subsequent data during the training phase; through self-attention, the features of each category can be further optimized to more accurately express relevant information in subsequent processing; after the self-attention layer, the supporting class features and the supporting target features are fused through residual connection (Add), and then processed by normalization (Norm) and region of interest (ROI) in turn to obtain the supporting class features.
[0040] Furthermore, the query image encoder includes a basic feature extraction layer, a cross-attention layer, a feedforward neural network and a feature fusion module; the basic feature extraction layer uses the base large model DINOv2 to extract the query image features of the query image; the cross-attention layer introduces cross-attention and uses the same weights as the self-attention layer of the support set encoder to calculate the correlation between the query image features to obtain the query image feature matrix; the feature fusion module fuses the query image feature matrix and the support class features to highlight the information of the support set image features in the query image to obtain the refined query image features; the feedforward neural network uses linear layers and activation function layers to perform nonlinear transformation on the refined query image features to obtain the target query feature representation.
[0041] The technical effect of the above technical solution is that the key task of the query image encoder is to extract the global features and contextual information of the query image, and to highlight the support set image feature information in the query image through the feature interaction module; the efficiency and consistency of the calculation are achieved by weight sharing, and the modules with the same name in the two encoders use the same weights, that is, the query image and the support set image are sent to the encoder and inferred in parallel. After passing through the same parameter-containing modules, such as the basic feature extraction layer, the self-attention layer and the cross-attention layer, the support set image features are refined by the support set encoder, aiming to refine the features to highlight their importance. The shared weight mechanism promotes the collaborative work of the query features and the support features through effective information interaction, thereby enhancing the learning ability and accuracy of the model in the case of few samples. This information sharing design enables the target detection algorithm to better cope with visual differences in different fields and ensure the accuracy and robustness of detection; by introducing cross-attention, the network can pay attention to other feature vectors in the entire input sequence when processing feature vectors, thereby capturing long-distance dependencies. For query image features, the attention mechanism automatically guides the network to focus on important areas in the image, thereby ignoring some interfering background information, and then fuses the extracted query image features and the query image feature matrix through residual connection (Add) and normalizes (Norm); the feature fusion module fuses the query image feature matrix and the support class features through residual connection (Add), and then transmits them to the feedforward neural network after normalization (Norm), and integrates them through residual connection and normalization to ensure the balance between different features and avoid information loss and overfitting; the feedforward neural network (FNN) follows the cross-attention operation and introduces nonlinear transformations by combining linear layers and activation functions to refine the feature representation capabilities of the encoder.
[0042] Throughout the encoder workflow, effective interaction between the support set image features and the query image features is crucial. Through a series of complex connections and calculations, the support set encoder and the query image encoder achieve full information sharing, enabling the network to better understand and process the correspondence between the support set information and the query image. This design not only improves the model's overall expressiveness but also enhances its adaptability to input data, enabling the encoder to maintain high performance and accuracy in complex visual environments, such as those with varying object sizes, shapes, or colors.
[0043] Furthermore, the feature fusion module includes a dot product layer, a channel rearrangement layer and a region of interest extraction layer; the dot product layer performs a dot product operation on the query image feature matrix and the support class features to obtain a product result; the channel rearrangement layer performs channel rearrangement on the product result to obtain subspace features; the region of interest extraction layer extracts the region of interest in the subspace features to obtain refined query image features.
[0044] Furthermore, the decoder includes a self-attention layer, a cross-attention layer and a feedforward neural network; the self-attention layer applies a self-attention mechanism to the input K target queries to obtain the target query distribution; the cross-attention layer uses cross-attention to integrate the target query feature representation and the target query distribution at different levels of information, and fuses the target query feature representation and the target query distribution to obtain the target integrated feature; the feedforward neural network converts the target integrated feature into the target high-level semantic feature that describes the target category and target position information.
[0045] The technical effect of the above technical solution is that the decoder network is responsible for aggregating and generating high-level semantic features for subsequent detection tasks. The self-attention layer combines each object query with the encoder's output features. When processing an object query, it is not limited to focusing on the current input features, but integrates the global feature information extracted by the encoder and can also obtain additional contextual information from the encoder output. This ensures that the decoder can attend to various relevant information when processing the object, while also structuring the network's response to improve object detection accuracy. Each object query corresponds to a potential object. The decoder uses the attention mechanism to model the relationship between each object query's features and the encoder's output features, ultimately generating a category prediction and corresponding position feature information for each potential object. The decoder can attend to various relevant information when processing the object, while also structuring the network's response to improve object detection accuracy. Through this hierarchical design and implementation, the decoder effectively transforms the abstract information output by the encoder into explicit category and position predictions, significantly improving the object detection capabilities of the entire network. The sophisticated design of the encoder and decoder complement each other, making the entire cross-domain few-shot object detection method not only have good performance but also exhibit strong robustness to diverse input data.
[0046] The high-level semantic features of the target not only provide the category information of the target, but also mark its position in the query image, providing support for decision-making and actions in practical applications. Through the design of this encoding-decoding network, useful information can be effectively extracted and aggregated, thereby achieving accurate target detection in complex visual scenes.
[0047] Furthermore, the detection head includes a position regression head and a category regression head; the position regression head predicts the target position based on the target high-level semantic features through a linear layer, and the category regression head predicts the target category based on the target high-level semantic features; the target position and target category constitute the target detection result.
[0048] Furthermore, the category regression head determines the corresponding target category by parsing the structured information of the input support set image, associating the class probability predicted by the target high-level semantic features with the category position in the structured information of the support set image; the structured information is the category sequence of the support set image arranged in order, including the target category and category position.
[0049] Furthermore, the process of predicting the target category by the category regression head is:
[0050] S21: The support set contains N image samples, and the target domain dataset contains C categories. If , then the support set is filled with background placeholders, and the image samples outside the C class in the support set are filled with background placeholders. The image samples in the filled support set are expressed as , pos represents the image sample;
[0051] S22: Apply the sigmoid activation function to the target high-level semantic features corresponding to each target query output by the decoder to predict the class probability of the category position of the corresponding support set image ;
[0052] S23: Define the output sequence format of the category regression head, map the single-stage network to the meta-learning task, and map the corresponding relationship between the category position and class probability in the support set image according to the output sequence format. , S n Indicates the category position of the nth image sample;
[0053] S24: Filter the maximum class probability according to the set threshold, determine the class position based on the corresponding relationship, and output the corresponding target class based on the class position.
[0054] The technical effect of the above technical solution is that the position regression head is a position-mapped detection head, mapping the category information in each query to a specific position in the support set sequence, rather than relying directly on the actual category label. This mapping mechanism enables the network to focus on the relative position of the object in space when recognizing the object, while ignoring the differences in specific categories. In this way, the network model not only reduces its reliance on category information but also effectively improves its ability to adapt to different environments and scenarios, thereby enhancing the performance and domain robustness of the object detection system. The position regression head focuses on features at specific locations, simplifying the complex object detection task into a more straightforward matching problem. Position mapping not only improves adaptability to new objects or changing environments, but also enhances the model's robustness to changes in data distribution. In particular, when there are significant visual differences between the support set and the query image, position mapping can effectively guide the network to perform matching in different domains, rather than being restricted to fixed category definitions. This flexibility provides a new approach for the network, enabling object detection systems to maintain efficient and accurate performance in changing environments.
[0055] Furthermore, the interactive method of support class features and query image features is characterized as follows: first, a highly efficient and accurate feature subspace is constructed. The design of this subspace is intended to significantly reduce the accuracy difference between the base class and the new class. This allows for a deep understanding and efficient extraction of target features during target detection. To ensure the effectiveness of this process, the support images of each category are modeled, their support class features are calculated, and the average vector of these features is made to accurately represent the characteristics of the category. In this context, we set is a set of category prototypes containing C categories, where D represents the dimension of the feature. The features of the query image are represented by representation. Under this setting, a method is constructed that treats the interaction between query image features and class prototype features as a means of compressing feature redundancy, effectively extracting the most essential information and achieving a more compact feature representation. First, during model design, only supporting prototypes from specific classes are used. This strategy ensures that the generated supporting features have greater specificity and independence. Using prototypes from all classes would result in a mixture of features, causing confusion in feature information and failing to accurately capture the unique characteristics of each class. Therefore, a separate subspace is designed to specifically characterize the object characteristics of each class, ensuring that each class can be effectively described within a unique feature domain. Furthermore, a background class is introduced to maintain robust independence between classes and ensure feature specificity. This design enables the model to intelligently ignore irrelevant interference during feature learning, effectively focusing on feature extraction from the supporting image, thereby improving learning efficiency and accuracy. The introduction of the background class not only reduces interference from irrelevant information but also helps enhance the model's ability to extract features from the target class, making the learning process more focused and clear. Furthermore, in the structural design of the feature subspace, multiple different feature mapping methods are designed to further enrich feature representation. This process ensures that each category has unique feature representations under specific conditions, improving the model's adaptability and accuracy in complex and dynamic environments. These design strategies collectively enable the model to maintain high detection performance and target recognition capabilities under diverse visual inputs. Regarding the interaction method between support class features and query image features, it involves not only the extraction of single-class features, but also in-depth mining of the correlations between features. Through an iterative learning process, the model can continuously improve its understanding of query image features by strengthening its focus on support class features, making the ultimately extracted feature representation more comprehensive and accurate. This systematic feature interaction mechanism provides strong support for cross-domain few-shot object detection tasks by constructing an ordered relationship between support class features, query features, and background features, thereby significantly improving the accuracy and robustness of target detection. This comprehensive approach not only enhances the model's ability to cope with new categories, but also lays a good foundation for subsequent detection tasks.
[0056] In a specific embodiment, the training process of the cross-domain few-shot object detection network is as follows:
[0057] S1: The cross-domain few-shot object detection network receives a query image, which is the image to be detected;
[0058] S2: The query image first passes through the basic feature extraction layer of the query image encoder, using the base large model DINOv2 to extract features and obtain query image features. This model uses the advanced capabilities of deep learning to extract key feature representations from the query image. This feature extraction process not only focuses on local information of the image, but also comprehensively considers global context information to form a more comprehensive feature representation.
[0059] S3: The query image features enter the cross-attention layer. This module learns multiple subspaces in parallel to obtain the query image feature matrix, effectively capturing the relationship between the query image's internal features and enhancing the model's attention and understanding of the target object.
[0060] S4: The attention query features are input into the feature fusion module, which further integrates and optimizes the extracted key features with the support class features extracted from the support set image. This allows the resulting feature representation to better reflect the specific information of the target object in subsequent detection steps, obtaining refined query image features. This is then processed into the target query feature representation by a feedforward neural network.
[0061] The support set images are processed as auxiliary samples and subjected to feature extraction using the same large base model to generate feature representations. These features are further refined using a multi-head self-attention mechanism to capture shared information and feature correlations between support class samples. After processing, the support target features are integrated by adding and normalizing modules to ensure information consistency. The support target features are also processed through residual connections, normalization, and region of interest (ROI) for targeted optimization, thereby enhancing the effectiveness of the feature representation and obtaining support class features.
[0062] S5: The decoder aggregates the target query feature representation and K target queries to generate target high-level semantic features;
[0063] S6: The high-level semantic features of the target are predicted in the detection head, and the target location and target category are output.
[0064] Furthermore, the support class features and the query image of the present invention are fused in the feature fusion module as follows: Figure 2 The paper describes a method for constructing feature subspaces in a few-shot object detection system for unseen classes. This approach aims to alleviate the generalization issues caused by training networks based solely on base classes. This method builds a low-dimensional subspace based on pre-trained ViT features to reduce the accuracy gap between base classes and novel classes, thereby improving the model's performance on novel classes.
[0065] First, the class prototype is defined as the class representative of each support image. Given the support samples of each class, the ViT features extracted by the base large model DINOv2 are calculated and the average features are obtained. The class prototype is constructed using the average features of each target class. , where C represents all classes and D is the channel dimension. At the same time, by extracting the features of the query image , and normalize it to length normalization as the unit to form the projection of the feature realization subspace, and the obtained feature subspace projection is P S The present invention takes into account the two main limitations of feature subspace construction. First, using only the prototype set C may not be able to fully capture the feature information, so a set of predefined background classes B is introduced to supplement the information P B , fill in the class prototype P S , thereby effectively constructing the subspace. Secondly, by classifying each class in class C and reordering other classes to resolve the arrangement ambiguity, the construction process of the feature subspace is further optimized. Specifically, the subspace feature proposed in the present invention is constructed by the following steps: First, calculate the class prototype P of class C S , and reorder the channels to form a more accurate feature representation. This process is expressed in the feature subspace projection as Expression, this formula combines the reordered features and prototype information to enhance the performance of target detection. In the present invention, the design of this method ensures that while retaining information between categories, it effectively improves the detection accuracy and inference speed of the model on new categories, overcoming the shortcomings of existing methods in processing samples of unseen categories. The present invention innovatively constructs a low-dimensional feature subspace, combines background class information and reordering strategies, and provides an efficient solution for cross-domain few-sample target detection. This solution has good application prospects and is suitable for various visual recognition tasks.
[0066] Furthermore, the present invention reinterprets and defines the detection head of the single-stage method and introduces background placeholders to support reasoning with an arbitrary number of target classes when the number of classes is unknown. When the support set contains N samples and the target domain dataset contains C classes, the background placeholders are used. To complete the unfilled part of the support set sequence, when In the category regression head, a sigmoid activation function is applied to the class prediction embedding vector of each query object to predict the class probability of the support set sequence position. By redefining the output sequence format of the class regression head, the single-stage network is mapped to a meta-learning task that corresponds to the class probability at a specific position. We achieve category-agnostic meta-learning by reinterpreting and redefining the detection head, alleviating domain bias and enhancing the domain robustness of the network.
[0067] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0068] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A cross-domain few-shot target detection method based on a single-stage network, characterized in that: The following steps are involved: Step 1: Collect query images and support sets, where the support sets include support set images of various categories; Step 2: Build a cross-domain few-shot object detection network based on a single-stage network and train it using query images and support set images. The cross-domain few-shot object detection network includes a shared weight encoder, a decoder, and a detection head. The shared weight encoder extracts query image features and support class features and performs feature interaction on the query image features and support class features to obtain the target query feature representation. The decoder combines the target query feature representation with the input target query and outputs the target high-level semantic features; The detection head makes predictions based on the high-level semantic features of the target to obtain the target detection results; Step 3: Input the image to be detected into the trained cross-domain few-shot object detection network and output the object detection result.
2. The cross-domain few-sample target detection method based on a single-stage network according to claim 1 is characterized in that The shared weight encoder includes a query image encoder and a support set encoder that adopt a weight sharing mechanism; the query image encoder extracts query image features from the query image, and the support set encoder extracts support class features from the support set image. The layers with the same update parameters in the query image encoder and the support set encoder use the same weights, and the query image encoder fuses the query image features and the support class features to generate a target query feature representation.
3. The cross-domain few-sample target detection method based on a single-stage network according to claim 2 is characterized in that The support set encoder includes a basic feature extraction layer and a self-attention layer; The basic feature extraction layer uses the base large model DINOv2 to extract the support target features of the support set image; The self-attention layer adjusts the support target features through self-attention, obtains the support class features, and transmits them to the query image encoder.
4. The cross-domain few-sample target detection method based on a single-stage network according to claim 3 is characterized in that The query image encoder includes a basic feature extraction layer, a cross-attention layer, a feedforward neural network, and a feature fusion module; The basic feature extraction layer uses the base large model DINOv2 to extract query image features of the query image; The cross attention layer introduces cross attention and uses the same weight as the self-attention layer of the support set encoder to calculate the correlation between the query image features to obtain the query image feature matrix; The fusion module fuses the query image feature matrix and the support class features to highlight the information of the support set image features in the query image and obtain the refined query image features; the feedforward neural network uses linear layers and activation function layers to perform nonlinear transformation on the refined query image features to obtain the target query feature representation.
5. The cross-domain few-sample target detection method based on a single-stage network according to claim 4 is characterized in that: The feature fusion module includes a dot product layer, a channel rearrangement layer, and a region of interest extraction layer; the dot product layer performs dot product operations on the query image feature matrix and the support class features to obtain the product result; the channel rearrangement layer performs channel rearrangement on the product result to obtain subspace features; the region of interest extraction layer extracts the region of interest in the subspace features to obtain refined query image features.
6. The cross-domain few-sample target detection method based on a single-stage network according to claim 2, characterized in that The decoder consists of a self-attention layer, a cross-attention layer, and a feedforward neural network. The self-attention layer applies a self-attention mechanism to the input target query to obtain the target query distribution. The cross-attention layer uses cross-attention to integrate the target query feature representation and target query distribution information at different levels, fusing the target query feature representation and target query distribution to obtain the target integrated feature. The feedforward neural network transforms the target integrated features into target high-level semantic features.
7. The cross-domain few-sample target detection method based on a single-stage network according to claim 6, characterized in that The detection head includes a position regression head and a category regression head; the position regression head predicts the target position based on the target's high-level semantic features, and the category regression head predicts the target category based on the target's high-level semantic features; The target location and target category constitute the target detection result.
8. The cross-domain few-sample target detection method based on a single-stage network according to claim 7, characterized in that: The category regression head parses the structural information of the input support set image, associates the class probability predicted by the target high-level semantic features with the category position in the structural information of the support set image, and determines the corresponding target category; The structured information is the category sequence of the support set images arranged in order, including the target category and category position.
Citation Information
Patent Citations
Few-sample defect detection method based on bilateral guidance network from pixel to global
CN115908314A
Robot inspection defect detection method based on cross attention and multi-dimensional measurement
CN119295732A