Small sample target detection method based on target feature enhancement and semantic fusion perception

By generating highly discriminative prototypes through dynamic hypergraph and semantic fusion perception modules, the problems of background noise interference and missing semantic information in small sample target detection are solved, higher detection accuracy and robustness are achieved, and the generalization ability of the model in small sample scenarios is improved.

CN120747465APending Publication Date: 2025-10-03XIAMEN UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510843472.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In existing small-sample target detection methods, prototype generation is easily affected by background noise, the representative differences of supported samples are ignored, and semantic information guidance is missing, resulting in insufficient prototype representation ability, affecting the model's generalization ability and detection accuracy in small-sample scenarios.

Method used

Dynamic hypergraph is used to enhance target features, background noise is suppressed by constructing a dynamic hypergraph structure, and a semantic fusion perception module is combined to generate a highly discriminative prototype. The variational autoencoder is used to optimize the semantic expression, and the regional features are fused through the channel attention mechanism to improve the distinguishability and robustness of the prototype.

Benefits of technology

It significantly improves the accuracy and robustness of small-sample target detection, enhances prototype representation capabilities, improves detection accuracy and cross-modal semantic fusion accuracy in complex scenarios, and enhances the generalization ability of the model in small-sample scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747465A_ABST
    Figure CN120747465A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample target detection method based on target feature enhancement and semantic fusion perception, and relates to a computer vision technology. A data set is divided into a query set and a support set, after features are extracted through a backbone network, background noise in the support features is inhibited through a dynamic hypergraph construction module, and high-order semantic association of a target area is enhanced; fusing the category name text semantics and the image specific prototype by using a semantic fusion perception module to generate a high-discrimination category prototype; modeling semantic distribution by means of a variational auto-encoder, and extracting variational features; and fusing the region-of-interest features and the variation features through a channel attention mechanism to realize classification and regression. According to the method, the prototype characterization capability is effectively improved, and experiments show that the method remarkably improves the detection precision and is suitable for labeling sample scarce scenes. More accurate small sample target detection is realized by enhancing the feature expression of the support feature map and the semantic meaning of the category prototype, and higher robustness and recognition performance are shown in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer vision technology, and in particular to a small sample target detection method based on target feature enhancement and semantic fusion perception. Background Art

[0002] As a key task in computer vision, small-shot object detection has significant application value in scenarios where labeled samples are scarce or expensive to obtain, such as medical image analysis and autonomous driving. Traditional object detection methods typically rely on large amounts of labeled data and require expensive model retraining when processing new categories of objects, resulting in poor adaptability under small-shot conditions. Small-shot object detection aims to accurately detect new categories of objects using only a small number of labeled samples, significantly reducing reliance on large-scale annotations. It has important practical significance and application value. Inspired by the human brain's ability to make analogical inductions and rapid recognition based on a small amount of experience, researchers have recently begun to explore the introduction of brain-inspired mechanisms, drawing on the brain's processing methods in visual perception, semantic understanding, and inductive reasoning to improve the model's generalization ability and learning efficiency under small-shot conditions.

[0003] Among existing methods, transfer learning-based strategies primarily adapt to new detection tasks by fine-tuning the general knowledge in pre-trained models, but they often suffer from underfitting or overfitting problems caused by insufficient data. In contrast, meta-learning methods, inspired by the brain's ability to "learn how to learn," have garnered widespread attention in recent years due to their rapid adaptability and good generalization in small-shot environments. Meta-learning-based small-shot object detection is typically trained by constructing a support set and a query set task. The support set is used to generate category prototypes, which are then matched to the region-of-interest features in the query set. Although meta-learning-based small-shot object detection has made significant progress, its performance remains unsatisfactory, primarily due to the poor quality of the generated prototypes. Specifically, directly aggregating the content within the object bounding box in the support image to construct the prototype often introduces background noise and fails to accurately capture the object's morphological details. The root cause of this problem is that the object bounding box in the support image typically precisely encloses the object, often including irrelevant background regions and the diversity of object morphology, which interferes with the representation of the support features. Furthermore, existing methods often employ a simple averaging strategy to aggregate image-level prototypes, ignoring the differences in representativeness among different support samples. In reality, some support samples may provide more representative features than others. Furthermore, existing methods often overlook the guiding role of semantic information in prototype generation. The introduction of semantic information can strengthen the focus on meaningful features, further optimizing prototype quality. Overall, these limitations lead to insufficient prototype representation capabilities, thus limiting the generalization ability of existing methods on query images.

[0004] Combined with these innovative designs, the proposed method improves the accuracy and robustness of small-sample object detection while demonstrating enhanced prototype discrimination. Experimental results demonstrate that this method significantly outperforms existing mainstream methods for small-sample object detection, providing new directions and insights for research and application in this field. Summary of the Invention

[0005] The purpose of the present invention is to address the technical problems in the existing technology such as serious background noise interference, ignored representative differences in supporting samples, and insufficient prototype representation ability caused by lack of semantic information guidance in prototype generation. The present invention provides a small-sample target detection method based on target feature enhancement and semantic fusion perception, which enhances target feature representation through dynamic hypergraph and generates highly discriminative prototypes by fusing visual and semantic information, thereby significantly improving the target detection accuracy and robustness in small-sample scenarios.

[0006] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions.

[0007] The small sample target detection method based on target feature enhancement and semantic fusion perception includes the following steps:

[0008] 1) Dataset partitioning: The training dataset is divided into a query set and a support set, which contains multiple categories and target instances. The target annotations include location and category.

[0009] 2) Feature extraction: The query set and support set images are input into the pre-trained backbone network to extract deep visual features and obtain query feature maps and support feature maps;

[0010] 3) Query region generation: The query feature map is used to generate candidate regions of interest through a region proposal network, and the region of interest features are obtained through region alignment.

[0011] 4) Target feature enhancement: The support feature map is fed into the dynamic hypergraph construction module, and the target-level features are enhanced and then the image-specific prototypes are extracted through the shared detection head;

[0012] 5) Semantic Fusion Prototype Generation: The semantic information of the category name and the image-specific prototype is integrated to generate the category-specific prototype through the semantic fusion perception module;

[0013] 6) Variational Autoencoder Modeling: Category-specific prototypes are input into a variational autoencoder, which is then encoded into a latent space distribution through the encoder. The decoder reconstructs the semantic representation and optimizes it using KL divergence loss, reconstruction loss, and consistency loss.

[0014] 7) Variational feature extraction: Extract variational features from latent space encoding to enhance the semantic expression of category prototypes;

[0015] 8) Regional feature fusion: The features of the region of interest are fused with the variational features through the channel attention mechanism to generate semantically enhanced regional features;

[0016] 9) Object Detection and Optimization: Classification and bounding box regression are performed based on semantically enhanced regional features, the total loss is calculated, and the network parameters are updated.

[0017] In step 1), the specific steps of dividing the dataset can be as follows: given a small sample dataset containing images and target annotations, the dataset is divided into a query set and a support set, where the support set contains categories, each with target instances, a total of support images; the query set consists of images containing several objects of interest, each of which is annotated with the precise location of the object and the corresponding category label; to adapt to small-sample learning tasks, this division follows the typical N-way K-shot setting, which helps the model to generalize effectively under limited samples.

[0018] In step 2), the specific steps of feature extraction can be: input the query set image and the support set image in step 1) into the pre-trained ResNet-101 backbone network respectively, extract their deep visual features, and thus obtain query feature maps respectively. and support feature maps , providing a solid feature foundation for subsequent matching and detection tasks.

[0019] In step 3), the specific steps of generating the query region can be as follows: the query feature map in step 2) is used to generate a set of candidate regions of interest through the region proposal network, and the regions of interest are aligned to obtain the region of interest feature. ,in is the number of regions of interest.

[0020] In step 4), the target feature enhancement is to transform the support feature map in step 2) into Input a target enhancement module based on a dynamic hypergraph structure to extract and enhance the key features related to the category target in the image; the target enhancement module constructs a dynamic hypergraph by modeling the high-order relationship between support samples and performs feature aggregation operations on the graph structure to obtain support features for target perception; the model process of the target enhancement module is as follows: Input support feature graph Represents the instance features within the bounding box in the image. The feature map is first processed by several convolutional layers and then divided into several local image blocks by the expansion layer operation. regions; each of size The region is flattened into a dimension of The one-dimensional feature vector of these regions is represented as and , which are used to construct edge-related features and node-related features in the hypergraph respectively; the pairwise cosine similarity between nodes is calculated to form a similarity matrix :

[0021]

[0022] in are learnable weights, Represents the L2 norm; all nodes are also associated with a set of learnable parameters , used to adjust the threshold of dynamic hyperedge construction based on similarity; node The dynamic threshold is calculated as in It is The mean of row similarity, is its maximum value; if a node With the Similarity of nodes , then the node is included in the The hyperedges of nodes; the incidence matrix of the hypergraph Defined as:

[0023]

[0024] in is the indicator function; the similarity matrix and the hypergraph structure are jointly encoded as a mask matrix , used to guide the dynamic aggregation of node features; the initial features of each node are updated to Thus, an enhanced feature set containing semantic association information is obtained ; Based on node aggregation, a general hypergraph convolution operation is introduced to further capture high-order structural dependencies; let the degree matrices of vertices and hyperedges be and , No. The layer hypergraph convolution is defined as:

[0025]

[0026] in is the hyperedge weight diagonal matrix, is the learnable transformation matrix, is the activation function; this process realizes the information flow and normalized aggregation between nodes and hyperedges in the hypergraph structure; finally, The node features after layer convolution are restored to the spatial feature map structure through the folding layer operation and fused element by element with the original feature map to form the final enhanced representation: , in Represents Hadamard multiplication; this mechanism introduces structural consistency and high-order semantic context while maintaining local details, improving the model's perception and discriminative expression of the target area; finally, a shared detection head is used to extract image-specific prototypes .

[0027] In step 5), the specific steps of generating the semantic fusion prototype may be:

[0028] The category name in step 1) and the image-specific prototype obtained in step 4) are fused through the semantic fusion perception module to obtain the category-specific prototype by fusing the semantic information of language and visual modalities. The process of the semantic fusion perception module is as follows: the module first obtains the image-specific prototype from step 4) ,in 、 、 Represent the number of categories, the number of samples per category and the feature dimension respectively; in order to avoid the information loss caused by simple averaging, the semantic fusion perception module introduces a fully connected layer based The attention mechanism applies weights to the features of each sample, and the category The initial prototype of

[0029]

[0030] in Indicates the Class Image-specific prototypes of samples; category names are semantically characterized by CLIP text encoder At the same time, the image-specific prototype is averaged and fully connected to obtain , ensuring alignment with the language feature dimension; then, the semantic fusion perception module generates cross-modal fusion features:

[0031]

[0032] in represents the Hadamard product, is a nonlinear transformation. Finally, in order to make full use of both image and text information, and Get the final category prototype:

[0033]

[0034] The hyperparameters Control the fusion ratio of visual information and semantic information.

[0035] In step 6), the variational autoencoder model is to transform the category-specific prototype in step 5) Input a variational autoencoder to model the category semantic distribution; the encoder will Mapping to latent space distribution: , the decoder is used to calculate the latent variables Reconstruct and maintain semantic consistency; during the training process, the variational autoencoder uses KL divergence loss ( ) guides the model to learn a stable latent space distribution, and then the reconstruction loss ( ) to maintain semantic fidelity, and finally through the consistency loss ( ) maintains the distinctiveness of each prototype, thereby enhancing the expressiveness and distribution separability of category representation.

[0036] In step 7), the variational feature extraction is performed by extracting a variational feature for each category from the latent space encoding in step 6). , this feature is used as a semantically enhanced category representation for subsequent regional feature semantic guidance; these variational features have stronger semantic expression and category discrimination capabilities than the original prototype.

[0037] In step 8), the regional feature fusion is to integrate the regional features of interest obtained in step 3) and the variational features obtained in step H , through a channel attention mechanism, a semantically enhanced regional feature representation that integrates category information is generated:

[0038]

[0039] in Represents the channel attention mechanism; the fused features have semantic information that is more sensitive to categories, which helps the model to more accurately identify target categories and locations.

[0040] In step 9), the specific steps of the target detection and optimization may be: feeding the semantically enhanced regional feature representation into classification and bounding box regression respectively, outputting the prediction results, calculating the total loss and updating the network parameters; the total loss is calculated as follows:

[0041]

[0042] in is the classification loss, is the regression loss, and the following three are the loss functions of the variational autoencoder. This multi-task joint optimization strategy ensures that the model achieves a good balance between target detection accuracy and semantic modeling.

[0043] The present invention generates high-quality prototypes by improving feature representation capabilities and enhancing the distinguishability of prototypes. Compared with the existing technologies, the present invention has the following outstanding technical effects and advantages:

[0044] 1. Significantly improved prototype representation capabilities: The target enhancement module constructed through a dynamic hypergraph effectively suppresses background noise interference and improves the target semantic purity of prototype features compared to traditional average aggregation methods. Experiments show that in the new class set 1 of the PASCAL VOC dataset, the present invention achieves nAP50 of 60.3%, 65.7%, 65.4%, 68.2%, and 67.9% in 1-, 2-, 3-, 5-, and 10-shot respectively. In scenarios with foreground and background confusion, the invention's "cat" detection accuracy is increased to 66%, fully demonstrating its robustness in complex scenarios.

[0045] 2. High precision of cross-modal semantic fusion: The semantic fusion perception module combines CLIP text semantics with visual features to significantly improve the discriminative power of the generated category prototypes. Experiments show that under 30-shot conditions, the nAP, nAP50, and nAP75 indicators of the MS COCO dataset reach 19.8%, 40.0%, and 17.8%, respectively.

[0046] 3. Strong generalization ability in small sample sizes: By modeling semantic distribution using a variational autoencoder, the model's robustness in few-shot scenarios is significantly enhanced. For example, in a 1-shot setting, the proposed method achieves nAP50 scores of 60.3%, 41.5%, and 50.8% in PASCAL VOC novel class sets 1, 2, and 3, respectively. The average performance across these three novel class sets surpasses that of the similar meta-learning-based method FCT by 17.8 percentage points. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is an overall flow chart of an embodiment of the present invention.

[0048] Figure 2 Schematic diagram of the target enhancement module.

[0049] Figure 3 Schematic diagram of the semantic fusion perception module.

[0050] Figure 4 A visualization of the experimental results. DETAILED DESCRIPTION

[0051] The method of the present invention is described in detail below with reference to the accompanying drawings and examples. This example is implemented based on the technical solution of the present invention, and provides an implementation method and specific operating process. However, the protection scope of the present invention is not limited to the following examples.

[0052] The present invention is first inspired by the pulse discharge pattern of brain neurons and the spatial positioning mechanism of the cerebellum, and designs a target enhancement module based on dynamic hypergraph construction to enhance the representation ability of supporting features. This module regards non-overlapping areas in the supporting feature graph as nodes in the hypergraph, and dynamically constructs hyperedges between highly similar areas, thereby realizing the propagation and aggregation of high-order semantic information through hypergraph convolution to enrich the feature representation ability. In addition, when constructing hyperedges related to the target, the target enhancement module effectively suppresses the interference of background noise by excluding background areas with low similarity, thereby reducing the negative impact of background on feature extraction during the hypergraph convolution aggregation process. Subsequently, inspired by the orientation selectivity mechanism of the primary visual cortex of the brain, the present invention designs a semantic fusion perception module to generate more representative and discriminative category-specific prototypes. This module improves the accuracy and discriminability of prototype expression by fusing the weighted representation of image-specific prototypes with text-based semantic information.

[0053] An embodiment of a small sample target detection method based on target feature enhancement and semantic fusion perception includes the following steps:

[0054] A. Given a small sample dataset containing images and object annotations, the data is divided into a query set and a support set according to the typical N-way K-shot setting, where the support set contains categories, each with target instances, a total of Support images ( Figure 1 The "plane, bicycle, and person" in the image are examples of N categories. The query set consists of images containing several objects of interest, each annotated with the object's precise location and corresponding category label. This partitioning strategy simulates small-sample learning scenarios and helps the model learn good class generalization capabilities even with extremely limited sample size.

[0055] B. Input the query set image and support set image in step A into the pre-trained ResNet-101 backbone network respectively to extract their deep visual features, thereby obtaining query feature maps respectively. and support feature maps These multi-level convolutional features serve as high-level semantic representations, providing a solid feature foundation for subsequent class feature matching and small-sample target detection tasks.

[0056] C. The query feature map in step B is passed through the region proposal network to generate a set of candidate regions of interest to cover the possible objects in the image. Then, the region of interest alignment is performed on each candidate region to obtain the region of interest features. ,in is the number of regions of interest, M is set to 300.

[0057] D. The support feature map in step B Input a target enhancement module based on a dynamic hypergraph structure to extract and enhance the key features related to the category target in the image. The flowchart of the target enhancement module is as follows: Figure 2 This module constructs a dynamic hypergraph by modeling the high-order relationships between support samples and performs feature aggregation operations on the graph structure to obtain target-aware support features. The model flow of the target enhancement module is as follows: Input support feature graph Represents the instance features within the bounding box in the image. The feature map is first processed by several convolutional layers and then divided into several local image blocks by the expansion layer operation. regions, of which Set to 2. Each size is The region is flattened into a dimension of The one-dimensional feature vector of The size of is 4096. These regional feature vectors are represented as and , which are used to construct edge-related features and node-related features in the hypergraph respectively. After passing through different fully connected layers, the pairwise cosine similarity between nodes is calculated to form a similarity matrix :

[0058]

[0059] in are learnable weights, Represents the L2 norm. All nodes are also associated with a set of learnable parameters , used to adjust the threshold of dynamic hyperedge construction based on similarity. Figure 2 As shown, the similarity matrix and learnable parameters Enter the dynamic hypergraph generation module, for nodes All are generated through a dynamic hyperedge, so the node The dynamic threshold is calculated as in It is The mean of row similarity, is its maximum value, is element-wise addition, is element-wise subtraction, is the dot product, corresponding to the addition, subtraction, and multiplication operations in the above formula. Each element in After a gating unit, if a node With the Similarity of nodes , then the node is included in the The hyperedge of each node (such as Figure 2 As shown in the hyperedges in the dynamic hypergraph generation module, 、 and The similarity is relatively large, connecting the nodes to form a hyperedge), thus obtaining Figure 2 The hypergraph shown in , the incidence matrix corresponding to the hypergraph The vertex set of the hypergraph is , the edge set is This method ensures that only nodes with high semantic similarity are included in the same hyperedge. The similarity matrix and the hypergraph structure are jointly encoded into a mask matrix The definition is as follows:

[0060]

[0061] The initial features of each node are updated under the guidance of its hyperedge neighbors as Thus, an enhanced feature set containing semantic association information is obtained On the basis of node aggregation, a general hypergraph convolution operation is introduced to further capture high-order structural dependencies. Let the degree matrices of vertices and hyperedges be and , No. The layer hypergraph convolution is defined as

[0062]

[0063] in is the hyperedge weight diagonal matrix, is the learnable transformation matrix, is the activation function. This process realizes the information flow and normalized aggregation between nodes and hyperedges in the hypergraph structure. Specifically, it is expressed as the aggregation of hyperedge features to form hyperedge features, and then the hyperedge features update the node features through the hyperedge. Finally, The enhanced node features after the layer convolution are restored to the spatial feature map structure through the folding layer operation to obtain the enhanced support feature map, and are fused element by element with the original feature map to form the final enhanced representation: , in Represents Hadamard multiplication. The number of layers of hypergraph convolution used in this invention is set to 2. This mechanism introduces structural consistency and high-order semantic context while maintaining local details, effectively improving the model's perception and discriminative expression of the target area. Finally, a shared detection head module is used to extract image-specific prototypes. .

[0064] E. The category name in step A and the image-specific prototype obtained in step D are fused through the semantic fusion perception module to fuse the semantic information of language and visual modalities to obtain the category-specific prototype. Figure 3 As shown in the figure, the process of the semantic fusion perception module is as follows: the module first obtains the image specific prototype from step D ,in 、 、 Represents the number of categories, the number of samples per category and the feature dimension respectively. To avoid the information loss caused by simple averaging, the semantic fusion perception module introduces a fully connected layer based And the attention mechanism of the normalized exponential function, which applies weights to the features of each sample, category The initial prototype of

[0065]

[0066] in Indicates the Class Image-specific prototypes of samples, Figure 3 in is the dot product, which calculates the product of the feature and the attention score. The category name (such as "airplane, bicycle, person", etc.) is extracted through the CLIP text encoder to extract the category text features. At the same time, the image-specific prototype is obtained after the mean layer and full connection mapping , ensuring alignment with the language feature dimension. Subsequently, the semantic fusion perception module generates cross-modal fusion features:

[0067]

[0068] in represents the Hadamard product, is a nonlinear transformation. Finally, in order to make full use of both image and text information, and Get the final category prototype:

[0069]

[0070] The hyperparameters Control the fusion ratio of visual information and semantic information. The value of is set to 0.7, which can fully control the fusion ratio of features. Figure 3 in Represents element-by-element addition, computing and The corresponding elements in are added together to obtain the final category-specific prototype.

[0071] F. Convert the category-specific prototypes from step E to Input a variational autoencoder to further model the category semantic distribution. The encoder will Mapping to latent space distribution: , the decoder is used to calculate the latent variables Reconstruct and maintain semantic consistency. During the training process, the variational autoencoder uses KL divergence loss ( ) guides the model to learn a stable latent space distribution, and then the reconstruction loss ( ) to maintain semantic fidelity, and finally through the consistency loss ( ) maintains the distinctiveness of each prototype, thereby enhancing the expressiveness and distribution separability of category representation.

[0072] H. Extract a variational feature for each category from the latent space encoding in step F This feature serves as a semantically enhanced category representation. Compared to the initial prototype, these features extracted from latent variables incorporate the latent semantic structure between categories, possessing stronger generalization and discriminative performance. They can better capture the essential differences between categories and effectively improve the semantic perception capabilities of subsequent detection tasks.

[0073] I. The region of interest features obtained in step C and the variational features obtained in step H , through a channel attention mechanism, a semantically enhanced regional feature representation that integrates category information is generated:

[0074]

[0075] in This mechanism adaptively adjusts the importance of each channel, effectively integrating regional features with category semantic information, and improving the model's perception of target categories and positioning accuracy.

[0076] J. The semantically enhanced regional features obtained in step I The input is sent to the classification head and the bounding box regression head respectively to predict the category of the target and its position bounding box. During the training phase, the model is optimized by the total loss function, which integrates multiple task objectives:

[0077]

[0078] in is the classification loss, is the regression loss, and the following three variational autoencoder-related losses jointly promote the stability, semantic consistency, and discriminability of category representation. This multi-task joint optimization strategy ensures that the model strikes a good balance between object detection accuracy and semantic modeling.

[0079] Figure 4 A visual comparison of object detection results using DPENet and a baseline method shows that the proposed method can better identify objects of interest in query images than existing methods. The proposed method significantly outperforms the baseline method in challenging scenarios such as foreground-background confusion, scale variation, and low confidence. For example, in the case of foreground-background confusion, the proposed method successfully detects a cat (with a confidence of 66%), while the baseline method fails. Regarding scale variation, the proposed method can detect sheep of varying sizes, while the baseline method misses the smaller one. Regarding low confidence, the proposed method improves dog detection confidence (from 66% and 82%, respectively) and bicycle detection performance from 83% to 90%, significantly outperforming the baseline method. Tables 1 and 2 show the nAP, nAP50, and nAP75 performance of the proposed method and existing methods on the small-sample object detection datasets PASCAL VOC and MS COCO, respectively. These significant improvements demonstrate that the proposed method can generate sufficiently discriminative class prototypes, enabling more accurate object detection.

[0080] The method utilizes the target enhancement module to highlight supporting features and improve feature representation capabilities; and generates more discriminative category-specific prototypes through the semantic fusion perception module.

[0081] This paper designs an object enhancement module that enables high-level semantic information interaction between highly similar regions through dynamic hypergraph construction, thereby highlighting supporting features and enhancing feature representation capabilities. Furthermore, a semantic fusion perception module is proposed, which uses a novel fusion strategy to combine visual features with semantic embeddings to generate more accurate and discriminative category-specific prototypes.

[0082] To verify the technical effects achieved by the present invention, experiments were conducted on the benchmark small-sample object detection dataset PASCAL VOC. The experimental results are shown in Table 1.

[0083] Table 1 Comparative experimental results on the PASCAL VOC dataset

[0084]

[0085] Experiments show that the present invention has significantly improved performance compared to existing technologies. Specifically, the average nAP50 performance of the present invention on new class sets 1, 2, and 3 reached 65.5%, 48.9%, and 57.0%, respectively.

[0086] In addition, additional experiments are conducted on the MS COCO dataset, and the experimental results are shown in Table 2.

[0087] Table 2 Comparative experimental results on the MS COCO dataset

[0088]

[0089] As shown in Table 2, the proposed method achieves 16.7% and 19.8% NAP in the 10-shot and 30-shot settings, respectively. These results demonstrate that the proposed method can accurately identify targets in small sample scenarios, adapt to diverse object recognition, and achieve greater robustness. By enhancing the feature representation of the supporting feature graphs and the semantics of the category prototypes, the proposed method achieves more accurate small sample target detection and demonstrates higher robustness and recognition performance in complex scenarios.

[0090] The above embodiments are only preferred embodiments of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the present invention.

Claims

1. A small sample target detection method based on target feature enhancement and semantic fusion perception, characterized by The following steps are involved: 1) Dataset partitioning: The training dataset is divided into a query set and a support set, which contains multiple categories and target instances. The target annotations include location and category. 2) Feature extraction: The query set and support set images are input into the pre-trained backbone network to extract deep visual features and obtain query feature maps and support feature maps; 3) Query region generation: The query feature map is used to generate candidate regions of interest through a region proposal network, and the region of interest features are obtained through region alignment. 4) Target feature enhancement: The support feature map is fed into the dynamic hypergraph construction module, and the target-level features are enhanced and then the image-specific prototypes are extracted through the shared detection head; 5) Semantic Fusion Prototype Generation: The semantic information of the category name and the image-specific prototype is integrated to generate the category-specific prototype through the semantic fusion perception module; 6) Variational Autoencoder Modeling: Category-specific prototypes are input into a variational autoencoder, which is then encoded into a latent space distribution through the encoder. The decoder reconstructs the semantic representation and optimizes it using KL divergence loss, reconstruction loss, and consistency loss. 7) Variational feature extraction: Extract variational features from latent space encoding to enhance the semantic expression of category prototypes; 8) Regional feature fusion: The features of the region of interest are fused with the variational features through the channel attention mechanism to generate semantically enhanced regional features; 9) Object Detection and Optimization: Classification and bounding box regression are performed based on semantically enhanced regional features, the total loss is calculated, and the network parameters are updated.

2. A small sample target detection method based on target feature enhancement and semantic fusion perception as described in claim 1, characterized in that In step 1), the specific steps of data set division are: given a small sample data set containing images and target annotations, the data set is divided into a query set and a support set, where the support set contains categories, each with target instances, a total of support images; the query set consists of images containing several objects of interest, each of which is annotated with the precise location of the object and the corresponding category label; to adapt to small-sample learning tasks, this division follows the typical N-way K-shot setting, which helps the model to generalize effectively under limited samples.

3. A small sample target detection method based on target feature enhancement and semantic fusion perception as described in claim 1, characterized in that In step 2), the specific steps of feature extraction are: inputting the query set image and the support set image in step 1) into the pre-trained ResNet-101 backbone network respectively, extracting their deep visual features, and thus obtaining query feature maps respectively. and support feature maps , providing a solid feature foundation for subsequent matching and detection tasks.

4. A small sample target detection method based on target feature enhancement and semantic fusion perception as claimed in claim 1, characterized in that In step 3), the specific steps of generating the query region are as follows: the query feature map in step 2) is used to generate a set of candidate regions of interest through the region proposal network, and the regions of interest are aligned to obtain the region of interest feature. ,in is the number of regions of interest.

5. A small sample target detection method based on target feature enhancement and semantic fusion perception as claimed in claim 1, characterized in that In step 4), the target feature enhancement is to transform the support feature map in step 2) into Input an object enhancement module built based on a dynamic hypergraph structure to extract and enhance key features related to the category target in the image; The target enhancement module constructs a dynamic hypergraph by modeling high-order relationships between support samples and performs feature aggregation operations on the graph structure to obtain target-aware support features. The model process of the target enhancement module is as follows: Input support feature map Represents the instance features within the bounding box in the image. The feature map is first processed by several convolutional layers and then divided into several local image blocks by the expansion layer operation. regions; each of size The region is flattened into a dimension of The one-dimensional eigenvectors of The regional feature vectors are expressed as and , which are used to construct edge-related features and node-related features in the hypergraph respectively; The pairwise cosine similarity between nodes is calculated to form a similarity matrix : in are learnable weights, Represents the L2 norm; all nodes are also associated with a set of learnable parameters , used to adjust the threshold of dynamic hyperedge construction based on similarity; node The dynamic threshold is calculated as in It is The mean of row similarity, is its maximum value; If a node With the Similarity of nodes , then the node is included in the The hyperedges of nodes; the incidence matrix of the hypergraph Defined as: in is the indicator function; the similarity matrix and the hypergraph structure are jointly encoded as a mask matrix , used to guide the dynamic aggregation of node features; the initial features of each node are updated to Thus, an enhanced feature set containing semantic association information is obtained ; Based on node aggregation, a general hypergraph convolution operation is introduced to further capture high-order structural dependencies; let the degree matrices of vertices and hyperedges be and , No. The layer hypergraph convolution is defined as: in is the hyperedge weight diagonal matrix, is the learnable transformation matrix, is the activation function; this process realizes the information flow and normalized aggregation between nodes and hyperedges in the hypergraph structure; finally, The node features after layer convolution are restored to the spatial feature map structure through the folding layer operation and fused element by element with the original feature map to form the final enhanced representation: , in Represents Hadamard multiplication; this mechanism introduces structural consistency and high-order semantic context while maintaining local details, effectively improving the model's perception and discriminative expression of the target area; finally, a shared detection head is used to extract image-specific prototypes .

6. A small sample target detection method based on target feature enhancement and semantic fusion perception as claimed in claim 1, characterized in that In step 5), the specific steps of generating the semantic fusion prototype are: The category name in step 1) and the image-specific prototype obtained in step 4) are fused through the semantic fusion perception module to obtain the category-specific prototype by fusing the semantic information of language and visual modalities. The process of the semantic fusion perception module is as follows: the module first obtains the image-specific prototype from step 4) ,in 、 、 Represents the number of categories, the number of samples in each category and the feature dimension respectively; In order to avoid the information loss caused by simple averaging, the semantic fusion perception module introduces a fully connected layer The attention mechanism applies weights to the features of each sample, and the category The initial prototype of in Indicates the Class Image-specific prototypes of samples; category names are semantically characterized by CLIP text encoder At the same time, the image-specific prototype is averaged and fully connected to obtain , ensuring alignment with the language feature dimension; Subsequently, the semantic fusion perception module generates cross-modal fusion features: in represents the Hadamard product, is a nonlinear transformation. Finally, in order to make full use of both image and text information, and Get the final category prototype: The hyperparameters Control the fusion ratio of visual information and semantic information.

7. A small sample target detection method based on target feature enhancement and semantic fusion perception as claimed in claim 1, characterized in that In step 6), the variational autoencoder model is to transform the category-specific prototype in step 5) Input a variational autoencoder to model the category semantic distribution; the encoder will Mapping to latent space distribution: , the decoder is used to calculate the latent variables Reconstruct and maintain semantic consistency; during the training process, the variational autoencoder uses KL divergence loss ( ) guides the model to learn a stable latent space distribution, and then the reconstruction loss ( ) to maintain semantic fidelity, and finally through the consistency loss ( ) maintains the distinctiveness of each prototype, thereby enhancing the expressiveness and distribution separability of category representation.

8. A small sample target detection method based on target feature enhancement and semantic fusion perception as claimed in claim 1, characterized in that In step 7), the variational feature extraction is performed by extracting a variational feature for each category from the latent space encoding in step 6). , this feature is used as a semantically enhanced category representation for subsequent regional feature semantic guidance; these variational features have stronger semantic expression and category discrimination capabilities than the original prototype.

9. A small sample target detection method based on target feature enhancement and semantic fusion perception as claimed in claim 1, characterized in that In step 8), the regional feature fusion is to integrate the regional features of interest obtained in step 3) and the variational features obtained in step H , through a channel attention mechanism, a semantically enhanced regional feature representation that integrates category information is generated: in Represents the channel attention mechanism; the fused features have semantic information that is more sensitive to categories, which helps the model to more accurately identify target categories and locations.

10. A small sample target detection method based on target feature enhancement and semantic fusion perception as claimed in claim 1, characterized in that In step 9), the specific steps of the target detection and optimization are: feeding the semantically enhanced regional feature representation into classification and bounding box regression respectively, outputting the prediction results, calculating the total loss and updating the network parameters; the total loss is calculated as follows: in is the classification loss, is the regression loss, and the following three are the loss functions of the variational autoencoder. This multi-task joint optimization strategy ensures that the model achieves a good balance between target detection accuracy and semantic modeling.

Citation Information

Cited By

  • OCR (optical character recognition) method and system based on small sample image cutting data

    CN120954001A

  • Pipeline crack detection method, device, equipment and medium under condition of few sample data

    CN121388406A

  • Method and device for obtaining key semantic enhancement features and medium

    CN121527598A

  • SAR directed target detection method based on multi-scale context sensing

    CN121582554A

  • A sar directed target detection method based on multi-scale context perception

    CN121582554B