Target detection method, device, electronic device and storage medium

Through the adaptive semantic-spatial joint attention mechanism and clustering knowledge distillation technology, the detection categories are dynamically adjusted, which solves the problem that the existing technology can only detect fixed categories, realizes the fusion of image and text information, and improves the flexibility and accuracy of target detection.

CN120375101BActive Publication Date: 2025-09-05FIBOCOM WIRELESS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510860069.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-05
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The target detection model in the existing technology can only detect fixed-category targets that are manually labeled and cannot break through the detection limitations of fixed categories.

Method used

By obtaining the image information and text information to be detected, the adaptive semantic-spatial joint attention mechanism is used for feature fusion, combined with clustering and knowledge distillation, the detection category is dynamically adjusted and the target category is automatically expanded.

Benefits of technology

It realizes the detection of targets outside of fixed categories and dynamically adjusts the detection categories, breaking through the limitations of traditional methods and improving the flexibility and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375101B_ABST
    Figure CN120375101B_ABST
Patent Text Reader

Abstract

The present application relates to a target detection method, device, electronic device and storage medium, which includes: obtaining image information and text information to be detected; performing feature extraction and fusion on the image information and text information to be detected respectively to obtain a candidate area and a feature vector V_raw corresponding to the candidate area; encoding the feature vector V_raw and text information corresponding to the candidate area respectively to obtain an embedding vector V corresponding to the candidate area and an embedding vector T corresponding to the text information, and calculating the similarity between the embedding vector V corresponding to the candidate area and the embedding vector T corresponding to the text information to obtain a similarity score Similarity corresponding to the candidate area; clustering and knowledge distillation of the feature vector V_raw corresponding to the candidate area to obtain a semantic embedding Linear corresponding to the candidate area; classifying the candidate area into target categories based on the similarity score Similarity corresponding to the candidate area and the semantic embedding Linear corresponding to the candidate area, thereby solving the problem of detection limitations of fixed categories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a target detection method, device, electronic device, and storage medium. Background Art

[0002] With the continuous development of computer vision technology, object detection plays a vital role in many application fields, such as autonomous driving and intelligent monitoring.

[0003] However, related technologies typically use traditional models such as YOLOv12 for object detection. However, this approach can only detect objects of fixed categories that are manually annotated in the model, and therefore cannot overcome the detection limitations of fixed categories. Therefore, how to detect objects of categories other than fixed categories has become a pressing technical problem. Summary of the Invention

[0004] The present application provides a target detection method, device, electronic device and storage medium to solve the technical problem in the prior art that it can only detect targets of fixed categories manually marked in the model, and therefore cannot break through the detection limitations of fixed categories.

[0005] In a first aspect, an embodiment of the present application provides a target detection method, the method comprising:

[0006] Acquire image information and text information to be detected, wherein the text information is used to describe the expected category of the target object to be detected;

[0007] Feature extraction is performed on the image information to be detected and the text information respectively, and the extracted features are fused based on the adaptive semantic-spatial joint attention mechanism to obtain a candidate region and a feature vector V_raw corresponding to the candidate region;

[0008] Encode the feature vector V_raw and the text information corresponding to the candidate region respectively to obtain the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information, and calculate the similarity between the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information to obtain the similarity score Similarity corresponding to the candidate region;

[0009] Performing clustering and knowledge distillation on the feature vector V_raw corresponding to the candidate region to obtain a semantic embedding Linear corresponding to the candidate region, wherein the semantic embedding Linear corresponding to the candidate region is used to characterize the potential category to which the candidate region belongs, and the potential category includes the desired category and an extended category related to the desired category;

[0010] Based on the similarity score Similarity corresponding to the candidate area and the semantic embedding Linear corresponding to the candidate area, the candidate area is classified into target categories to obtain a target detection result, wherein the target detection result is used to characterize the detected target object and the category to which the target object belongs, and the category to which the target object belongs is the expected category, the extended category and / or the unknown category.

[0011] Optionally, the feature extraction is performed on the image information to be detected and the text information respectively, and the extracted features are fused based on an adaptive semantic-spatial joint attention mechanism to obtain a candidate region and a feature vector V_raw corresponding to the candidate region, including:

[0012] Performing feature extraction on the image information to be detected and the text information respectively to obtain a feature map corresponding to the image information to be detected and an embedding vector T corresponding to the text information;

[0013] Based on the adaptive semantic-spatial joint attention mechanism, the spatial attention weight W_s and the semantic attention weight W_t are obtained, and feature fusion is performed according to the spatial attention weight W_s and the semantic attention weight W_t to obtain a fused feature map, wherein the spatial attention weight W_s is obtained by performing a convolution operation on the feature map corresponding to the image information to be detected, and the semantic attention weight W_t is captured from the feature map corresponding to the image information to be detected and the embedding vector T corresponding to the text information through a multi-head self-attention mechanism;

[0014] Perform candidate region detection on the fused feature map to obtain a feature vector V_raw corresponding to the candidate region.

[0015] Optionally, performing feature fusion according to the spatial attention weight W_s and the semantic attention weight W_t to obtain a fused feature map includes:

[0016] Obtaining the mean of the feature map corresponding to the image information to be detected and the variance of the embedding vector T corresponding to the text information;

[0017] Determining a value of an adaptive coefficient α based on a mean value of a feature map corresponding to the image information to be detected and a variance of an embedding vector T corresponding to the text information;

[0018] Using the value of the adaptive coefficient α, the spatial attention weight W_s and the semantic attention weight W_t are fused to obtain a joint attention weight W_j;

[0019] Based on the joint attention weight W_j, the feature map corresponding to the image information to be detected is updated to obtain the fused feature map.

[0020] Optionally, clustering and knowledge distilling the feature vector V_raw corresponding to the candidate region to obtain a semantic embedding Linear corresponding to the candidate region includes:

[0021] Using a preset clustering algorithm and pre-set initial cluster centers, cluster the feature vectors V_raw corresponding to the candidate regions to obtain updated cluster centers, wherein the number of the updated cluster centers can be adaptively adjusted according to the variance of the batch features participating in each clustering;

[0022] Mapping the feature vector V_raw corresponding to the candidate region and the updated cluster center to the same space, and calculating the distillation loss value between the candidate region and each cluster center in the updated cluster center;

[0023] Determining a potential category corresponding to the candidate region based on the distillation loss value;

[0024] According to the potential category belonging to the candidate region, a semantic embedding Linear corresponding to the candidate region is determined.

[0025] Optionally, determining the potential category corresponding to the candidate region based on the distillation loss value includes:

[0026] Determine whether there is a target cluster center among the updated cluster centers, the distillation loss value between which the target cluster center and the candidate region is less than a first preset threshold, wherein the target cluster center is any cluster center among the updated cluster centers;

[0027] When there is a target cluster center among the updated cluster centers whose distillation loss value with the candidate region is less than the first preset threshold, assigning the potential category corresponding to the candidate region to the category corresponding to the target cluster center;

[0028] If there is no target cluster center in the updated cluster centers whose distillation loss value with the candidate area is less than the first preset threshold, a new cluster center is added based on the updated cluster centers, and the potential category corresponding to the candidate area is assigned to the category corresponding to the new cluster center.

[0029] Optionally, the target category classification of the candidate region based on the similarity score corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region to obtain the target detection result includes:

[0030] Compare the similarity score Similarity corresponding to the candidate region with a second preset threshold;

[0031] Determine the target category of the candidate area whose similarity score Similarity is greater than the second preset threshold as the expected category, and calculate the similarity score Sim_C corresponding to the candidate area whose similarity score Similarity is less than or equal to the second preset threshold, wherein the similarity score Sim_C is used to represent the similarity between the semantic embedding Linear of the cluster center to which the candidate area whose similarity score Similarity is less than or equal to the second preset threshold belongs and the embedding vector T corresponding to the text information;

[0032] Comparing the similarity score Sim_C with a third preset threshold;

[0033] The target category of the candidate area whose similarity score Sim_C is greater than the third preset threshold is determined as the extended category, and the target category of the candidate area whose similarity score Sim_C is less than or equal to the third preset threshold is determined as the unknown category.

[0034] Optionally, after classifying the candidate region into target categories based on the similarity score corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region to obtain the target detection result, the method further includes:

[0035] The target object corresponding to the expected category, the target object corresponding to the extended category, and the target object corresponding to the unknown category are visually output.

[0036] In a second aspect, an embodiment of the present application further provides a target detection device, the device comprising:

[0037] An acquisition module, configured to acquire image information and text information to be detected, wherein the text information is used to describe the expected category of the target object to be detected;

[0038] A feature extraction and fusion module is used to extract features from the image information to be detected and the text information respectively, and fuse the extracted features based on an adaptive semantic-spatial joint attention mechanism to obtain a candidate region and a feature vector V_raw corresponding to the candidate region;

[0039] An encoding module is used to encode the feature vector V_raw corresponding to the candidate region and the text information respectively, obtain the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information, and calculate the similarity between the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information to obtain the similarity score Similarity corresponding to the candidate region;

[0040] A clustering and knowledge distillation module is used to perform clustering and knowledge distillation on the feature vector V_raw corresponding to the candidate region to obtain a semantic embedding Linear corresponding to the candidate region, wherein the semantic embedding Linear corresponding to the candidate region is used to characterize the potential category to which the candidate region belongs, and the potential category includes an extended category related to the desired category;

[0041] The target category classification module is used to classify the candidate area into target categories based on the similarity score corresponding to the candidate area and the semantic embedding Linear corresponding to the candidate area, so as to obtain a target detection result, wherein the target detection result is used to characterize the detected target object and the category to which the target object belongs, and the category to which the target object belongs is the expected category, the extended category and / or the unknown category.

[0042] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0043] Memory for storing computer programs;

[0044] The processor is used to implement the target detection method described in any one of the embodiments of the first aspect when executing the program stored in the memory.

[0045] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the target detection method as described in any embodiment of the first aspect.

[0046] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:

[0047] The method provided by the embodiment of the present application obtains image information and text information to be detected, wherein the text information is used to describe the expected category of the target object to be detected; performs feature extraction on the image information to be detected and the text information respectively, and fuses the extracted features based on the adaptive semantic-spatial joint attention mechanism to obtain a candidate area and a feature vector V_raw corresponding to the candidate area; encodes the feature vector V_raw corresponding to the candidate area and the text information respectively to obtain an embedding vector V corresponding to the candidate area and an embedding vector T corresponding to the text information, and calculates the similarity between the embedding vector V corresponding to the candidate area and the embedding vector T corresponding to the text information to obtain a similarity score Similar corresponding to the candidate area ity; clustering and knowledge distillation are performed on the feature vector V_raw corresponding to the candidate region to obtain the semantic embedding Linear corresponding to the candidate region, wherein the semantic embedding Linear corresponding to the candidate region is used to characterize the potential category affiliation of the candidate region, and the potential category includes the expected category and the extended category related to the expected category; based on the similarity score Similarity corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region, the candidate region is classified into target categories to obtain a target detection result, wherein the target detection result is used to characterize the detected target object and the category to which the target object belongs, and the category to which the target object belongs is the expected category, the extended category and / or the unknown category. Through the above method, the features of the image information to be detected and the text information can be dynamically fused so that the detected target category can be dynamically adjusted according to the description of the text information, rather than being limited to the fixed category manually labeled; and by clustering and knowledge distilling the feature vector V_raw that fuses the image information to be detected and the text information, the target category can be automatically expanded, thereby further solving the problem of the inability to break through the detection limitations of fixed categories in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0049] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0050] Figure 1 A schematic diagram of a flow chart of a base target detection method provided in an embodiment of the present application;

[0051] Figure 2 A network structure of an ASSJA module provided in an embodiment of the present application;

[0052] Figure 3 A schematic diagram of the structure of a target detection device provided in an embodiment of the present application;

[0053] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0055] See also Figure 1 , Figure 1 This is a flow chart of a base target detection method provided in an embodiment of the present application. Figure 1 As shown, the target detection method may include the following steps:

[0056] Step S101: Acquire image information and text information to be detected, wherein the text information is used to describe the expected category of the target object to be detected.

[0057] Specifically, the image information to be detected can be any image or image sequence that requires target detection. For example, in the field of autonomous driving, the image information to be detected can be image information collected by a vehicle camera; in the field of intelligent monitoring, the image information to be detected can be image information collected by a camera in a monitoring system, etc. The image information to be detected can be a color image with three color channels: red, green, and blue, or an image in another format. The text information can be used to describe the desired category of the target object to be detected, such as text content input by the user, such as "pedestrian wearing a hat" or "pedestrian wearing red clothes."

[0058] Step S102: Feature extraction is performed on the image information and text information to be detected respectively, and the extracted features are fused based on the adaptive semantic-spatial joint attention mechanism to obtain the candidate region and the feature vector V_raw corresponding to the candidate region.

[0059] Specifically, learning models such as the Region Convolutional Neural Network (RCNN) series, the YOLO (You Only Look Once) series, the Scalable and Efficient Object Detection (EfficientDet), and the Mask Region-based Convolutional Neural Network (Mask R-CNN) series can be used to extract features from the image information to be detected. Furthermore, models such as the Contrastive Language-Image Pre-Training (CLIP) model and the Bidirectional Encoder Representations from Transformers (BERT) model can be used to extract features from text information. Of course, the aforementioned algorithms can also be improved and integrated into a new model that can be used to extract features from both the image information to be detected and the text information.

[0060] The Adaptive Semantic-Spatial Joint Attention (ASSJA) mechanism is used to fuse features extracted from the image and text to be detected. By dynamically fusing the spatial distribution information of the image with the semantic description of the text, it optimizes the feature extraction process, enabling the system to accurately focus on the user-specified target in complex scenes (such as those with densely populated objects or strong background interference). The candidate regions are areas in the image that may contain objects. These regions are typically pre-extracted by an algorithm and used for subsequent object detection and category determination. The number of candidate regions can be one or multiple, depending on the actual image content. Each candidate region corresponds to a feature vector V_raw. For example, when the detection head detects the fused features, it obtains N candidate box coordinates (x, y, w, h) and N feature vectors V_raw, where N is a positive integer greater than or equal to 1.

[0061] Step S103: Encode the feature vector V_raw and text information corresponding to the candidate region respectively to obtain the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information, and calculate the similarity between the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information to obtain the similarity score Similarity corresponding to the candidate region.

[0062] Specifically, the CLIP visual encoder can be used to encode the feature vector V_raw corresponding to the candidate region to obtain the embedding vector V of the candidate region. This CLIP visual encoder can adopt a visual transformer (ViT) structure, which can include multiple transformer layers. Through linear projection, the feature vector V_raw corresponding to the candidate region is mapped to the semantic space to generate the embedding vector V. To ensure computational efficiency, the features of the ViT input are divided into patches of fixed size, and spatial information is preserved through positional encoding. In addition, the CLIP text encoder can be used to encode the text information to obtain the embedding vector T corresponding to the text information. Based on the Transformer architecture, the CLIP text encoder can convert the text sequence corresponding to the text information into the embedding vector T. During the processing, the text is segmented and special tags (such as [CLS]) are added. The global semantics are extracted through the self-attention mechanism. The output embedding vector T can represent the semantic representation of the entire description.

[0063] Next, we can calculate the similarity between the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information to obtain the similarity score of the candidate region. Specifically, we can use the following formula to calculate:

[0064] Similarity = (V · T) / (||V|| × ||T||);

[0065] Here, Similarity represents the similarity score corresponding to the candidate region, with a value range of [-1, 1]. V represents the embedding vector V corresponding to the candidate region, and T represents the embedding vector T corresponding to the text information. ||V|| and ||T|| are scalars representing the L2 norm of V and T, respectively. The Similarity score corresponding to the candidate region indicates the semantic consistency between the object and the description. To improve robustness, if the input text is empty (e.g., no user description), the system defaults to generating an embedding vector T for a general description (e.g., "any object") to ensure uninterrupted matching.

[0066] Step S104: Cluster and perform knowledge distillation on the feature vector V_raw corresponding to the candidate region to obtain the semantic embedding Linear corresponding to the candidate region, wherein the semantic embedding Linear corresponding to the candidate region is used to characterize the potential category affiliation of the candidate region, and the potential category includes the desired category and the extended category related to the desired category.

[0067] Specifically, during the clustering process, the feature vectors V_raw corresponding to the candidate regions can be clustered using algorithms such as the K-means clustering algorithm (also known as the K-means algorithm) or the Mini-Batch K-Means clustering algorithm (also known as the Mini-Batch K-Means algorithm). After clustering, knowledge distillation can be performed using the resulting cluster centers and the embedding vectors V corresponding to the candidate regions. By calculating the distillation loss between the embedding vectors V corresponding to the candidate regions and each cluster center, it is determined whether the candidate regions belong to the cluster centers obtained after clustering. If not, the cluster centers are expanded, ultimately obtaining the semantic linear embeddings corresponding to the candidate regions. These semantic linear embeddings can be used to characterize the potential category affiliation of the candidate regions. This allows for automatic discovery and expansion of target categories without the need for manual labeling by the user.

[0068] Step S105: Based on the similarity score corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region, the candidate region is classified into target categories to obtain target detection results, wherein the target detection results are used to characterize the detected target objects and the categories to which the target objects belong. The categories to which the target objects belong are expected categories, extended categories and / or unknown categories.

[0069] Specifically, the open-world classification head can integrate the similarity scores corresponding to the candidate regions and the semantic embeddings corresponding to the candidate regions to classify the candidate regions into target categories, thereby obtaining target detection results, namely the detected target objects and the categories to which the target objects belong. The categories to which the target objects belong can be any one or a combination of the expected category corresponding to the text information, an extended category related to the expected category, and an unknown category.

[0070] Through the above method, the features of the image information to be detected and the text information can be dynamically fused, so that the detected target category can be dynamically adjusted according to the description of the text information, rather than being limited to the fixed category manually labeled; and by clustering and knowledge distilling the feature vector V_raw that fuses the image information to be detected and the text information, the target category can be automatically expanded, thereby further solving the problem of the inability to break through the detection limitations of fixed categories in related technologies.

[0071] In an optional embodiment, the above step S102, performing feature extraction on the image information and text information to be detected respectively, and fusing the extracted features based on the adaptive semantic-spatial joint attention mechanism to obtain the candidate region and the feature vector V_raw corresponding to the candidate region, includes:

[0072] Feature extraction is performed on the image information to be detected and the text information respectively, and the feature map corresponding to the image information to be detected and the embedding vector T corresponding to the text information are obtained;

[0073] Based on the adaptive semantic-spatial joint attention mechanism, the spatial attention weight W_s and the semantic attention weight W_t are obtained, and feature fusion is performed based on the spatial attention weight W_s and the semantic attention weight W_t to obtain the fused feature map. The spatial attention weight W_s is obtained by performing a convolution operation on the feature map corresponding to the image information to be detected, and the semantic attention weight W_t is captured from the feature map corresponding to the image information to be detected and the embedding vector T corresponding to the text information through the multi-head self-attention mechanism;

[0074] Perform candidate region detection on the fused feature map to obtain the feature vector V_raw corresponding to the candidate region.

[0075] Specifically, the image to be detected (e.g., with a resolution of 640×640, 3 channels, and RGB format) can be extracted using the YOLOv12 backbone network (i.e., the Residual Efficient Layer Aggregation Network (R-ELAN)). This network utilizes an efficient layer aggregation design to generate a feature map corresponding to the image to be detected. The feature map corresponding to the image to be detected can be a three-layer multi-scale feature map, such as 80×80×256, 40×40×512, and 20×20×1024, corresponding to the detection requirements of small, medium, and large objects, respectively. The number of channels in each layer of the feature map (256, 512, and 1024) was experimentally optimized to ensure a balance between information representation and computational efficiency.

[0076] Subsequently, the ASSJA module can be used to process the feature map. The network structure of the ASSJA module is as follows Figure 2 As shown in , it consists of two branches: a spatial branch and a semantic branch. The spatial branch is used to perform a convolution operation on the feature map corresponding to the image information to be detected, obtaining the spatial attention weight W_s. This spatial attention weight W_s can reflect the spatial distribution characteristics of the target, such as the priority of dense or isolated areas. For example, the spatial branch can convolve the feature map with a 3×3 convolution kernel (stride 1, padding 1) to generate the spatial attention weight W_s (dimension 80×80×256). For ease of understanding, the processing process is explained here using an 80×80×256 feature map as an example. It can be expressed using the following formula:

[0077] W_s = softmax(Conv2D(Feature));

[0078] Where W_s represents the spatial attention weight, Conv2D represents a 3×3 convolution operation (stride 1, padding 1), Feature represents the feature map, and softmax is the normalization function. The above method can be applied to each feature layer in the feature map corresponding to the image information to be detected. Three layers of multi-scale feature maps can generate three corresponding spatial attention weights W_s.

[0079] The semantic branch uses a multi-head self-attention mechanism to capture semantic attention weights W_t from the feature map corresponding to the image to be detected and the embedding vector T corresponding to the text. The semantic branch receives the embedding vector T output by the CLIP text encoder (e.g., 512-dimensional, based on the user's natural language description, such as "pedestrian wearing a hat") and generates semantic attention weights W_t using a multi-head self-attention mechanism (e.g., 4 heads, each with 64 dimensions). For ease of understanding, the processing is explained using an 80×80×256 feature map as an example. It can be expressed as follows:

[0080] W_t = MultiHead(T, Feature);

[0081] Where W_t represents the semantic attention weight, MultiHead represents the multi-head self-attention function, T represents the text embedding vector, and Feature represents the feature map. The above method can be used to process each feature layer in the feature map corresponding to the image information to be detected. Three layers of multi-scale feature maps can generate three corresponding semantic attention weights W_t.

[0082] After obtaining the spatial attention weight W_s and the semantic attention weight W_t, the ASSJA module can continue to perform feature fusion based on the spatial attention weight W_s and the semantic attention weight W_t to obtain a fused feature map, and then perform candidate region detection on the fused feature map to obtain the feature vector V_raw corresponding to the candidate region.

[0083] In this embodiment, an adaptive semantic-spatial joint attention mechanism dynamically fuses the spatial information of an image with the semantic description of text, significantly improving the accuracy and flexibility of object detection. Specifically, this mechanism introduces a two-branch design during the feature extraction phase of YOLOv12: the spatial branch utilizes 3×3 convolutions to generate spatial attention weights W_s, capturing the spatial distribution characteristics of objects; the semantic branch uses multi-head self-attention to generate semantic attention weights W_t, enhancing response to user-specified objects. Feature fusion is then performed based on the spatial attention weights W_s and the semantic attention weights W_t to produce a fused feature map. Compared to traditional static attention mechanisms, the adaptive semantic-spatial joint attention mechanism can adaptively adjust weight distribution in complex scenarios (e.g., densely populated objects or with strong background interference). This overcomes the limitation of existing methods that rely on a single feature and provides a new feature optimization paradigm for open-world detection.

[0084] In an optional embodiment, the above steps of performing feature fusion according to the spatial attention weight W_s and the semantic attention weight W_t to obtain a fused feature map include:

[0085] Obtain the mean of the feature map corresponding to the image information to be detected and the variance of the embedding vector T corresponding to the text information;

[0086] Determine the value of the adaptive coefficient α based on the mean of the feature map corresponding to the image information to be detected and the variance of the embedding vector T corresponding to the text information;

[0087] Using the value of the adaptive coefficient α, the spatial attention weight W_s and the semantic attention weight W_t are fused to obtain the joint attention weight W_j;

[0088] Based on the joint attention weight W_j, the feature map corresponding to the image information to be detected is updated to obtain the fused feature map.

[0089] Specifically, when performing feature fusion based on the spatial attention weight W_s and the semantic attention weight W_t to obtain the fused feature map, we can first obtain the mean of the feature map corresponding to the image information to be detected and the variance of the embedding vector T corresponding to the text information. Then, based on the mean of the feature map corresponding to the image information to be detected and the variance of the embedding vector T corresponding to the text information, we can determine the value of the adaptive coefficient α. This can be expressed as follows:

[0090] α = MLP(mean(Feature), var(T));

[0091] Where α represents the adaptive coefficient, ranging from [0 to 1]. MLP stands for multi-layer perceptron. mean(Feature) represents the mean of the feature map. var(T) represents the variance of the embedding vector T corresponding to the text information. As an alternative embodiment, the multi-layer perceptron can be a three-layer perceptron with an input dimension of 257 and 128 hidden layers. The above method can be used to process each feature layer in the feature map corresponding to the image information to be detected. Three layers of multi-scale feature maps can yield three corresponding feature map means, mean(Feature), and three corresponding adaptive coefficients α.

[0092] Then, using the adaptive coefficient α, the spatial attention weight W_s and the semantic attention weight W_t are fused to obtain the joint attention weight W_j. Specifically, it can be expressed as follows:

[0093] W_j = α × W_s + (1 - α) × W_t;

[0094] Where W_j represents the joint attention weight, α represents the adaptation coefficient, W_t represents the semantic attention weight, and W_s represents the spatial attention weight. Here, each feature map fuses its corresponding spatial attention weight W_s and semantic attention weight W_t. Thus, three layers of multi-scale feature maps can generate three corresponding W_j.

[0095] Afterwards, the feature map corresponding to the image information to be detected can be updated based on the joint attention weight W_j to obtain the fused feature map. Specifically, it can be expressed as follows:

[0096] Feature' = Feature × W_j;

[0097] Where W_j represents the joint attention weight, Feature' represents the adjusted feature map, and Feature represents the feature map. The above method can be used to process each feature layer in the feature map corresponding to the image information to be detected. Three layers of multi-scale feature maps can be used to obtain three corresponding fused feature maps, Feature'.

[0098] Finally, the adjusted feature map Feature' is input to the detection head in YOLOv12 to generate N candidate box coordinates (x, y, w, h) and N feature vectors V_raw, where N is an integer greater than or equal to 1.

[0099] In an optional embodiment, the above step S104, clustering and knowledge distillation of the feature vector V_raw corresponding to the candidate region to obtain the semantic embedding Linear corresponding to the candidate region, includes:

[0100] Using the preset clustering algorithm and the pre-set initial cluster centers, the feature vectors V_raw corresponding to the candidate regions are clustered to obtain updated cluster centers. The number of updated cluster centers can be adaptively adjusted according to the variance of the batch features participating in each clustering.

[0101] Map the feature vector V_raw corresponding to the candidate region and the updated cluster center to the same space respectively, and calculate the distillation loss value between the candidate region and each cluster center in the updated cluster center;

[0102] Based on the distillation loss value, determine the potential category corresponding to the candidate region;

[0103] According to the potential category corresponding to the candidate region, the semantic embedding Linear corresponding to the candidate region is determined.

[0104] Specifically, the Mini-Batch K-Means algorithm can be used to cluster the feature vectors V_raw corresponding to the candidate regions. Assuming that the number of initial cluster centers K = 10, one batch is processed per frame (batch size 32), and the update formula for the dynamic update of the cluster center C_i is:

[0105] C_i = (1 - η) × C_i '+ η × mean(V_raw_batch);

[0106] Where C_i represents the updated cluster center, C_i' represents the cluster center before the update, η represents the learning rate (such as 0.01), and V_raw_batch represents the mean of the current batch features (such as dimension 256 and batch size 32).

[0107] It should be noted that the number of cluster centers K can be adaptively adjusted according to the variance of the current batch features. If the variance of the current batch features exceeds a threshold (such as 0.5), the number of cluster centers is increased by one, with an upper limit of 50. This ensures the ability to distinguish categories. Single-frame clustering takes about 2 milliseconds.

[0108] Next, the feature vector V_raw corresponding to the candidate region and the updated cluster center can be mapped to the same space, and the distillation loss between the candidate region and each of the updated cluster centers can be calculated. The specific method is: the feature vector V_raw corresponding to the candidate region is input into the CLIP visual encoder to generate a semantic embedding V (dimension 512). Then, C_i is mapped to the same space through a linear transformation layer (parameters 256×512). The distillation loss is then calculated using the following formula:

[0109] Loss_distill = ||CLIP(V_raw) - Linear(C_i)||2;

[0110] Loss_distill represents the distillation loss, a scalar. CLIP(V_raw) represents the output of the CLIP visual encoder (dimension 512). Linear(C_i) represents the linearly transformed C_i (dimension 512, transformation matrix 256×512). ||·||2 represents the L2 norm. Loss optimization uses stochastic gradient descent (learning rate 0.001), updating C_i every 10 frames to ensure semantic representation. This takes approximately 1 millisecond.

[0111] Then, based on the distillation loss value, the potential category belonging to the candidate region can be determined, and according to the potential category belonging to the candidate region, the semantic embedding Linear corresponding to the candidate region can be determined.

[0112] In this embodiment, the Unsupervised Dynamic Category Distillation (UDCD) module performs online feature clustering and knowledge distillation, enabling the automatic discovery and expansion of unknown categories. This overcomes the efficiency bottleneck of traditional methods that rely on manual labeling or retraining. UDCD expands new categories without user intervention, enabling automatic recognition of unknown objects, making it more practical in dynamic environments (such as new obstacles in autonomous driving).

[0113] In an optional embodiment, the above step of determining the potential category corresponding to the candidate region based on the distillation loss value includes:

[0114] Determine whether there is a target cluster center among the updated cluster centers, the distillation loss value between the target cluster center and the candidate region being less than a first preset threshold, wherein the target cluster center is any cluster center among the updated cluster centers;

[0115] If there is a target cluster center among the updated cluster centers whose distillation loss value with the candidate region is less than a first preset threshold, the potential category corresponding to the candidate region is assigned to the category corresponding to the target cluster center;

[0116] If there is no target cluster center in the updated cluster centers whose distillation loss value with the candidate area is less than the first preset threshold, a new cluster center is added based on the updated cluster centers, and the potential category corresponding to the candidate area is assigned to the category corresponding to the new cluster center.

[0117] Specifically, the first preset threshold can be set according to actual needs, such as 0.5, etc., and is not specifically limited here.

[0118] When determining the potential category corresponding to the candidate area based on the distillation loss value, it can be determined based on the calculated distillation loss value whether there is a target cluster center in the updated cluster center whose distillation loss value with the candidate area is less than a first preset threshold. If there is a target cluster center in the updated cluster center whose distillation loss value with the candidate area is less than the first preset threshold, the potential category corresponding to the candidate area is assigned to the category corresponding to the target cluster center; if there is no target cluster center in the updated cluster center whose distillation loss value with the candidate area is less than the first preset threshold, a new cluster center is added based on the updated cluster center, and the potential category corresponding to the candidate area is assigned to the category corresponding to the new cluster center.

[0119] It should be noted that to avoid redundancy, the distance between cluster centers is checked every 100 frames. If the distance between cluster centers is less than 0.3, similar cluster centers are merged, and the value of K is reduced after merging. Compared with traditional methods, this approach can achieve unsupervised category discovery with an automation rate of 90%, providing efficient support for subsequent identification of new categories.

[0120] In an optional embodiment, the above step S105, based on the similarity score corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region, classifies the candidate region into target categories to obtain the target detection result, including:

[0121] Compare the similarity score corresponding to the candidate region with a second preset threshold;

[0122] The target category of the candidate area whose similarity score Similarity is greater than the second preset threshold is determined as the expected category, and the similarity score Sim_C corresponding to the candidate area whose similarity score Similarity is less than or equal to the second preset threshold is calculated, wherein the similarity score Sim_C is used to represent the similarity between the semantic embedding Linear of the cluster center to which the candidate area whose similarity score Similarity is less than or equal to the second preset threshold belongs and the embedding vector T corresponding to the text information;

[0123] Comparing the similarity score Sim_C with a third preset threshold;

[0124] The target category of the candidate region whose similarity score Sim_C is greater than the third preset threshold is determined as the extended category, and the target category of the candidate region whose similarity score Sim_C is less than or equal to the third preset threshold is determined as the unknown category.

[0125] Specifically, the second preset threshold and the third preset threshold can be set according to actual needs, such as 0.7, etc., and are not specifically limited here.

[0126] Specifically, the open-world classification can be used to integrate the similarity score Similarity corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region to determine the target category. After receiving the similarity score Similarity corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region, the open-world classification can compare the similarity score Similarity corresponding to the candidate region with a second preset threshold, determine the target category of the candidate region whose similarity score Similarity is greater than the second preset threshold as the expected category, and calculate the similarity score Sim_C corresponding to the candidate region whose similarity score Similarity is less than or equal to the second preset threshold, compare the similarity score Sim_C with a third preset threshold, then determine the target category of the candidate region whose similarity score Sim_C is greater than the third preset threshold as the extended category, and determine the target category of the candidate region whose similarity score Sim_C is less than or equal to the third preset threshold as the unknown category.

[0127] For example, assuming both the second and third preset thresholds are 0.7, if the similarity score Similarity>0.7, the target category of the candidate region is determined to be a known category, and the user description (such as "pedestrian wearing a hat") is directly used as the label. If the similarity score Similarity≤0.7, the cluster center to which the candidate region belongs is determined, and then the similarity score Sim_C between the semantic embedding Linear of the cluster center and the embedding vector T corresponding to the text information is calculated. Specifically, the calculation can be performed using the following formula:

[0128] Sim_C = (Linear(C_i) · T) / (||Linear(C_i)|| × ||T||);

[0129] Among them, Sim_C represents the extended category similarity, Linear(C_i) is the semantic embedding of cluster center C_i, T is the embedding vector T corresponding to the text information, ||Linear(C_i)|| and ||T|| represent the L2 norm (scalar) of Linear(C_i) and T respectively.

[0130] Next, the similarity score Sim_C is compared with a third preset threshold of 0.7. If the similarity score Sim_C exceeds 0.7, the candidate region is determined to be an extended category, with the label C_i (e.g., "Category 3"). Otherwise, it is determined to be an unknown category. The classification results include three states: known category, extended category, and unknown category. Known categories directly map to user input, extended categories are automatically generated via UDCD, and unknown categories are retained for subsequent processing. The entire classification process takes approximately 0.5 milliseconds per frame, ensuring real-time performance. The classification logic supports parallel computing. When N=50, the total latency does not exceed 1 millisecond, providing efficient output.

[0131] Through the above method, the target category of the candidate region can be accurately classified based on the similarity score corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region.

[0132] In an optional embodiment, after performing target category classification on the candidate regions based on the similarity scores of the candidate regions and the semantic embeddings of the candidate regions in step S105 to obtain target detection results, the method further includes:

[0133] The target objects corresponding to the expected category, the target objects corresponding to the extended category, and the target objects corresponding to the unknown category are visualized and output.

[0134] Specifically, the object detection results from the open-world classification head are combined with the candidate bounding box coordinates (x, y, w, h) output by YOLOv12 to generate visual annotations. Known categories are represented by green bounding boxes with user-defined labels (e.g., "pedestrian wearing a hat") and a Similarity score (e.g., 0.85); extended categories are represented by blue bounding boxes with temporary labels generated by UDCD (e.g., "Category 3") and a Sim_C score; and unknown categories are represented by yellow bounding boxes with the text "Unknown Object." Visualization is implemented using an image rendering module that supports OpenCV or similar frameworks. Drawing the candidate bounding boxes takes approximately 0.3 milliseconds. Results are displayed through a user interface consisting of an image area (displaying the annotation results) and a text input box (receiving a description). If the object is of an unknown category, the system can optionally prompt the user to enter a supplementary description (e.g., "What is this"), but manual intervention is not mandatory. The output resolution matches the input (640×640), supports real-time streaming, and takes approximately 2 milliseconds per frame. For easy verification, the resulting data can be stored in JSON format, including candidate bounding box coordinates, category labels, and similarity scores, for subsequent analysis. The intuitive output interface ensures that users or detection devices can clearly determine the target status at a glance, meeting practical application needs.

[0135] For ease of understanding, this application example uses an intelligent monitoring scenario as an example to specifically illustrate how this application is implemented in practice, demonstrating how to achieve dynamic target screening and unknown category expansion through the adaptive semantic-spatial joint attention mechanism ASSJA and the unsupervised dynamic category distillation module UDCD. Specifically, the following steps may be included:

[0136] 1. Experimental Preparation

[0137] Hardware environment: NVIDIA RTX 3090 GPU (24GB video memory), Intel i9-12900K CPU, 64GB memory, running Ubuntu 20.04 system.

[0138] Software environment: Based on the PyTorch 2.0 framework, pre-installed OpenCV 4.5.5 for image processing and FAISS 1.7.2 for fast clustering.

[0139] Input data: Video stream captured by a surveillance camera, with a resolution of 640×640 and a frame rate of 30 fps. The scene is a shopping mall entrance and contains common objects such as pedestrians and shopping bags.

[0140] Pre-trained model: Use the pre-trained weights of YOLOv12 (based on the COCO dataset) and the pre-trained model of CLIP (ViT-B / 32 version).

[0141] 2. Implementation steps

[0142] Step 1: System initialization and data input

[0143] Specific steps: Start the system, load the YOLOv12 and CLIP pre-trained models into GPU memory, and initialize the cluster centers of the UDCD module (K=10, dimension 256). Real-time video streams are collected from surveillance cameras, and frame data is input into the system in RGB format (640×640×3).

[0144] Technical Specifications: Image preprocessing includes normalization (pixel value range [0, 1]) and mean subtraction (mean [0.485, 0.456, 0.406], standard deviation [0.229, 0.224, 0.225]). Initialization takes approximately 5 seconds.

[0145] Step 2: Feature extraction and ASSJA processing

[0146] Specific steps: Each frame of image is input into the YOLOv12 detection module, and the backbone network (R-ELAN) generates three layers of feature maps (80×80×256, 40×40×512, and 20×20×1024). The user enters a text description of "pedestrians in red clothes" through the interface, and the CLIP text encoder (12-layer Transformer, dimension 512) generates an embedding vector T (dimension 512). The ASSJA module receives the feature map and embedding vector T. The spatial branch calculates the spatial attention weight W_s through 3×3 convolution (stride 1, padding 1), and the semantic branch calculates the semantic attention weight W_t through multi-head self-attention (4 heads, each head dimension 64). The adaptive coefficient α is generated by the MLP (input dimension 257, hidden layer 128). The formula for feature fusion design is:

[0147] W_s=softmax(Conv2D(Feature));

[0148] W_t=MultiHead(T,Feature);

[0149] α=MLP(mean(Feature),var(T));

[0150] W_j=α×W_s+(1-α)×W_t;

[0151] Feature'=Feature×W_j;

[0152] Since the parameters and meanings of the formulas have been explained in the previous text, they will not be repeated here.

[0153] Technical parameters: Convolution kernel weights are initialized to a Xavier distribution, the multi-head attention dropout rate is 0.1, and the MLP activation function is ReLU. The detection head outputs 50 candidate boxes (x, y, w, h) and a feature vector V_raw (dimension 256), taking 1.2 milliseconds.

[0154] Step 3: Image and text matching

[0155] Specific actions: Input the feature vector V_raw into the CLIP visual encoder (ViT, 12 layers, 8 heads) to generate an embedding vector V (dimension 512). Then calculate the cosine similarity between the embedding vector V and the embedding vector T, the calculation formula is as follows:

[0156] Similarity=(V·T) / (||V||×||T||);

[0157] Since the various parameters in the formula and the meaning of the formula have been explained in the previous text, they will not be repeated here.

[0158] This will output 50 similarity scores. If there is no text input, the default embedding vector T is the predefined embedding of "any target".

[0159] Technical parameters: ViT patch size is 16×16, position encoding uses a sine function, and similarity calculation takes 0.5 milliseconds.

[0160] Step 4: Unsupervised Category Distillation

[0161] Specific actions: The UDCD module receives the feature vector V_raw and uses Mini-Batch K-Means (batch size 32, η = 0.01) to update the cluster centers C_i (dimension 256). The update process is expressed as follows:

[0162] C_i=(1-η)×C_i+η×mean(V_raw_batch);

[0163] Knowledge distillation maps C_i to CLIP space (Linear layer, 256×512), and then calculates the distillation loss. The calculation formula is as follows:

[0164] Loss_distill=||CLIP(V_raw)-Linear(C_i)||2;

[0165] Since the parameters and meanings of the formulas have been explained in the previous text, they will not be repeated here.

[0166] Technical parameters: K capped at 50, Stochastic Gradient Descent (SGD) optimizer, learning rate 0.001, distillation update frequency 10 frames, and time taken 3 milliseconds.

[0167] Step 5: Open World Classification Head

[0168] Specific actions: The open world classification head receives the similarity score and the semantic embedding Linear output by UDCD. If the similarity score is > 0.7, it is a known category; otherwise, the similarity between Linear (C_i) and the embedding vector T is calculated using the following formula:

[0169] Sim_C=(Linear(C_i)·T) / (||Linear(C_i)||×||T||);

[0170] If the similarity score Sim_C>0.7, it is an extended category; otherwise, it is an unknown category.

[0171] Technical parameters: Classification takes 0.5 milliseconds.

[0172] Step 6: Output

[0173] Specific actions: Draw a green box (known, such as "pedestrian in red clothes"), a blue box (extended, such as "category 5"), and a yellow box (unknown) on the image, render it using OpenCV, and save it to a JSON file (containing coordinates, labels, and scores).

[0174] Technical parameters: Rendering resolution 640×640, time taken 0.3 milliseconds, total time taken per frame is about 2 milliseconds.

[0175] 3. Implementation Results

[0176] Output example: When the input is "pedestrian in red clothes", the system detects 3 targets (similarity scores are 0.89, 0.85, and 0.72 respectively), marked as green boxes; detects 2 new targets (similarity scores Sim_C=0.75 and 0.71), marked as blue boxes (extended category); and 1 target (Similarity=0.45) is marked as a yellow box (unknown).

[0177] Performance indicators: Average precision reached 42.5, inference time was 2 milliseconds, category expansion took 5 milliseconds, and automation rate was 90%.

[0178] It can be seen that this application achieves three major advantages in new category detection through ASSJA and UDCD. First, in terms of high-precision interaction, ASSJA dynamically integrates spatial and semantic information, enabling the system to accurately locate targets based on natural language descriptions input by users (such as "pedestrians wearing hats"), improving detection accuracy in complex scenarios (such as dense crowds or changes in lighting) while ensuring real-time performance. Second, in terms of automated category expansion, UDCD achieves rapid induction of new categories (such as "new hats") through unsupervised clustering and distillation, enabling rapid identification of unknown targets. Finally, in practical applications, this application takes into account both real-time and openness, such as quickly screening specific targets in intelligent monitoring and automatically identifying new obstacles in autonomous driving. It breaks through the fixed category limitations of traditional detection and provides an efficient and flexible solution with broad market potential and technical competitiveness.

[0179] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of a target detection device provided in an embodiment of the present application. Figure 3 As shown, the target detection device 300 includes:

[0180] An acquisition module 301 is used to acquire image information and text information to be detected, wherein the text information is used to describe the expected category of the target object to be detected;

[0181] The feature extraction and fusion module 302 is used to extract features from the image information and text information to be detected respectively, and fuse the extracted features based on the adaptive semantic-spatial joint attention mechanism to obtain the candidate region and the feature vector V_raw corresponding to the candidate region;

[0182] The encoding module 303 is used to encode the feature vector V_raw and the text information corresponding to the candidate region respectively, obtain the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information, and calculate the similarity between the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information to obtain the similarity score Similarity corresponding to the candidate region;

[0183] Clustering and knowledge distillation module 304 is used to cluster and perform knowledge distillation on the feature vector V_raw corresponding to the candidate region to obtain a semantic embedding linear corresponding to the candidate region, wherein the semantic embedding linear corresponding to the candidate region is used to represent the potential category to which the candidate region belongs, and the potential category includes an extended category related to the desired category;

[0184] The target category classification module 305 is used to classify the candidate area into target categories based on the similarity score corresponding to the candidate area and the semantic embedding Linear corresponding to the candidate area, and obtain the target detection result, wherein the target detection result is used to characterize the detected target object and the category to which the target object belongs. The category to which the target object belongs is the expected category, the extended category and / or the unknown category.

[0185] Furthermore, the feature extraction and fusion module 302 includes:

[0186] The feature extraction submodule is used to extract features of the image information to be detected and the text information respectively, and obtain the feature map corresponding to the image information to be detected and the embedding vector T corresponding to the text information;

[0187] The feature fusion submodule is used to obtain the spatial attention weight W_s and the semantic attention weight W_t based on the adaptive semantic-spatial joint attention mechanism, and perform feature fusion based on the spatial attention weight W_s and the semantic attention weight W_t to obtain the fused feature map. The spatial attention weight W_s is obtained by performing a convolution operation on the feature map corresponding to the image information to be detected, and the semantic attention weight W_t is captured from the feature map corresponding to the image information to be detected and the embedding vector T corresponding to the text information through the multi-head self-attention mechanism;

[0188] The detection submodule is used to detect candidate regions on the fused feature map and obtain the feature vector V_raw corresponding to the candidate region.

[0189] Furthermore, the feature fusion submodule includes:

[0190] An acquisition unit, configured to acquire the mean of the feature map corresponding to the image information to be detected and the variance of the embedding vector T corresponding to the text information;

[0191] a determination unit, configured to determine a value of an adaptive coefficient α based on a mean value of a feature map corresponding to the image information to be detected and a variance of an embedding vector T corresponding to the text information;

[0192] The fusion unit is used to fuse the spatial attention weight W_s and the semantic attention weight W_t using the value of the adaptive coefficient α to obtain the joint attention weight W_j;

[0193] The updating unit is used to update the feature map corresponding to the image information to be detected based on the joint attention weight W_j to obtain the fused feature map.

[0194] Furthermore, the clustering and knowledge distillation module 304 includes:

[0195] The clustering submodule is used to cluster the feature vectors V_raw corresponding to the candidate regions using a preset clustering algorithm and pre-set initial cluster centers to obtain updated cluster centers. The number of updated cluster centers can be adaptively adjusted according to the variance of the batch features participating in each clustering.

[0196] The distillation submodule is used to map the feature vector V_raw corresponding to the candidate region and the updated cluster center to the same space, and calculate the distillation loss value between the candidate region and each cluster center in the updated cluster center;

[0197] The first determination submodule is used to determine the potential category corresponding to the candidate region based on the distillation loss value;

[0198] The second determination submodule is used to determine the semantic embedding Linear corresponding to the candidate region according to the potential category belonging to the candidate region.

[0199] Furthermore, the first determining submodule includes:

[0200] a judging unit, configured to judge whether there is a target cluster center among the updated cluster centers, the distillation loss value between the target cluster center and the candidate region being less than a first preset threshold, wherein the target cluster center is any cluster center among the updated cluster centers;

[0201] A first assigning unit is configured to assign the potential category corresponding to the candidate region to the category corresponding to the target cluster center when there is a target cluster center among the updated cluster centers whose distillation loss value with the candidate region is less than a first preset threshold;

[0202] The second attribution unit is used to add a new cluster center based on the updated cluster center when there is no target cluster center in the updated cluster center whose distillation loss value with the candidate area is less than the first preset threshold, and to attribute the potential category corresponding to the candidate area to the category corresponding to the new cluster center.

[0203] Furthermore, the target category classification module 305 includes:

[0204] A first comparison submodule is used to compare the similarity score Similarity corresponding to the candidate region with a second preset threshold;

[0205] A calculation submodule is used to determine the target category of the candidate area whose similarity score Similarity is greater than the second preset threshold as the expected category, and calculate the similarity score Sim_C corresponding to the candidate area whose similarity score Similarity is less than or equal to the second preset threshold, wherein the similarity score Sim_C is used to represent the similarity between the semantic embedding Linear of the cluster center to which the candidate area whose similarity score Similarity is less than or equal to the second preset threshold belongs and the embedding vector T corresponding to the text information;

[0206] A second comparison submodule, configured to compare the similarity score Sim_C with a third preset threshold;

[0207] The third determination submodule is configured to determine the target category of the candidate area whose similarity score Sim_C is greater than a third preset threshold as an extended category, and determine the target category of the candidate area whose similarity score Sim_C is less than or equal to the third preset threshold as an unknown category.

[0208] Furthermore, the target detection device 300 further includes:

[0209] The output module is used to visualize the target objects corresponding to the expected category, the target objects corresponding to the extended category, and the target objects corresponding to the unknown category.

[0210] It should be noted that the target detection device 300 can implement the steps of the target detection method provided by any of the aforementioned method embodiments and can achieve the same technical effects, which will not be described in detail here.

[0211] like Figure 4 As shown, the embodiment of the present application further provides an electronic device, including a processor 411, a communication interface 412, a memory 413 and a communication bus 414, wherein the processor 411, the communication interface 412, and the memory 413 communicate with each other through the communication bus 414.

[0212] Memory 413, for storing computer programs;

[0213] In one embodiment of the present application, the processor 411 is configured to implement the target detection method provided by any one of the aforementioned method embodiments when executing a program stored in the memory 413 .

[0214] The electronic devices provided in the embodiments of the present application may specifically be modules capable of implementing communication functions or terminal devices containing such modules, and the terminal devices may be mobile terminals or smart terminals. A mobile terminal may specifically be at least one of a mobile phone, a tablet computer, and a laptop computer; a smart terminal may specifically be a terminal containing a wireless communication module, such as a smart car, a smart watch, a shared bicycle, or a smart cabinet; and a module may specifically be a wireless communication module, such as any one of a 2G communication module, a 3G communication module, a 4G communication module, a 5G communication module, and an NB-IOT communication module.

[0215] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the target detection method provided in any of the aforementioned method embodiments is implemented.

[0216] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0217] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.

Claims

1. A target detection method, characterized in that: The method comprises: Acquire image information and text information to be detected, wherein the text information is used to describe the expected category of the target object to be detected; Feature extraction is performed on the image information to be detected and the text information respectively, and the extracted features are fused based on the adaptive semantic-spatial joint attention mechanism to obtain a candidate region and a feature vector V_raw corresponding to the candidate region; Encode the feature vector V_raw and the text information corresponding to the candidate region respectively to obtain the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information, and calculate the similarity between the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information to obtain the similarity score Similarity corresponding to the candidate region; Performing clustering and knowledge distillation on the feature vector V_raw corresponding to the candidate region to obtain a semantic embedding Linear corresponding to the candidate region, wherein the semantic embedding Linear corresponding to the candidate region is used to characterize the potential category to which the candidate region belongs, and the potential category includes the desired category and an extended category related to the desired category; Based on the similarity score corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region, the candidate region is classified into target categories to obtain a target detection result, wherein the target detection result is used to characterize the detected target object and the category to which the target object belongs, and the category to which the target object belongs is the expected category, the extended category and / or the unknown category; The feature extraction of the image information to be detected and the text information is performed respectively, and the extracted features are fused based on the adaptive semantic-spatial joint attention mechanism to obtain the candidate region and the feature vector V_raw corresponding to the candidate region, including: Performing feature extraction on the image information to be detected and the text information respectively to obtain a feature map corresponding to the image information to be detected and an embedding vector T corresponding to the text information; Based on the adaptive semantic-spatial joint attention mechanism, the spatial attention weight W_s and the semantic attention weight W_t are obtained, and feature fusion is performed according to the spatial attention weight W_s and the semantic attention weight W_t to obtain a fused feature map, wherein the spatial attention weight W_s is obtained by performing a convolution operation on the feature map corresponding to the image information to be detected, and the semantic attention weight W_t is captured from the feature map corresponding to the image information to be detected and the embedding vector T corresponding to the text information through a multi-head self-attention mechanism; Perform candidate region detection on the fused feature map to obtain a feature vector V_raw corresponding to the candidate region.

2. The target detection method according to claim 1, wherein: The feature fusion is performed according to the spatial attention weight W_s and the semantic attention weight W_t to obtain a fused feature map, including: Obtaining the mean of the feature map corresponding to the image information to be detected and the variance of the embedding vector T corresponding to the text information; Determining a value of an adaptive coefficient α based on a mean value of a feature map corresponding to the image information to be detected and a variance of an embedding vector T corresponding to the text information; Using the value of the adaptive coefficient α, the spatial attention weight W_s and the semantic attention weight W_t are fused to obtain a joint attention weight W_j; Based on the joint attention weight W_j, the feature map corresponding to the image information to be detected is updated to obtain the fused feature map.

3. The target detection method according to claim 1, wherein: The clustering and knowledge distillation of the feature vector V_raw corresponding to the candidate region to obtain the semantic embedding Linear corresponding to the candidate region includes: Using a preset clustering algorithm and pre-set initial cluster centers, cluster the feature vectors V_raw corresponding to the candidate regions to obtain updated cluster centers, wherein the number of the updated cluster centers can be adaptively adjusted according to the variance of the batch features participating in each clustering; Mapping the feature vector V_raw corresponding to the candidate region and the updated cluster center to the same space, and calculating the distillation loss value between the candidate region and each cluster center in the updated cluster center; Determining a potential category corresponding to the candidate region based on the distillation loss value; According to the potential category belonging to the candidate region, a semantic embedding Linear corresponding to the candidate region is determined.

4. The target detection method according to claim 3, wherein: The determining, based on the distillation loss value, the potential category corresponding to the candidate region includes: Determine whether there is a target cluster center among the updated cluster centers, the distillation loss value between which the target cluster center and the candidate region is less than a first preset threshold, wherein the target cluster center is any cluster center among the updated cluster centers; When there is a target cluster center among the updated cluster centers whose distillation loss value with the candidate region is less than the first preset threshold, assigning the potential category corresponding to the candidate region to the category corresponding to the target cluster center; If there is no target cluster center in the updated cluster centers whose distillation loss value with the candidate area is less than the first preset threshold, a new cluster center is added based on the updated cluster centers, and the potential category corresponding to the candidate area is assigned to the category corresponding to the new cluster center.

5. The target detection method according to claim 1, wherein: The target detection result is obtained by classifying the candidate region into target categories based on the similarity score corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region, including: Compare the similarity score Similarity corresponding to the candidate region with a second preset threshold; Determine the target category of the candidate area whose similarity score Similarity is greater than the second preset threshold as the expected category, and calculate the similarity score Sim_C corresponding to the candidate area whose similarity score Similarity is less than or equal to the second preset threshold, wherein the similarity score Sim_C is used to represent the similarity between the semantic embedding Linear of the cluster center to which the candidate area whose similarity score Similarity is less than or equal to the second preset threshold belongs and the embedding vector T corresponding to the text information; Comparing the similarity score Sim_C with a third preset threshold; The target category of the candidate area whose similarity score Sim_C is greater than the third preset threshold is determined as the extended category, and the target category of the candidate area whose similarity score Sim_C is less than or equal to the third preset threshold is determined as the unknown category.

6. The target detection method according to claim 1, wherein: After classifying the candidate region into target categories based on the similarity scores corresponding to the candidate regions and the semantic embedding Linear corresponding to the candidate regions to obtain target detection results, the method further includes: The target object corresponding to the expected category, the target object corresponding to the extended category, and the target object corresponding to the unknown category are visually output.

7. A target detection device, characterized in that: The device comprises: An acquisition module, configured to acquire image information and text information to be detected, wherein the text information is used to describe the expected category of the target object to be detected; A feature extraction and fusion module is used to extract features from the image information to be detected and the text information respectively, and fuse the extracted features based on an adaptive semantic-spatial joint attention mechanism to obtain a candidate region and a feature vector V_raw corresponding to the candidate region; An encoding module is used to encode the feature vector V_raw corresponding to the candidate region and the text information respectively, obtain the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information, and calculate the similarity between the embedding vector V corresponding to the candidate region and the embedding vector T corresponding to the text information to obtain the similarity score Similarity corresponding to the candidate region; A clustering and knowledge distillation module is used to perform clustering and knowledge distillation on the feature vector V_raw corresponding to the candidate region to obtain a semantic embedding Linear corresponding to the candidate region, wherein the semantic embedding Linear corresponding to the candidate region is used to characterize the potential category to which the candidate region belongs, and the potential category includes an extended category related to the desired category; A target category classification module is used to classify the candidate region into target categories based on the similarity score corresponding to the candidate region and the semantic embedding Linear corresponding to the candidate region, thereby obtaining a target detection result, wherein the target detection result is used to characterize the detected target object and the category to which the target object belongs, and the category to which the target object belongs is the expected category, the extended category, and / or the unknown category; Wherein, the feature extraction and fusion module includes: A feature extraction submodule is used to extract features from the image information to be detected and the text information respectively, to obtain a feature map corresponding to the image information to be detected and an embedding vector T corresponding to the text information; A feature fusion submodule is used to obtain the spatial attention weight W_s and the semantic attention weight W_t based on the adaptive semantic-spatial joint attention mechanism, and perform feature fusion according to the spatial attention weight W_s and the semantic attention weight W_t to obtain a fused feature map, wherein the spatial attention weight W_s is obtained by performing a convolution operation on the feature map corresponding to the image information to be detected, and the semantic attention weight W_t is captured from the feature map corresponding to the image information to be detected and the embedding vector T corresponding to the text information through a multi-head self-attention mechanism; The detection submodule is used to perform candidate region detection on the fused feature map to obtain a feature vector V_raw corresponding to the candidate region.

8. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; The processor is configured to implement the target detection method according to any one of claims 1 to 6 when executing a program stored in the memory.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target detection method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Remote sensing small target detection method for model lightweight design

    CN112329721A

  • Knowledge distillation-based target detection model training method

    CN119295739A