Semi-supervised camouflage target detection method based on adaptive data selection and text fusion
The method addresses the issue of random data selection in semi-supervised camouflage target detection by using adaptive data selection and text fusion to enhance the training process, achieving improved performance in camouflage target detection.
Patent Information
- Application Number
- CN202510788597.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing semi-supervised camouflage object detection method selects data as labeled data through random sampling, and does not consider the data quality, which makes it difficult for the model to learn the deep characteristics of the camouflage object, lacks sufficient mining of labeled data information, and has poor performance.
Adaptive data enhancement module and adaptive data selection module are adopted to build high-value data sets by rating and manually labeling the label-free images, and combining image-level directive text, model training is used using the teacher-student framework and text fusion module to generate high-quality segmentation masks.
The precise segmentation of the camouflage targets under a small amount of manual labeling data is achieved, which surpasses the performance of the existing methods and improves the detection accuracy and segmentation effect of the model.
Smart Images

Figure CN120318661A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, specifically the field of semi-supervised camouflaged object detection, and particularly refers to a semi-supervised camouflaged object detection method based on adaptive data selection and text fusion. Background Art
[0002] Camouflaged Object Detection (COD) is a highly challenging task. Its core objective is to train a model using a large amount of pixel-level annotated data to learn how to accurately detect and segment camouflaged objects that have a high degree of similarity to the background. Compared with traditional salient object detection, camouflaged object detection faces more complex challenges: Camouflaged objects usually have characteristics that interfere with detection and localization, such as small size, partial occlusion, high concealment of blending with the background, and even the ability of self-camouflage to dynamically change their appearance to enhance concealment. This high similarity in color, texture, and morphology between the object and the background significantly increases the difficulty of camouflaged object detection. Since segmenting and annotating camouflaged object images is a relatively difficult task, studying how to efficiently train a camouflaged object detection model using a small amount of data has great application value.
[0003] Compared with traditional fully supervised camouflaged object detection methods, semi-supervised camouflaged object detection trains a model to segment camouflaged objects by using a small amount of labeled data and a large amount of unlabeled data, effectively alleviating the problem of dependence on labeled data. Most existing semi-supervised camouflaged object detection methods adopt a teacher-student network framework and train the model by designing data augmentation strategies or constructing regularization loss terms to alleviate the noise problem of pseudo-labels in the teacher model. However, these methods do not consider the impact of the quality of labeled data on model training and only randomly sample a certain proportion of data as labeled data; in addition, these methods lack sufficient exploration of the information in labeled data, resulting in poor performance.
[0004] The directed object detection task further refines the scope of the object detection field. This task focuses on identifying and locating specific objects in an image based on specific cues or descriptions. Such methods not only need to analyze the visual content of the image but also understand the descriptions or instructions related to the object to accurately lock the object in complex scenes. Directed camouflaged object detection combines the research results of two fields: camouflaged object detection and directed object detection. Existing image-based directed object detection methods capture the common representation of target objects from reference images composed of salient objects by using the reference images of camouflaged objects as directed information for identifying specified camouflaged objects. In addition, some methods attempt to use the way of visual language model question answering. By designing a series of question and answer prompt templates, they guide the visual language model to locate camouflaged objects, so as to obtain relevant descriptions. The performance of this method is limited by the ability of the visual language model to identify specified camouflaged objects from images and accurately describe their attributes. If the visual language model itself cannot locate the specified camouflaged object in the image, or the number of located camouflaged objects is incorrect, it will have a great impact on the accuracy of the final prediction of the model. Therefore, accurate and appropriate descriptions of target images play a crucial role in the quality of training models.
[0005] In summary, in the prior art, most existing semi-supervised camouflaged object detection methods select data as labeled data by random sampling, without considering the quality of the selected data; for the small amount of selected labeled data, it is difficult for the model to learn knowledge related to camouflage and lacks sufficient mining of the information of labeled data, resulting in poor performance. Summary of the Invention
[0006] The main purpose of the present invention is to provide a semi-supervised camouflaged object detection method based on adaptive data selection and text fusion, to solve the problems existing in the prior art. By using adaptive data augmentation technology and text fusion strategies, model training is carried out under the condition of only requiring a small amount of manually labeled data, guiding the model to learn the deep features of camouflaged objects, realizing accurate segmentation of camouflaged objects that are very similar to the background, and generating high-quality segmentation masks.
[0007] To achieve the above object, the solution of the present invention is: A semi-supervised camouflaged object detection method based on adaptive data selection and text fusion, which includes an adaptive data augmentation module, an adaptive data selection module and a text fusion module, and includes the following steps: Step 1: Use the adaptive data selection module to score the value of unlabeled images, select high-value pictures suitable for model training according to a preset scoring threshold for manual annotation, and construct a labeled data set containing image-level directed text and real image labels; Step 2: First, preprocess the manually annotated images to be detected, and then use the adaptive data augmentation module to obtain adaptively data-augmented images; Step 3: Input the adaptively data-augmented images into the teacher-student framework, and extract the image depth features by the visual feature extraction backbone network; the teacher-student framework includes a teacher model and a student model; Step 4: Input the image depth features and the image-level directional text into the text fusion module to generate text-visual fusion features, and generate the teacher model prediction result and the student model prediction result by the decoder; Step 5: Calculate the total loss of the entire network using the real image labels and the student model prediction results, and then adjust the parameters of the visual feature extraction backbone network, the adaptive data augmentation module, the text fusion module, and the decoder through backpropagation.
[0008] Step 1 includes a construction step of a camouflaged object detection dataset RefTextCOD containing image-level directional text: Step 1.1.1: Perform directional text annotation on four mainstream camouflaged object detection datasets, namely CHAMELEON, CAMO, COD10K, and NC4K, and obtain the precise foreground area of the camouflaged object in the input image according to the dataset ground truth map; Step 1.1.2: Construct corresponding graph-prompt pairs according to the preset prompt rules; Step 1.1.3: Input the original image, the foreground area of the camouflaged object, and the prompt words into the Qwen-VL-Chat model and the GPT4-Vision model respectively to obtain multiple groups of candidate prompt words for each group of images; Step 1.1.4: Screen and correct the candidate prompt words in an artificial trimming manner to obtain precise directional text; Step 1.1.5: Organize the results into the format of a complete dataset image-text pair.
[0009] Preferably, the preset prompt rules are: by given the original image, and using the corresponding segmentation mask to extract the camouflaged object as an auxiliary positioning image, guiding the model to gradually identify the foreground of the camouflaged object and its corresponding background physical characteristics through prompt words, and guiding the VLM to aggregate this information to generate the final annotation.
[0010] In Step 1, an adaptive data selection module is used to incrementally expand the annotated dataset, specifically including: Step 1.2.1 For the model trained for the first time, the adaptive data selection module extracts features in the color, texture, and frequency domains, classifies the data using K-Means clustering, selects the samples close to the cluster centers as the initial labeled data, and proceeds to Step 2; for the model trained again, based on the structural similarity index and the mean absolute error, calculate the data scores between the predicted results of the teacher model after all unlabeled images are input into the model and the predicted results of the student model : ; where represents the index of the unlabeled image; and represent the structural similarity index and the mean absolute error respectively; the subscript refers to unlabeled, the superscript refers to image, the subscript refers to teacher, and the subscript refers to student; Step 1.2.2 Normalize the data scores of all unlabeled images, scale the data scores to the numerical interval [0, 1] using min-max normalization, and sort the data according to the normalization results; Step 1.2.3 Select data symmetrically starting from the middle position of the sorted results, manually label the selected data, and add it to the labeled dataset.
[0011] In the said Step 2, scale the original pixel values of the manually labeled images to be detected to the interval [0, 1], and normalize them using the preset channel means and variances; then scale the pictures to a size of 640×640, and the semantic segmentation masks corresponding to the images are also scaled to a size of 640×640; use a trainable color enhancer and a trainable geometric transformation enhancer to perform data augmentation on the scaled images.
[0012] In the said Step 3, both the teacher model and the student model of the teacher-student framework are composed of a visual feature extraction backbone network, a text fusion module, and a decoder, and the exponential moving average is performed on the student model parameters to generate the teacher model parameters; the visual feature extraction backbone network uses the Swin pre-trained model to extract the image depth features; the decoder uses the BiRefBlock decoder; the visual feature extraction backbone network, the text fusion module, and the decoder all participate in the model training.
[0013] The process of fusing the image depth feature with the image-level directional text input text into the text-visual fusion module in step 4 specifically includes the following steps: Step 4.1: Encode the image-level directional text using a directional text encoder to obtain a text feature ; The directional text encoder uses a pre-trained CLIP text encoder with frozen parameters, which does not participate in model training and is only used to encode text information; The superscript represents Text; The subscript represents Labeled, and the subscript represents the index of labeled data; Step 4.2: For the image depth feature , generate a clue attention query using a clue attention mechanism : ; ; ; Among them, the subscript represents the th feature in the multi-level feature; represents the clue query vector; represents the learnable query; represents the clue key vector; represents the clue value vector; , and represent the linear projection weights; represents the clue attention map; represents transpose; represents the dimension of the attention head; represents the softmax activation function; Step 4.3: Calculate the cosine similarity score between the clue attention query and the codebook : ; Among them, represents the cosine similarity; represents the element in the codebook; represents a codebook of a set of feature vectors storing category-related information; Step 4.4: Based on the cosine similarity score , perform weighted aggregation on the codebook to generate a clue vector : ; Among them, Denotes calculating the number of elements in a set; Step 4.5 Insert the clue vector into the text feature to obtain an enhanced text feature : ; Wherein, Denotes a concatenation operation; Step 4.6 Use a cross-attention mechanism for the enhanced text feature and the image depth feature to obtain a text-visual fusion feature : ; Wherein, Denotes cross-attention; Step 4.7 Input the text-visual fusion feature into the decoder to generate the prediction result of the teacher model and the prediction result of the student model .
[0014] Preferably, in the said step 5, the total loss of the network is calculated as: ; ; ; ; ; ; ; Wherein, Denotes the supervised loss, used for supervising the prediction result of the student model by the true image label; Denotes a hyperparameter, used for balancing the supervised loss and the unsupervised loss; Denotes the unsupervised loss, used for supervising the prediction result of the student model by the prediction result of the teacher model; Denotes the supervised segmentation loss, Denotes the unsupervised segmentation loss, both of which are composed of binary cross-entropy loss, intersection over union loss, and structural similarity loss; Denotes the adaptive enhancer loss; Denotes the binary cross-entropy loss; , Denote the prediction results of the teacher model and the student model respectively after the labeled image is input into the model; Denotes the true annotation of the image; Represents the loss of the text fusion module; Represents the clue attention map; Represents adjusting the resolution using the bilinear interpolation algorithm; Represents the mean squared error loss; Represents the CLIP pre-trained text encoder; Represents the class word of the image.
[0015] After adopting the above technical solutions, the present invention has the following technical effects: (1) The present invention has an adaptive data augmentation module and an adaptive data selection module, which adaptively select valuable samples for annotation through adversarial augmentation and sampling strategies; (2) The present invention proposes a text fusion module, which realizes the full utilization of labeled data by combining camouflage knowledge and visual-text interaction; (3) The present invention proposes a new dataset, called RefTextCOD, which contains image-level text annotations for existing mainstream camouflage object detection datasets; (4) Referring to the settings of existing semi-supervised camouflage object detection for simulation experiments, the experimental results show that the method proposed by the present invention surpasses all existing semi-supervised camouflage object detection methods, effectively proving its effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Is the overall flowchart of the method proposed in the specific embodiment of the present invention.
[0017] Figure 2 Is the predefined prompt word construction rule in the specific embodiment of the present invention.
[0018] Figure 3 Is the specific flowchart of dataset construction in the specific embodiment of the present invention.
[0019] Figure 4 Is the performance comparison experimental result between the specific embodiment of the present invention and the mainstream camouflage object detection method.
[0020] Figure 5 Is the ablation verification of the module effectiveness in the specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to further explain the technical solutions of the present invention, the present invention will be elaborated in detail through specific embodiments below.
[0022] Refer to Figure 1As shown in the figure, the present invention discloses a semi-supervised camouflaged target detection method based on adaptive data selection and text fusion, which includes an Adaptive Data Augmentation (ADA) module, an Adaptive Data Selection (ADS) module, and a Text Fusion Module (TFM), and the method comprises the following steps: Step 1: Use the adaptive data selection module to perform value scoring on unannotated images, select high-value pictures suitable for model training according to a preset scoring threshold for manual annotation, and construct an annotated dataset containing image-level directional text and real image labels; Step 2: First, preprocess the manually annotated image to be detected, and then use the adaptive data augmentation module to obtain an adaptively data-augmented image; Step 3: Input the adaptively data-augmented image into a teacher-student framework (including a teacher model and a student model), and extract the image depth features by a visual feature extraction backbone network; Step 4: Input the image depth features and the image-level directional text into the text fusion module to generate text-visual fusion features, and generate the prediction results of the teacher model and the student model by a decoder; Step 5: Calculate the total loss of the entire network (the complete network including all components) using the real image labels and the prediction results of the student model, and then adjust the parameters of the visual feature extraction backbone network, the adaptive data augmentation module, the text fusion module, and the decoder through backpropagation.
[0023] Reference Figures 2-3 As shown in the figure, the above Step 1 includes a construction step of a camouflaged target detection dataset RefTextCOD containing image-level directional text: Step 1.1.1: Perform directional text annotation on four mainstream camouflaged target detection datasets, namely CHAMELEON, CAMO, COD10K, and NC4K, and obtain the precise foreground area of the camouflaged target in the input image according to the ground truth map of the dataset; Step 1.1.2: Construct corresponding graphic-prompt pairs according to the preset prompt word rules (see Figure 2 ); where "{}" represents the answer result according to the previous prompt word, replacing the corresponding position; Step 1.1.3: Input the original image, the foreground area of the camouflaged target, and the prompt words into the Qwen-VL-Chat model and the GPT4-Vision model respectively to obtain multiple groups of candidate prompt words for each group of images; Step 1.1.4. Screen and correct the candidate prompt words manually to obtain accurate directional text; Step 1.1.5. Organize the results into the format of a complete dataset of image-text pairs.
[0024] Furthermore, the above preset prompt word rules are as follows: By given the original image and using the corresponding segmentation mask to extract the camouflaged target as an auxiliary positioning image, and guiding the model to gradually identify the foreground of the camouflaged target and its corresponding background physical characteristics through prompt words, and guiding the VLM to aggregate this information to generate the final annotation.
[0025] In the above Step 1, an adaptive data selection module is used to incrementally expand the annotated dataset, which specifically includes: Step 1.2.1. For the model trained for the first time, the adaptive data selection module extracts features such as color, texture, and frequency domain, and uses K-Means clustering to classify the data, selects the samples close to the cluster center as the initial annotated data, and enters Step 2; for the model trained again, based on the structural similarity index and the mean absolute error, calculate the data scores between the prediction results of the teacher model after all unannotated images are input into the model and the prediction results of the student model : : ; where represents the index of the unannotated image; and represent the structural similarity index and the mean absolute error respectively; the subscript refers to unannotated, the superscript refers to image, the subscript refers to teacher, and the subscript refers to student; Step 1.2.2. Normalize the data scores of all unannotated images, scale the data scores to the numerical interval [0,1] using min-max normalization, and sort the data according to the normalization results; Step 1.2.3: Select data symmetrically starting from the middle position of the sorting result, manually annotate the selected data, and then add it to the labeled dataset; among them, the number of selected data is determined according to the need to increase the number of labeled images. For example, if the current labeled data is 1% (41 images), and if it is necessary to increase it to 5% (202 images) of the labeled data, then 4% (161 images) of the data is symmetrically selected from the middle position; in this embodiment, the labeled data is incrementally selected from 1%, 5%, and 10% in ascending order.
[0026] In the above step 2, the original pixel values of the manually annotated images to be detected are scaled to the interval [0, 1], and normalized using the preset channel mean and variance (the mean and variance of the images in the ImageNet dataset are used in the embodiment); then the image is scaled to a size of 640×640, and the semantic segmentation mask corresponding to the image is also scaled to a size of 640×640; a trainable color enhancer and a trainable geometric transformation enhancer are used to perform data augmentation on the scaled image (the forms of data augmentation include but are not limited to brightness, contrast, color perturbation, and affine transformation, etc.).
[0027] In the above step 3, both the teacher model and the student model of the teacher-student framework are composed of a visual feature extraction backbone network, a text fusion module, and a decoder, and the exponential moving average (EMA) is performed on the parameters of the student model to generate the parameters of the teacher model; the visual feature extraction backbone network uses the Swin pre-trained model to extract the deep features of the image; the decoder uses the BiRefBlock decoder proposed by Peng Zheng et al. in their work "Bilateral Reference for High-Resolution Dichotomous Image Segmentation" published in "CAAI Artificial Intelligence Research" in 2024 (Peng Zheng, Dehong Gao, Deng-Ping Fan, Li Liu, Jorma Laaksonen, Wanli Ouyang, and Nicu Sebe. Bilateral reference for high-resolution dichotomous image segmentation. CAAI Artificial Intelligence Research, 2024.); the visual feature extraction backbone network, the text fusion module, and the decoder all participate in the model training.
[0028] In step 4 above, the process of fusing the image depth feature with the image-level directional text input text into the text-visual fusion feature specifically includes the following steps: Step 4.1 Encode the image-level directional text using a directional text encoder to obtain a text feature ; The directional text encoder uses a pre-trained CLIP text encoder with frozen parameters, which is not involved in model training and is only used to encode text information; The superscript refers to Text; The subscript refers to Labeled, and the subscript represents the index of the labeled data; Step 4.2 For the image depth feature , use the Clue Attention Mechanism to generate a clue attention query : ; ; ; Among them, the subscript represents the th feature in the multi-level feature; represents the clue query vector; represents the learnable query; represents the clue key vector; represents the clue value vector; , and represent the linear projection weights; represents the clue attention map; represents transpose of; represents the dimension of the attention head; represents the softmax activation function; Step 4.3 Calculate the cosine similarity score between the clue attention query and the Code Book : ; Among them, represents the cosine similarity; represents the element in the code book; represents a code book of a set of feature vectors storing category-related information; Step 4.4 Based on the cosine similarity score Perform weighted aggregation on the codebook to generate a clue vector : ; Among them, represents calculating the number of elements in the set; Step 4.5 Insert the clue vector into the text feature to obtain an enhanced text feature : ; Among them, represents a concatenation operation; Step 4.6 Use the Cross-Attention mechanism on the enhanced text feature and the image depth feature to obtain a text-visual fusion feature : ; Among them, represents cross-attention; Step 4.7 Input the text-visual fusion feature into the decoder to generate the prediction result of the teacher model and the prediction result of the student model .
[0029] In the above step 5, the total loss of the network is calculated as follows: ; ; ; ; ; ; ; Among them, represents the supervised loss, which is used to supervise the prediction result of the student model by the true image label; represents a hyperparameter, which is used to balance the supervised loss and the unsupervised loss; represents the unsupervised loss, which is used to supervise the prediction result of the student model by the prediction result of the teacher model; represents the supervised segmentation loss, represents the unsupervised segmentation loss, both of which are composed of binary cross-entropy loss, intersection over union loss, and structural similarity loss; represents the adaptive enhancer loss; Denotes binary cross-entropy loss; 、 Denote the prediction results of the teacher model and the student model after the labeled images are input into the model respectively; Denotes the true annotation of the image; Denotes the loss of the text fusion module; Denotes the clue attention map; Denotes adjusting the resolution using the bilinear interpolation algorithm; Denotes mean squared error loss; Denotes the CLIP pre-trained text encoder; Denotes the class word of the picture.
[0030] Through the above solution, the present invention has an adaptive data augmentation module and an adaptive data selection module, and adaptively selects valuable samples for annotation through adversarial augmentation and sampling strategies; at the same time, the present invention proposes a text fusion module to fully utilize the labeled data by combining camouflage knowledge and vision-text interaction; the present invention proposes a new dataset called RefTextCOD, which contains image-level text annotations for existing mainstream camouflage object detection datasets; simulation experiments are carried out with reference to the settings of existing semi-supervised camouflage object detection. The experimental results show that the method proposed by the present invention surpasses all existing semi-supervised camouflage object detection methods, effectively proving its effectiveness.
[0031] Experimental process: The present invention follows the common settings in the previous semi-supervised camouflage object detection field, uses the CAMO training set and the COD10K training set (excluding images without camouflage objects) combined as the basic training set, and divides based on the labeled data ratios of 1%, 5% and 10%. The division of the labeled data is carried out according to the adaptive data augmentation module proposed by the present invention. For the selection of the test set, the present invention conducts tests on a total of four datasets: CHAMELEON (a total of 76 test images), CAMO-Test (a total of 250 test images), COD10K-Test (a total of 2026 test images), and NC4K (a total of 4121 test images). During the test, two different settings are used for testing respectively: (1) All images use fixed directional text; (2) All images use fine directional text.
[0032] The experimental results are as Figure 4As shown, the experimental results indicate that the present invention is superior to the existing semi-supervised camouflaged object detection models in multiple metrics. Compared with the previous methods, it has improved by 52.0% in the mean absolute error (MAE) and by 19.1% in the structural similarity score (S-Measure), demonstrating the effectiveness of the model proposed in this paper.
[0033] Figure 5 For the functional ablation experiment of the relevant modules of the method proposed by the present invention, it can be seen that the present invention is of great help in improving the performance of semi-supervised camouflaged object detection, further demonstrating the effectiveness of the method proposed by the present invention.
[0034] The above embodiments and diagrams do not limit the product form and style of the present invention. Any appropriate changes or modifications made by those of ordinary skill in the art shall be regarded as not departing from the patent scope of the present invention.
Claims
1. A semi-supervised camouflaged target detection method based on adaptive data selection and text fusion, characterized in that It includes an adaptive data augmentation module, an adaptive data selection module, and a text fusion module, and the following steps are included: Step 1: Use the adaptive data selection module to perform value scoring on unlabeled images, select high-value pictures suitable for model training according to a preset scoring threshold for manual annotation, and construct a labeled dataset containing image-level directional text and real image labels; Step 2: First, preprocess the manually annotated images to be detected, and then use the adaptive data augmentation module to obtain adaptively data-augmented images; Step 3: Input the adaptively data-augmented images into the teacher-student framework, and extract the image depth features by the visual feature extraction backbone network; the teacher-student framework includes a teacher model and a student model; Step 4: Input the image depth features and the image-level directional text into the text fusion module to generate text-visual fusion features, and generate the teacher model prediction result and the student model prediction result by the decoder; Step 5: Calculate the total loss of the entire network using the real image label and the student model prediction result, and then adjust the parameters of the visual feature extraction backbone network, the adaptive data augmentation module, the text fusion module, and the decoder through backpropagation.
2. The semi-supervised camouflaged target detection method based on adaptive data selection and text fusion according to claim 1, characterized in that Step 1 includes a construction step of a camouflaged object detection dataset RefTextCOD containing image-level directional text: Step 1.1.1 Perform directional text annotation on four mainstream camouflaged object detection datasets CHAMELEON, CAMO, COD10K, and NC4K, and obtain the precise foreground area of the camouflaged object in the input image according to the dataset ground truth map; Step 1.1.2 Construct corresponding graph-prompt pairs according to the preset prompt rules; Step 1.1.3 Input the original image, the foreground area of the camouflaged object, and the prompt into the Qwen-VL-Chat model and the GPT4-Vision model respectively to obtain multiple groups of candidate prompts for each group of images; Step 1.1.4 Screen and correct the candidate prompts in an artificial trimming manner to obtain precise directional text; Step 1.1.5 Organize the results into the format of a complete dataset image-text pair.
3. The semi-supervised camouflaged target detection method based on adaptive data selection and text fusion according to claim 2, characterized in that The preset prompt rules are: By giving the original image and using the corresponding segmentation mask to extract the camouflaged object as an auxiliary positioning image, guiding the model to gradually recognize the foreground of the camouflaged object and its corresponding background physical characteristics through prompts, and guiding the VLM to aggregate this information to generate the final annotation.
4. The semi-supervised camouflaged target detection method based on adaptive data selection and text fusion according to claim 1, characterized in that In Step 1, the adaptive data selection module is used to incrementally expand the labeled dataset, specifically including: Step 1.2.1 For the model trained for the first time, the adaptive data selection module extracts features in the color, texture, and frequency domains, classifies the data using K-Means clustering, selects samples close to the cluster centers as the initial labeled data, and proceeds to Step 2; for the model trained again, based on the structural similarity index and the mean absolute error, calculate the data scores of the prediction results of the teacher model after all unlabeled images are input into the model and the prediction results of the student model : ; Among them, represents the index of the unannotated image; and represent the structural similarity index and the mean absolute error respectively; the subscript refers to unannotated, the superscript refers to image, the subscript refers to teacher, the subscript refers to student; Step 1.2.2 Normalize the data scores of all unlabeled images, scale the data scores to the numerical interval [0,1] using min-max normalization, and sort the data according to the normalization results; Step 1.2.3 Symmetrically select data starting from the middle position of the sorting results, perform manual annotation on the selected data, and then add it to the labeled dataset.
5. The semi-supervised camouflaged target detection method based on adaptive data selection and text fusion according to claim 1, wherein: In the said step 2, the original pixel values of the manually labeled images to be detected are scaled to the interval [0, 1], and normalized using the preset channel means and variances; then the images are scaled to a size of 640×640, and the corresponding semantic segmentation masks of the images are also scaled to a size of 640×640; a trainable color enhancer and a trainable geometric transformation enhancer are used to perform data augmentation on the scaled images.
6. The semi-supervised camouflaged target detection method based on adaptive data selection and text fusion according to claim 1, wherein: In the step 3, both the teacher model and the student model of the teacher-student framework are composed of a visual feature extraction backbone network, a text fusion module and a decoder, and the exponential moving average is performed on the parameters of the student model to generate the parameters of the teacher model; the visual feature extraction backbone network uses the Swin pre-trained model to extract the image depth features; The decoder uses the BiRefBlock decoder; the visual feature extraction backbone network, the text fusion module and the decoder are all involved in the model training.
7. The semi-supervised camouflaged target detection method based on adaptive data selection and text fusion according to claim 1, characterized in that The process of inputting the image depth features and the image-level directional text into the text fusion module to generate the text-visual fusion features in the step 4 specifically includes the following steps: Step 4.1 Encode the image-level directional text using a directional text encoder to obtain text features ; The directional text encoder uses a pre-trained CLIP text encoder with frozen parameters. This text encoder does not participate in model training and is only used to encode text information; The superscript refers to Text; The subscript refers to Labeled, the subscript represents the index of labeled data; Step 4.2 For the image depth features , use the clue attention mechanism to generate clue attention queries : ; ; ; Among them, the subscript represents the th feature in the multi-level features; represents the clue query vector; represents the learnable query; represents the clue key vector; represents the clue value vector; , and represent the linear projection weights; represents the clue attention map; represents transpose of; represents the dimension of the attention head; represents the softmax activation function; Step 4.3 Calculate the cosine similarity score between the clue attention query and the codebook : ; Among them, represents the cosine similarity; represents an element in the codebook; represents a codebook of feature vectors storing a set of information related to storage categories; Step 4.4 Based on the cosine similarity score Perform weighted aggregation on the codebook to generate a cue vector : ; Among them, represents calculating the number of elements in a set; Step 4.5 Insert the clue vector into the text feature to obtain the enhanced text feature : ; Among them, represents a splicing operation; Step 4.
6. For the enhanced text features and the image depth features use the cross-attention mechanism to obtain the text-visual fusion features : ; Among them, represents cross-attention; Step 4.7 Input the text-visual fusion feature into the decoder to generate the prediction result of the teacher model and the prediction result of the student model .
8. The semi-supervised camouflaged target detection method based on adaptive data selection and text fusion according to claim 7, characterized in that In step 5, the total loss of the network The calculation formula is as follows: ; ; ; ; ; ; ; Among them, represents the supervised loss, which is used to supervise the prediction results of the student model by the real image labels; represents a hyperparameter used to balance the supervised loss and the unsupervised loss; represents the unsupervised loss, which is used to supervise the prediction results of the student model by the prediction results of the teacher model; represents the supervised segmentation loss, represents the unsupervised segmentation loss, both of which are composed of binary cross-entropy loss, intersection over union loss, and structural similarity loss; represents the adaptive enhancer loss; represents the binary cross-entropy loss; 、 respectively represent the prediction results of the teacher model and the student model after the labeled image is input into the model; represents the real annotation of the image; represents the text fusion module loss; represents the clue attention map; represents adjusting the resolution using the bilinear interpolation algorithm; represents the mean squared error loss; represents the CLIP pre-trained text encoder; represents the class word of the picture.
Citation Information
Patent Citations
Camouflage painting camouflage target detection method and system based on YOLOv5 algorithm
CN115953668A
Teacher-student network method for semi-supervised directive target detection
CN116563687A
Camouflage target detection method based on edge information adaptive feature fusion network
CN118071998A
Multi-modal guide feature enhanced camouflage target detection method
CN119273887A
Authenticity determination support system
US20250014368A1