Knowledge guidance-based camouflage target detection method and system
By using the multi-level semantic knowledge description generated by the multi-modal large language model, combined with the multi-level knowledge aggregation module and the knowledge-guided semantic enhancement adapter module, the problem of identification of complex backgrounds and high intrinsic consistency targets in camouflage object detection is solved, and high-quality camouflage target segmentation and good performance under low data conditions is achieved.
Patent Information
- Application Number
- CN202510097090.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing camouflage object detection methods are insufficient in accuracy and robustness when dealing with complex backgrounds and high intrinsic consistency goals, and face the problems of scarcity of data and high labeling costs.
A knowledge-guided camouflage object detection method is proposed, using a multi-modal large language model to generate multi-level semantic knowledge descriptions, and through a multi-level knowledge aggregation module and a knowledge-guided semantic enhancement adapter module, the model's understanding of camouflage images is enhanced and the dependence on large-scale training data is reduced.
High-quality disguised target segmentation is achieved, the model's performance under low data conditions is improved, the labeling cost is reduced, and the model's semantic understanding ability is enhanced.
Smart Images

Figure CN120070855A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer and information services, and particularly relates to a method and system for detecting camouflaged objects in camouflaged images. Background Art
[0002] Object detection is a fundamental task in the field of computer vision, aiming to identify and locate objects in images. Camouflaged Object Detection (COD) is a fine-grained sub-task in the field of object detection, which aims to detect imperceptible objects hidden in the environment. Camouflaged scenes widely appear in human production activities. For example, there are subtle surface defects in industrial products, it is difficult to distinguish early pathological tissues from normal tissues in medical images, and pests with protective colors in farmland are difficult to be found by the naked eye. Thus, it can be seen that camouflaged object detection has broad application prospects in practical production activities such as industrial defect detection, medical image segmentation, and agricultural pest detection, and thus has attracted more and more attention from researchers in the field of computer vision.
[0003] Different from general objects and salient objects, camouflaged objects usually have similar shapes, colors, and textures to their surrounding environment, and there is a high degree of internal consistency between the object and the environment, resulting in sparse representation features for object perception. Therefore, the camouflaged object detection task is more difficult than general object detection and salient object detection, and requires more complex detection strategies.
[0004] Traditional camouflaged object detection methods rely on low-level features such as brightness, intensity, color, texture, optical flow, etc. These methods have achieved reliable performance in specific scenarios, but they rely on manually designed operators, with limited feature extraction capabilities, and are difficult to handle complex backgrounds and significant object changes, resulting in insufficient accuracy and robustness in different scenarios.
[0005] In recent years, significant progress has been made in deep learning-based camouflaged object detection methods. Researchers have proposed various strategies to address the problem of the inherent consistency between objects and backgrounds, which can be roughly classified into: methods based on multi-scale feature mining, iterative methods, edge-guided methods, frequency-domain information-based methods, uncertainty learning-based methods, graph learning-based methods, etc. Methods based on multi-scale feature mining extract image context information at different scales and enhance representation by aggregating low-level detailed features and high-level semantic features. Iterative methods simulate the human visual mechanism, divide camouflage detection into multiple stages such as discovery, localization, and refinement perception, and decompose difficult tasks into multiple simple steps to improve detection performance. Edge-guided methods enable the model to perceive and predict the contour edges of camouflaged objects, and use edge supervision for constraint to achieve more accurate segmentation results. Frequency-domain information-based methods are based on the idea that different frequency-domain features of images help to discover hidden objects. They use Fourier transform to convert images from the spatial domain to the frequency domain and enhance the model's detection ability by absorbing discriminative information in the frequency-domain features. Uncertainty learning-based methods model the uncertainty of model predictions, strengthen the loss penalty for low-confidence regions, and make the model focus on confusing regions that are difficult to discriminate. Graph learning-based methods project the pixel space into the graph space and mine high-level associations between regional pixels through the graph structure.
[0006] Most deep learning-based methods focus on designing complex network structures and learning strategies to extract perceptual cues of camouflaged objects, rely on large-scale training datasets for training, and have achieved good performance results. However, on the one hand, camouflaged samples are far fewer than general samples, and on the other hand, the difficulty of camouflaged objects themselves makes the annotation cost higher. Camouflaged samples in specific fields such as medical images require professional personnel for annotation. These factors lead to the problem of insufficient high-quality annotated data in real-world applications, and the performance of complex structures specifically designed for camouflage scenarios is often limited under data-scarce conditions.
[0007] As the foundation models trained on large-scale datasets demonstrate impressive capabilities, leveraging these models to address the COD problem has become one of the hotspots of interest among researchers in the field. This mainly falls into two aspects: directly using the foundation models to solve the COD task and alleviating the problem of data scarcity. Some methods directly utilize the powerful representation and generalization capabilities of the foundation models, fine-tuning them on the camouflage dataset to achieve significant performance improvements. For example, CamoDiff regards camouflaged object detection as a conditional generation task, leveraging the high generalization ability of the diffusion model to handle complex, diverse, and significantly different camouflage images, and promoting the model's generalization ability through uncertain supervision. Other methods aim to explore using the foundation models to alleviate the data scarcity problem. For instance, LAKE-RED clusters the target features and uses them as anchors for the background environment, generating an image background similar to the foreground target using the diffusion model, thereby generating camouflage images from general images to mitigate the problem of data shortage. WSCOS uses the SAM model to obtain segmentation pseudo-labels to reduce the annotation cost of camouflage images and learns robust representations from noise supervision. ProMaC utilizes the rich internal knowledge of the multimodal large language model to mine the semantic information of the target, proposing an iterative prompt-mask generation framework to generate prediction masks without training. However, the methods that directly use the foundation models to solve the COD task neglect the problem of data scarcity, while the methods that alleviate data scarcity through the foundation models face the problem of noise signals. Summary of the Invention
[0008] To overcome the deficiencies of existing methods, the present invention proposes a knowledge-guided camouflaged object detection method for achieving high-quality segmentation of camouflaged objects and reducing dependence on training data. For the problem of sparse perceptual features of camouflaged objects, the method of the present invention uses a multimodal large language model to generate knowledge descriptions for camouflage images. Through highly abstract and generalized knowledge representations, it strengthens the model's understanding ability of semantic targets and camouflage scenes and enhances the semantic information representation ability. Specifically, the multi-level knowledge aggregation module generates multi-level semantic knowledge vectors by aggregating the consistent information in multi-level text descriptions, thereby effectively suppressing the noise interference caused by incorrect descriptions and preventing the model features from being defocused due to overly rich text descriptions. To maintain the internal knowledge and general capabilities of the foundation model and improve its performance under low data resources, the method of the present invention proposes a knowledge-guided semantic enhancement adapter module. By freezing the original parameters of the foundation model and adjusting the image visual feature representation in the low-rank space, with the guidance of semantic knowledge for model adjustment, while integrating the semantic information of camouflage images, it maintains the general knowledge and segmentation ability of the visual model.
[0009] The present invention uses the image descriptions generated by a multimodal large language model as auxiliary information to guide the model to better understand camouflaged images, achieving high-quality camouflaged object segmentation. At the same time, based on an efficient parameter fine-tuning method for domain adaptation, the intrinsic knowledge of the visual base model is maintained, effectively alleviating the dependence on large-scale training data.
[0010] The technical solution of the present invention is a knowledge-guided camouflaged object detection method, which includes the following steps:
[0011] Step 1, dataset construction: Select a camouflaged object dataset and divide the dataset into a training set, a validation set, and a test set;
[0012] Step 2, query the multimodal large language model using the defined task instructions to extract multi-level semantic knowledge descriptions;
[0013] Step 3, construct a camouflaged object detection model. The specific processing process is as follows: Encode the multi-level semantic knowledge descriptions into description vectors through a text encoder, and then use the multi-level knowledge aggregation module to input the description vectors into the multi-level knowledge aggregation module to aggregate them into multi-level knowledge vectors. Perform semantic guidance on the multi-level knowledge vectors through several knowledge-guided semantic enhancement adapter modules and the parallel blocks of the encoding layer in the visual base model to obtain semantically enhanced image features. Finally, decode the output prediction mask through the decoding layer in the visual base model;
[0014] Step 4, input the training images and corresponding knowledge descriptions in the training set into the camouflaged object detection model, perform supervised training using the annotations of the training images, test the performance on the validation set, and save the model that performs best on the validation set;
[0015] Step 5, input the test images and corresponding knowledge descriptions in the test set into the best-performing model, and the model predicts the camouflaged object to output a segmentation mask.
[0016] Furthermore, the dataset includes a dataset under normal settings and a dataset under low-data conditions; in the dataset under normal settings, following the general setting method of the dataset in the field of camouflaged object detection, the camouflaged images in the training sets of the two datasets COD10K and CAMO are used as the training set, and the validation sets and test sets of the three datasets COD10K, CAMO, and NC4K are used as the validation set and test set respectively;
[0017] In the dataset under low data conditions, several categories are randomly selected from the COD10K dataset as the base classes, and the samples of other categories are used as new classes. For the base class categories, multiple samples are randomly selected from the base class samples in the COD10K training set to form a small-scale training set, and all the base class samples in the COD10K test set are used as the base class test set; for the new class categories, all the new class samples in the COD10K test set are used to form the new class test set, and the base class test set is used as the validation set; model training and validation are both carried out on the base classes.
[0018] Furthermore, when using a multimodal large language model to query the text description of a camouflaged image, the task instruction consists of three parts: [Task Prompt], [Query], and [Format Instruction]; specifically, [Task Prompt] prompts the multimodal large language model to activate the ability to process a specific task by outlining the content of the specific task; [Query] is used to ask about the key information that humans need to perceive when identifying camouflaged objects, including category, quantity, color, texture, shape, and environment; [Format Instruction] is used to constrain the output response of the multimodal large language model so that it outputs a description content with a fixed format.
[0019] Furthermore, for the obtained knowledge description, rich knowledge description information is obtained by querying the multimodal large language model through the task instruction. The description information includes six aspects: category, quantity, color, texture, shape, and environment. Then, the output text with a fixed template removes the fixed prefix in the instruction and is encoded into a description vector using the text encoder of the CLIP model, and the description vector is input into the multi-level knowledge aggregation module to be aggregated into a multi-level knowledge vector.
[0020] Furthermore, the specific processing process of the multi-level knowledge aggregation module is as follows: Let D a 、D b 、D c 、D d 、D e 、 respectively represent the six embedded description vectors of category, quantity, color, texture, shape, and environment. The category and quantity vectors are concatenated to obtain the complete category information of the entire image, and are aggregated into a category knowledge vector through a layer of multi-layer perceptron; the category description information is multiplied by the color, texture, and shape descriptions respectively to strengthen the consistency information and suppress the interference of inconsistent noise information. After being activated by the sigmoid activation function, they are integrated and finally aggregated into a target knowledge vector through a multi-layer perceptron; the category description information is removed from the environment description information to exclude the influence of the target semantic information, and is integrated into an environment knowledge vector through a multi-layer perceptron;
[0021] The aggregation process of the category knowledge vector and the environment knowledge vector is expressed as:
[0022] K a= MLP(Cat(LN(D a ), LN(D b )))
[0023] K c = MLP(LN(D f ) - LN(D a ))
[0024] where LN(·) and MLP(·) are the linear layer and the multi - layer perceptron respectively, and Cat(·) represents the concatenation operation;
[0025] The aggregation process of the target knowledge vector is expressed as:
[0026] Agg(D) = σ(N(D a ) * LN(D))
[0027] K b = MLP(Agg(D c ) + Agg(D d ) + Agg(D e ))
[0028] where D represents any one of the color, texture, and shape description vectors, Agg(·) represents the consistency information enhancement process, and σ(·) is the sigmoid activation function.
[0029] 6. The knowledge - guided camouflage target detection method according to claim 4, characterized in that: the overall process of semantic guidance is expressed as:
[0030] x l = E l (x l-1 ) + KSEA l (x l-1 )
[0031] where, represents the output feature of the (l - 1) - th parallel block, which is obtained by adding the output feature of the (l - 1) - th encoding layer and the output feature of the (l - 1) - th semantic enhancement adapter module. h, w, and c respectively represent the length, width, and number of channels of the output feature. E l (·) is the encoding layer in the l - th parallel block, and KSEA l (·) is the knowledge - guided semantic enhancement adapter module in the l - th parallel block; the input of the first parallel block is the embedded feature of the camouflage image, which is obtained by performing patch embedding on the input camouflage image, and the patch embedding is implemented through a convolutional layer.
[0032] Furthermore, the specific processing process of the knowledge - guided semantic enhancement adapter module KSEA is as follows:
[0033] Project the input features so that the projected features emphasize semantic information at different levels respectively. After a 1×1 convolution, map the features to a low-rank space, and restore the local details through two 3×3 convolutions to enhance the feature representations at different semantic levels, thereby obtaining the enhanced feature x guided by class knowledge a = Conv3(Conv3(Conv1(K a *x))), the enhanced feature x guided by object knowledge b = Conv3(Conv3(Conv1(K b *x))), the enhanced feature x guided by background knowledge c = Conv3(Conv3(Conv1(K c *x))), where K a , K b , K c represent the knowledge vectors at the three levels of aggregated class, object, and environment respectively; add the enhanced features x a , x b guided by class and object knowledge and pass through two 3×3 convolutions to fuse the global semantic features and object detail features, obtaining Concatenate the enhanced feature x ab with the feature x c guided by background knowledge to integrate the difference representations of foreground and background. Adjust the channels through a 1×1 convolution, and use channel attention and spatial attention to strengthen the difference between foreground and background, expressed as:
[0034] x′ = SA(CA(Conv1(Cat(x ab , x c ))))
[0035] where Conv3 represents a 3×3 convolution, Conv1 represents a 1×1 convolution, is the output feature of the knowledge-guided semantic enhancement adapter module, h, w, and c represent the length, width, and number of channels of the output feature respectively, CA(·) and SA(·) represent the channel attention module and the spatial attention module respectively.
[0036] Furthermore, use the cross-entropy loss L BCE and the intersection over union loss L IOU for training, and the overall loss function is:
[0037] Loss = L BCE + L IOU
[0038] The cross-entropy loss and the intersection over union loss are respectively expressed as:
[0039] LBCE = -∑(P(x, y) * ln(G(x, y)) + (1 - P(x, y)) * ln(1 - G(x, y)))
[0040]
[0041] Where P(x, y) and G(x, y) represent the predicted mask value and the ground truth mask value at the position of coordinates (x, y) respectively.
[0042] Furthermore, during the training process, random cropping, random rotation, and random selection are performed on the input samples to achieve data augmentation; specifically, random flipping horizontally flips the image with a certain probability, random cropping does not exceed k pixel units in length and width, randomly samples an integer from [H - k, H] as the length of the cropped image, randomly samples an integer from [W - k, W] as the width of the cropped image, where H and W represent the length and width of the original image, and crops with the center of the original image as the center; random rotation rotates with a certain probability, and randomly samples an integer angle within the angle range to rotate the cropped image.
[0043] The augmented image is input into the constructed KGDA camouflage target detection model to generate a camouflage target prediction mask, calculate the loss through the loss function, and perform backpropagation to update the parameters. After several rounds of training, the model effect is tested on the validation set, and the model with the best performance on the validation set is saved, that is, the model with the smallest mean absolute error between all prediction masks and ground truth masks.
[0044] The present invention also provides a knowledge-guided camouflage target detection system, including:
[0045] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the knowledge-guided camouflage target detection method as described in the above technical solution.
[0046] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0047] a) A knowledge-guided camouflage target detection method is proposed, which uses a multimodal large language model to obtain multi-level semantic knowledge to help the model understand camouflage images and segment targets.
[0048] b) A multi-level knowledge aggregation module is proposed to integrate noisy text representations into multi-scale knowledge representation vectors, alleviating the impact of ambiguity and inaccuracy in text descriptions, thereby providing effective knowledge guidance.
[0049] c) A knowledge-guided semantic enhancement adapter module is proposed to align text representations with visual representations and inject multi-level semantic knowledge into the visual model to enhance the model's semantic understanding.
[0050] d) By freezing the pre-trained parameters of the visual backbone and fine-tuning a small number of parameters to maintain the general performance of the model, good performance under low-data conditions is achieved. Description of the Drawings
[0051] Figure 1 is the KGDA model structure of the present invention.
[0052] Figure 2 is the multi-level knowledge aggregation module structure proposed by the present invention.
[0053] Figure 3 is the knowledge-guided semantic enhancement adapter module structure proposed by the present invention.
[0054] Figure 4 is the visualization result of the comparative experiment of the present invention. The first column is the input sample, the second column is the ground truth mask, the third column is the predicted mask of the proposed model, and the fourth to tenth columns are the predicted masks of other models. Our method obtains more accurate prediction outputs. Detailed Embodiments
[0055] The technical solution of the present invention will be further described below in conjunction with the drawings and specific embodiments.
[0056] As Figure 1 shown, an embodiment of the present invention provides a knowledge-guided camouflage target detection method, including the following steps:
[0057] Step 1, dataset construction: Select a camouflage target dataset and divide the dataset into a training set, a validation set, and a test set;
[0058] Step 2, description text extraction: Use the defined task instructions to query a multi-modal large language model to extract multi-level semantic knowledge descriptions;
[0059] Step 3, model construction: Construct a KGDA camouflage target detection model. The specific processing process is as follows: Encode the multi-level semantic knowledge description into a description vector through a text encoder, and then use the multi-level knowledge aggregation module MLKA to input the description vector into the multi-level knowledge aggregation module to aggregate it into a multi-level knowledge vector. Perform semantic guidance on the multi-level knowledge vector through a number of knowledge-guided semantic enhancement adapter modules KSEA and the parallel blocks of the encoding layer in the visual base model to obtain semantically enhanced image features, and finally decode and output the predicted mask through the decoding layer in the visual base model; Step 4, model training: Input the training images and corresponding knowledge descriptions in the training set into the KGDA camouflage target detection model, use the annotations of the training images for supervised training, test the performance on the validation set, and save the model with the best performance on the validation set;
[0060] Step 5, Model Testing: Input the test images and corresponding knowledge descriptions in the test set into the best-performing model, and the model predicts the camouflaged target to output a segmentation mask.
[0061] Further, step 1 includes the following steps:
[0062] In the conventional setting, following the general method for setting up the dataset in the field of camouflaged target detection, the camouflaged images in the training sets of the two datasets COD10K and CAMO are used as the training set, and the validation sets and test sets of the three datasets COD10K, CAMO, and NC4K are used as the experimental validation set and test set respectively. Under the low-data test, we randomly select 36 categories from the COD10K dataset as the base classes, and the samples of other categories as the new classes. For the base class categories, 8 samples are randomly selected from the base class samples in the COD10K training set to form a small-scale training set, and all the base class samples in the COD10K test set are used as the base class test set. For the new classes, all the new class samples in the COD10K test set are used to form the new class test set. The same applies to the validation set. Under the low-data setting, both model training and validation are carried out on the base class dataset.
[0063] Further, step 2 includes the following:
[0064] Input each image sample in the dataset into the multi-modal large language model GPT-4o, and use a unified task instruction to query the large model to make it answer the multi-level knowledge description of the input camouflaged image. Among them, the task instruction is designed starting from the cognitive path of humans describing camouflaged images and consists of three parts: [Task Prompt], [Query], and [Format Instruction]. Specifically, [Task Prompt] prompts the model to activate the ability to process specific tasks by outlining the content of a specific task. [Query] is used to ask about the key information that humans need to perceive when identifying camouflaged objects, including category, quantity, color, texture, and background environment. [Format Instruction] is used to constrain the output response of the model so that it outputs descriptive content with a fixed format. The specific instruction used in the present invention is as follows:
[0065]
[0066] Further, the KGDA camouflaged target detection model in step 3 includes the following:
[0067] The KGDA camouflaged target detection model includes a visual foundation model SAM (Segment Anything Model), a text encoder CLIP, a multi-level knowledge aggregation module, and a knowledge-guided semantic enhancement adapter.
[0068] For the multi-level knowledge aggregation module, in order to suppress the noise introduced by possible misdescriptions or hallucination descriptions in the generated text descriptions and avoid the over-abundant semantic knowledge from distracting the model's focus, this module integrates the extracted rich text semantic descriptions into fixed multi-level knowledge vectors, including aggregation information at three levels: category, target, and environment. As Figure 1 shown, we query the multi-modal large language model through the task instructions to obtain rich knowledge description information, and its description information includes six aspects: category, quantity, color, texture, shape, and environment. Then, the output text with a fixed template removes the fixed prefix in the instructions and is encoded into a description vector using the text encoder of the CLIP model. These rich text description information are input into the multi-level knowledge aggregation module to be aggregated into multi-level knowledge vectors, thereby enhancing the consistency information and suppressing the noise interference. Specifically, let D a , D b , D c , D d , D e , represent the six embedded description vectors of category, quantity, color, texture, shape, and environment respectively, where c represents the number of channels. Let K a , K b , represent the aggregated knowledge vectors at the three levels of category, target, and environment respectively, where c′ = c / 2. As Figure 2 shown, the category and quantity vectors are concatenated to obtain the complete category information of the entire image, and are aggregated into a category knowledge vector through a layer of multi-layer perceptron. The category description information is multiplied by the color, texture, and shape descriptions respectively to strengthen the consistency information and suppress the interference of inconsistent noise information. After being activated by the sigmoid activation function and integrated, it is finally aggregated into a target knowledge vector through a multi-layer perceptron. The category description information is removed from the environment description information to exclude the influence of the target semantic information, and is integrated into an environment knowledge vector through a multi-layer perceptron.
[0069] The aggregation process of the category knowledge vector and the environment knowledge vector is expressed as:
[0070] K a = MLP(Cat(LN(D a ), LN(D D )))
[0071] K c = MLP(LN(D f ) - LN(D a ))
[0072] where LN(·) and MLP(·) are the linear layer and the multi-layer perceptron respectively, and Cat(·) represents the concatenation operation.
[0073] The aggregation process of the target knowledge vector is expressed as:
[0074] Agg(D) = σ(N(D a ) * LN(D))
[0075] K b = MLP(Agg(D c ) + Agg(D d ) + Agg(D e ))
[0076] Where D represents any one of the color, texture, and shape description vectors, and Agg(·) represents the consistency information enhancement process. σ(·) is the sigmoid activation function.
[0077] The above is the processing process of knowledge description. Next, the processing process of image input is introduced. For the input camouflage image First, patch embedding is performed. Specifically, a convolutional layer with a convolutional kernel size of 16*16 and a stride of 16 is used to embed the image into the feature space to obtain the embedded features Where H and W represent the length and width of the original image, and h and w represent the length and width of the embedded features, and h = H / 16, w = W / 16.
[0078] For the knowledge-guided semantic enhancement adapter module, we insert this module as a parallel module of the visual base model into the model. By freezing the backbone parameters of the model, the inherent knowledge and general capabilities of the base model pre-trained with large-scale data are maintained. By training a small number of parameters of the semantic enhancement adapter module and the multi-level knowledge aggregation module, the model's adaptation to the camouflage detection task is achieved, realizing efficient fine-tuning under low data resources and accurate recognition under general settings. The overall process of semantic guidance is expressed as:
[0079] x l = E l (x l-1 ) + KSEA l (x l-1 )
[0080] Where, represents the output feature of the (l - 1)-th parallel block, which is obtained by adding the output feature of the (l - 1)-th encoding layer and the output feature of the (l - 1)-th semantic enhancement adapter module. h, w, and c respectively represent the length, width, and number of channels of the output feature. E l (·) is the encoding layer in the l-th parallel block, and KSEA l (·) is the knowledge-guided semantic enhancement adapter module in the l-th parallel block; the input of the first parallel block is the embedded feature of the camouflage image.
[0081] The knowledge-guided semantic enhancement adapter applies multi-level semantic knowledge to consistently guide the model, enhancing the model's ability to understand camouflaged images by injecting semantic knowledge. Specifically, as Figure 3 shown, we project the input features using the multi-level knowledge vectors aggregated by the multi-level knowledge aggregation module, enabling the projected features to emphasize semantic information at different levels. After mapping the features to a low-rank space through 1*1 convolution, local details are restored through two 3*3 convolutions to enhance the feature representations at different semantic levels. The enhanced feature x guided by category knowledge is x a = Conv3(Conv3(Conv1(K a *x))), the enhanced feature x b guided by object knowledge is x b = Conv3(Conv3(Conv1(K c *x))), and the enhanced feature x c guided by background knowledge is x a = Conv3(Conv3(Conv1(K b *x))). Where x c′ = c / 16. The enhanced features x a and x b guided by category and object knowledge are added together and passed through two 3*3 convolutions to fuse global semantic features and object detail features The enhanced feature x ab is concatenated with the feature x c guided by background knowledge to integrate the difference representations of the foreground and background. After adjusting the channels through 1*1 convolution, channel attention and spatial attention are used to strengthen the distinction between the foreground and background. It is expressed as:
[0082] x′ = SA(CA(Conv1(Cat(x ab , x c ))))
[0083] Where is the output feature of the knowledge-guided semantic enhancement adapter, and CA(·) and SA(·) represent the channel attention module and the spatial attention module respectively.
[0084] After passing through several encoding layers and the semantic enhancement adapter module, the final image encoding features are obtained and decoded by the decoder of the SAM model to output the predicted mask. The cross-entropy loss L BCE and the intersection over union loss L IOU are used for training, and the overall loss function is:
[0085] Loss = L BCE + L IOU
[0086] The cross - entropy loss and the intersection - over - union loss are respectively expressed as:
[0087] L BCE = -∑(P(x, y)*ln(G(x, y))+(1 - P(x, y))*ln(1 - G(x, y)))
[0088]
[0089] where P(x, y) and G(x, y) respectively represent the predicted mask value and the ground - truth mask value at the position of coordinates (x, y).
[0090] Furthermore, the model training in step 4 includes the following steps:
[0091] During the training process, in order to increase data diversity, we perform data augmentation on the input samples through steps such as random cropping and random rotation. Specifically, random flipping flips the image horizontally with a probability of 0.5. Random cropping does not exceed 30 pixel units in length and width. Random integers are sampled from [H - 30, H] as the length of the cropped image and from [W - 30, W] as the width of the cropped image, where H and W represent the length and width of the original image, and the cropping is centered on the center of the original image. Random rotation rotates the cropped image by randomly sampling an integer from - 15° to 15° with a probability of 0.2.
[0092] The augmented images are input into the constructed KGDA model to generate predicted masks for camouflage targets, calculate the loss through the loss function, and perform backpropagation to update the parameters. After every 3 rounds of training, the model performance is tested on the validation set, and the model with the best performance on the validation set is saved, that is, the model with the smallest mean absolute error between all predicted masks and ground - truth masks.
[0093] The embodiment of the present invention provides a knowledge - guided camouflage target detection method, which uses descriptive knowledge to promote the model to understand camouflage images, aggregates rich image descriptions through a multi - level knowledge aggregation module and suppresses noise interference, injects semantic guidance knowledge through a knowledge - guided semantic enhancement adapter module, and finally achieves a more accurate segmentation result. In addition, this method has more significant advantages under low - data conditions.
[0094] The following uses a specific example to illustrate the implementation process and experimental results of the present invention.
[0095] Step S1, prepare the conventional - setting and low - data - setting datasets for model training.
[0096] a. Prepare three widely used camouflage target datasets, namely CAMO, COD10K, and NC4K. CAMO contains 1000 training images and 250 test images. COD10K contains 10,000 images, among which 5066 are camouflage images, including 3040 training images and 2026 test images in the camouflage images. NC4K only contains a test set with 4121 camouflage test samples.
[0097] b. Prepare the dataset under the normal setting. Combine the training sets of COD10K and CAMO as the training set, with a total of 4040 training images. Use the test set of COD10K as the validation set for the experiment. The test sets of COD10K, CAMO, and NC4K are used as the test sets for the experiment.
[0098] c. Prepare the dataset under the low-data condition. Randomly select 36 categories from the COD10K dataset as the base classes, and the samples of other categories as the new classes. For the base class categories, randomly select 8 samples from the base class samples in the COD10K training set to form a small-scale training set, with a total of 288 training images. All the base class samples in the COD10K test set are used as the base class test set. For the new classes, all the new class samples in the COD10K test set are used to form the new class test set. Use the base class test set as the validation set. Under the low-data setting, both model training and validation are carried out on the base class dataset.
[0099] The purpose of setting the dataset under the low-data condition is to obtain a dataset with low data, that is, the training data is much less than the normal setting, to verify the generalization ability of the model in the case of a small number of training samples.
[0100] Step S2: Input the dataset samples into the multi-modal large language model to query the knowledge description of the camouflage images as the text content for subsequent semantic guidance.
[0101] a. Our multi-modal large language model selects GPT-4o, and prepare the image samples and task instructions.
[0102] b. Access the GPT-4o model by calling the application programming interface API provided by OpenAI, query the knowledge description of each camouflage image, and generate the text and save it in a txt file.
[0103] Step S3: Set a text encoder for encoding the text description, and set a visual base model for encoding the image features and decoding the output prediction results.
[0104] a. The text encoder adopts the text encoder of the CLIP model to encode the knowledge description into a representation vector.
[0105] b. The visual foundation model adopts the SAM model, which is pre-trained on more than 11 million images, has rich visual knowledge and powerful segmentation capabilities, and is used to predict the segmentation mask of the camouflaged target;
[0106] c. Freeze the decoder parameters of SAM and do not update the parameters during the training process. The decoder parameters are trained normally.
[0107] d. The prediction mask finally generated by the model is constrained by the joint loss of cross-entropy loss and intersection over union loss to ensure that during the training process, the model prediction results gradually approach the true value.
[0108] Step S4: Construct a multi-level knowledge aggregation module. Integrate the noisy text representation into the multi-scale knowledge representation vector, suppress the influence of irrelevant or incorrect information by aggregating consistent information, alleviate the interference caused by ambiguity and inaccuracy in the text description, and provide effective knowledge guidance.
[0109] a. Obtain a multi-dimensional description vector. Input the text description obtained in Step S2 into the token processor of CLIP to map the tokens, and the tokens are encoded into a description vector by the CLIP text encoder.
[0110] b. Aggregate the description vectors into multi-level knowledge vectors. Construct a linear layer and a multi-layer perceptron layer. The category description and quantity description are aggregated into category knowledge, the category description, color description, texture description, and shape description are aggregated into target knowledge, and the category description and environment description are aggregated into environment knowledge. Aggregate the rich text descriptions into a compact knowledge vector, and use consistent information to reduce noise interference and provide effective knowledge guidance.
[0111] Step S5: Construct a knowledge-guided semantic enhancement adapter module to align the text representation with the visual representation, integrate multi-level semantic knowledge into the visual model, use semantic knowledge to guide the model to adjust the feature space, and help the model understand the camouflaged image through semantic injection, so as to better adapt to the downstream task of camouflaged target detection.
[0112] a. Construct a branch knowledge-guided semantic enhancement adapter module in each encoding layer of the SAM model encoder to adjust the image feature space of each layer.
[0113] b. Multiply the category knowledge vector, target knowledge vector, and environment knowledge vector with the image features respectively to enhance the semantic information of different levels.
[0114] c. Construct one layer of 1*1 convolution and two layers of 3*3 convolution to map the semantic enhancement features to the low-rank space and enhance the features in the low-rank space.
[0115] d. The category enhancement features and the target enhancement features are fused with each other and then concatenated with the environmental enhancement features, thereby enriching the semantic information of the camouflage target and distinguishing the differential features between the target and the environment.
[0116] e. Construct a 1*1 convolution layer, a channel attention layer and a spatial attention layer to enhance the differential features between the target and the background and improve the camouflage perception ability of the model.
[0117] The model training is implemented on the Ubuntu operating system. The model is built through the Pytorch deep learning framework and the GPU is used for calculation. The specific software and hardware parameters used in this example are shown in the following table. In this example, the Adam optimizer is used for optimization, the batch size is set to 6, the learning rate is set to 0.0002, and 60 epochs of optimization are performed.
[0118]
[0119] During model training, the training set, validation set and test set are divided according to step 1, a conventional dataset and a low-data setting dataset are constructed, and experiments are carried out under the datasets of both settings. Data augmentation is performed on the input images. The size of the augmented images is uniformly adjusted to 704*704 and the corresponding description texts are input into the constructed model.
[0120] (1) After the image is input into the model, a prediction mask of 1*704*704 is finally obtained. Calculate the cross-entropy loss and the intersection over union loss between the prediction mask and the ground truth mask, and the sum of the two losses is used as the overall loss:
[0121] Loss = L BCE + L IOU
[0122] L BCE = -∑(P(x, y)*ln(G(x, y))+(1 - P(x, y))*ln(1 - G(x, y)))
[0123]
[0124] where P(x, y) and G(x, y) represent the prediction mask and the ground truth mask respectively.
[0125] During the training process, after every three rounds of training of the model, the model is verified on the validation set, the mean of the mean absolute error of all samples of the model on the validation set is calculated, and the current test result is compared with the previous test result. If it is better than the previous performance, the model is saved.
[0126] After the model is trained, its performance is tested on the test set. During the test, the camouflaged image samples are scaled to a size of 704*704 without data augmentation operations. The model performance is evaluated using subjective and objective evaluations.
[0127] The objective evaluation is the calculation of metrics. The S-measure (S α ), F-measure (F β ), Mean Absolute Error (MAE), and E-measure of the predicted mask and the ground truth mask are calculated respectively to evaluate the segmentation accuracy of the model.
[0128] The S-measure quantifies the spatial structural similarity between the predicted mask and the ground truth mask, which combines the object-aware evaluation S o and the region-aware evaluation S r , and is expressed as follows:
[0129] S α = α * S o + (1 - α) * S r
[0130] where α ∈ [0, 1] is the weight factor, usually set to 0.5.
[0131] The F-measure is used to calculate the relationship between precision and recall. It adjusts the mask to the range [0, 255] and divides it into a binary mask by a threshold, and is expressed as follows:
[0132]
[0133] where M(T) represents binarizing the predicted mask with threshold T. |·| represents the total area of the mask.
[0134] The F-measure is expressed as:
[0135]
[0136] β 2 is usually set to 0.3.
[0137] The E-measure evaluates the local and global similarities between the predicted mask and the ground truth mask, and is expressed as:
[0138]
[0139] where is the enhanced alignment matrix, and W and H are the width and height of the input image. C and G are the predicted mask and the ground truth mask respectively.
[0140] The mean absolute error metric measures the mean absolute error per pixel between the normalized predicted mask and the ground truth mask, expressed as:
[0141]
[0142] where W and H are the height and width of the input image, and |·| represents the absolute value operation. C and G are the predicted mask and the ground truth mask respectively.
[0143] By calculating various metrics and comparing them with those of other models, the detection performance of the model is reflected. The method of the present invention achieves the optimal evaluation metrics among 15 comparison models.
[0144] The subjective evaluation compares the visual effects of the predicted masks. As can be seen from Figure 4 it, the method of the present invention achieves a more accurate segmentation effect and has obvious advantages in detecting small targets, confusing targets, multi-targets, etc.
[0145] On the other hand, the embodiment of the present invention also provides a knowledge-guided camouflaged target detection system, including:
[0146] a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the knowledge-guided camouflaged target detection method as described in the above technical solution.
[0147] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A knowledge-guided camouflaged target detection method, characterized in that: The following steps are involved: Step 1, data set construction: select the disguised target data set and divide the data set into training set, validation set and test set; Step 2: Use the defined task instructions to query the multimodal large language model and extract multi-level semantic knowledge descriptions; Step 3, constructing a disguised target detection model. The specific processing process is as follows: the multi-level semantic knowledge description is encoded into a description vector through a text encoder, and then the description vector is input into the multi-level knowledge aggregation module by a multi-level knowledge aggregation module to be aggregated into a multi-level knowledge vector. The multi-level knowledge vector is semantically guided by several knowledge-guided semantic enhancement adapter modules and the parallel blocks of the encoding layer in the visual basic model to obtain the semantically enhanced image features, and finally the prediction mask is decoded and output by the decoding layer in the visual basic model; Step 4: Input the training images and corresponding knowledge descriptions in the training set into the disguised target detection model, use the annotations of the training images for supervised training, test the performance on the validation set, and save the model with the best performance on the validation set; In step 5, the test images in the test set and the corresponding knowledge descriptions are input to the best performing model, and the model predicts the camouflaged target output segmentation mask.
2. The method for detecting camouflaged targets based on knowledge guidance according to claim 1, characterized in that: The datasets include datasets under conventional settings and datasets under low-data conditions. In the datasets under conventional settings, the general setting method of datasets in the field of camouflaged target detection is followed, and the camouflaged images of the training sets in the COD10K and CAMO datasets are used as training sets, and the validation sets and test sets of the COD10K, CAMO, and NC4K datasets are used as validation sets and test sets respectively. In the dataset under low data conditions, several categories are randomly selected from the COD10K dataset as base categories, and samples of other categories are used as new categories. For the base category, multiple samples are randomly selected from the base class samples of the COD10K training set to form a small-scale training set, and all base class samples of the COD10K test set are used as the base class test set; for the new category, all new class samples in the COD10K test set constitute the new class test set, and the base class test set is used as the verification set; model training and verification are both performed on the base class.
3. The method for detecting camouflaged targets based on knowledge guidance according to claim 1, characterized in that: When using a multimodal large language model to query the text description of a camouflaged image, the task instruction consists of three parts: [task prompt], [query], and [format instruction]. Specifically, [task prompt] prompts the multimodal large language model to activate the ability to handle specific tasks by outlining the content of the specific task; [query] is used to inquire about the key information that humans need to perceive when identifying camouflaged objects, including category, quantity, color, texture, shape, and environment; [format instruction] is used to constrain the output response of the multimodal large language model so that its output has a description content in a fixed format.
4. The method for detecting camouflaged targets based on knowledge guidance according to claim 1, characterized in that: For the acquired knowledge description, the multimodal large language model is queried through task instructions to obtain rich knowledge description information, which includes six aspects: category, quantity, color, texture, shape and environment. Then, the output text with a fixed template is removed from the fixed prefix in the instruction, and encoded into a description vector using the text encoder of the CLIP model. The description vector is then input into the multi-level knowledge aggregation module and aggregated into a multi-level knowledge vector.
5. The knowledge-guided camouflaged target detection method according to claim 1, characterized in that: The specific processing process of the multi-level knowledge aggregation module is as follows: a , D b , D c , D d , D e , Represent the six embedded description vectors of category, quantity, color, texture, shape and environment respectively, concatenate the category and quantity vectors to obtain the complete category information of the whole image, and aggregate them into a category knowledge vector through a layer of multi-layer perceptron; multiply the category description information with the color, texture and shape description respectively to strengthen the consistency information and suppress the interference of inconsistent noise information, integrate them after activation by the sigmoid activation function, and finally aggregate them into the target knowledge vector through a multi-layer perceptron; remove the category description information from the environment description information to eliminate the influence of the target semantic information, and integrate them into the environment knowledge vector through a multi-layer perceptron; The aggregation process of the category knowledge vector and the environment knowledge vector is expressed as: K a =MLP(Cat(LN(D a ),LN(D b ))) K c =MLP(LN(D f )-LN(D a )) Where LN(·) and MLP(·) are linear layers and multi-layer perceptrons, respectively, and Cat(·) represents the concatenation operation; The aggregation process of the target knowledge vector is expressed as: Agg(D)=σ(N(D a )*LN(D)) K b =MLP(Agg(D c )+Agg(D d )+Agg(D e )) Where D represents any one of the color, texture, and shape description vectors, Agg(·) represents the consistency information enhancement process, and σ(·) is the sigmoid activation function.
6. The knowledge-guided camouflaged target detection method according to claim 1, characterized in that: The overall process of semantic guidance is expressed as: x l =E l (x l-1 )+KSEA l (x l-1 ) in, represents the output feature of the l-1th parallel block, which is obtained by adding the output feature of the l-1th encoding layer and the output feature of the l-1th semantic enhancement adapter module. h, w, and c represent the length, width, and number of channels of the output feature, respectively. E l (·) is the coding layer in the lth parallel block, KSEA l (·) is the knowledge-guided semantic enhancement adapter module in the lth parallel block; the input of the first parallel block is the embedded feature of the disguised image, which is obtained by performing patch embedding on the input disguised image, and the patch embedding is implemented through a convolutional layer.
7. The knowledge-guided camouflaged target detection method according to claim 1, characterized in that: The specific processing process of the knowledge-guided semantic enhancement adapter module is as follows: For input features Projection is performed so that the projected features emphasize semantic information at different levels respectively. The features are mapped to a low-rank space through a 1*1 convolution. Local details are restored through two 3*3 convolutions to enhance the feature representations at different semantic levels, thereby obtaining enhanced features x guided by category knowledge. a =Conv3(Conv3(Conv1(K a *x))), enhanced features x guided by target knowledge b =Conv3(Conv3(Conv1(K b *x))), enhanced features guided by background knowledge x c =Conv3(Conv3(Conv1(K c *x))), where K a , K b , K c Represent the knowledge vectors of the three levels of category, target and environment after aggregation respectively; the enhanced features x guided by category and target knowledge are a 、x b Add and pass two 3*3 convolutions to fuse the global semantic features and target detail features to obtain The feature x will be enhanced ab The feature x guided by the same background knowledge c Splicing is used to integrate the difference between foreground and background. After 1*1 convolution, the channel is adjusted. Channel attention and spatial attention are used to strengthen the difference between foreground and background. It is expressed as: x′=SA(CA(Conv1(Cat(x ab ,x c )))) Among them, Conv3 represents 3*3 convolution, Conv1 represents 1*1 convolution, is the output feature of the knowledge-guided semantic enhancement adapter module, h, w, c represent the length, width, and number of channels of the output feature, respectively, and CA(·) and SA(·) represent the channel attention module and the spatial attention module, respectively.
8. The knowledge-guided camouflaged target detection method according to claim 1, characterized in that: Using cross entropy loss L BCE and intersection loss L IOI For training, the overall loss function is: Loss=L BCE +L IOU The cross entropy loss and intersection-over-union loss are expressed as: L BCE =-∑(P(x,y)*ln(G(x,y))+(1-P(x,y))*ln(1-G(x,y))) Where P(x, y) and G(x, y) represent the predicted mask value and the true mask value at the coordinate (x, y) position, respectively.
9. The knowledge-guided camouflaged target detection method according to claim 1, characterized in that: During the training process, the input samples are randomly cropped, randomly rotated, and randomly selected to achieve data enhancement. Specifically, the random flipping flips the image horizontally with a certain probability, the random cropping does not exceed k pixel units in length and width, and randomly samples integers from [Hk, H] as the cropped image length, and randomly samples integers from [Wk, W] as the cropped image width. H and W represent the length and width of the original image, and the cropping is performed with the center of the original image as the center. The random rotation rotates with a certain probability, and the cropped image is rotated by randomly sampling integer angles within the angle range. The enhanced image is input into the constructed camouflaged target detection model to generate a camouflaged target prediction mask. The loss is calculated through the loss function, and the parameters are updated by backpropagation. After several rounds of training, the model effect is tested on the validation set, and the best performing model on the validation set is saved, that is, the model with the smallest mean absolute error between all predicted masks and the true value masks.
10. The knowledge-guided camouflaged target detection system is characterized by: include: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the knowledge-guided camouflaged target detection method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Video target detection avoidance system and method
CN113743231A
Camouflage target detection method based on multi-task adapter fine tuning
CN116524183A
Camouflage target detection method based on space-frequency domain positioning and edge diffusion enhancement
CN117809338A
Camouflage object detection method based on multi-clue sliding window attention
CN118691837A
Counterfactual context-aware texture learning for camouflaged object detection
US20240312194A1
Cited By
Fabric matching method and system based on multi-modal information matching
CN120976585A
A fabric matching method and system based on multi-modal information matching
CN120976585B
Unified medical image segmentation method based on context hierarchical guidance
CN121073995A