Camouflage target detection method for pseudo label optimization based on point supervision
By using a point-supervised pseudo-label optimization method, leveraging a multi-mask weighted fusion mechanism of point annotation and semantic entropy, and combining a global-local perception module and bidirectional cross-layer fusion, the camouflage target detection model is optimized, solving the problems of high annotation cost and low detection accuracy in existing technologies, and achieving more efficient camouflage target detection.
Patent Information
- Application Number
- CN202511393612.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-28
AI Technical Summary
Existing methods for detecting camouflaged targets rely on fully supervised training, which requires a large amount of manual annotation, resulting in high annotation costs. They also suffer from problems such as over-response to background regions, incomplete activation of target regions, and inaccurate boundary detection. Existing weakly supervised methods still suffer from sparse supervision signals and accuracy gaps.
A pseudo-label optimization method based on point supervision is adopted to obtain bounding box annotations through point annotation. Combined with the multi-mask weighted fusion mechanism of the object segmentation model and semantic entropy, a global-local perception module is constructed to enhance feature interaction. The detection model is optimized by bidirectional cross-layer fusion and loss function.
While reducing annotation costs, it improves the accuracy and segmentation effect of camouflage target detection, solves the problems of unstable response and blurred boundaries, and improves the accuracy of weakly supervised camouflage target detection.
Smart Images

Figure CN120876840A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of camouflaged target detection, specifically to a camouflaged target detection method based on point supervision and pseudo-label optimization. Background Technology
[0002] Camouflaged target detection is a computer vision task aimed at identifying targets that blend seamlessly with the background and whose visual features are not readily apparent. Because camouflaged targets typically exhibit high similarity to their surroundings in terms of color, texture, and shape, their detection is far more challenging than general target detection tasks. Early camouflaged target detection methods relied primarily on manually designed shallow features; however, these handcrafted features have limited expressive power and struggle to handle target concealment in complex backgrounds, resulting in low detection accuracy. With the development of deep learning and the release of large-scale pixel-level labeled datasets, convolutional neural network-based camouflaged target detection methods have made significant progress. However, most existing high-performance camouflaged target detection models rely on fully supervised training, requiring a large amount of manual input of pixel-level masks as training labels. This label acquisition process is extremely time-consuming, with an average annotation cost of up to 60 minutes per image, severely limiting the scalability and practicality of the algorithms.
[0003] Due to the high visual similarity between the target and the background, existing deep learning methods still face the following core challenges in camouflaged target detection: (1) Over-reaction to background regions: some methods generate incorrect activation in the background region, resulting in false detection; (2) Incomplete activation of target regions: only local regions of the target can be identified, and the overall outline is unclear; (3) Inaccurate boundary detection: it is difficult to accurately distinguish between the target and the background edge, resulting in blurred segmentation.
[0004] To reduce the time cost of annotation, some studies have begun to explore weakly supervised camouflage target detection methods. Existing weakly supervised camouflage target detection methods still generally have the following problems: (1) The supervision signals are too sparse, which limits the model's learning ability; (2) There is still a significant gap in accuracy between weakly supervised methods and fully supervised methods; (3) Although existing improved methods have integrated multiple weak supervision signals, they still rely on a large number of points or boxes for manual annotation, which restricts their annotation efficiency and practical application promotion. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention provides a pseudo-label optimization method for detecting camouflaged targets based on point supervision, which improves the detection accuracy of camouflaged objects as much as possible while reducing the annotation time cost and achieving better segmentation results.
[0006] To achieve the above objectives, the technical solution adopted by this invention is: a camouflaged target detection method based on point supervision and optimized pseudo-labels, the detection method comprising the following steps: Step S1: Input the image into the backbone network for preprocessing to obtain the corresponding initial prediction map; Step S2: Based on the image and the initial prediction map from Step S1, obtain multiple prompts, input the multiple prompts into the object segmentation model, and obtain the mask; Step S3: The initial prediction map obtained in step S1 and the mask obtained in step S2 are weighted and fused using the semantic entropy multi-mask weighted fusion mechanism to obtain pseudo-labels; Step S4: Construct a global-local perception module to obtain global-local interaction enhancement features; Step S5: Perform bidirectional cross-layer fusion on the global-local interactive enhancement features obtained in step S4 to obtain fused features, and process the obtained fused features to obtain a prediction map. Step S6: Construct a camouflaged target detection model based on the pseudo-labels obtained in step S3 and the prediction map obtained in step S5.
[0007] Furthermore, in step S1, the image is input into the backbone network for preprocessing to obtain the corresponding initial prediction map; specifically: Using the point annotations in the image as the center and r as the radius, a circular supervision signal is obtained. This circular supervision signal is then used to supervise the backbone network to obtain the corresponding initial prediction map.
[0008] Furthermore, in step 2, multiple prompts are obtained based on the image and the initial prediction map from step S1. These multiple prompts are then input into the object segmentation model to obtain a mask; specifically: Step 21: Based on the maximum bounding matrix in the initial prediction map of step S1, obtain the bounding box annotations; combine the point annotations of the image and the obtained bounding box annotations to form a multi-cue. Step 22: Input the multi-cue input into the object segmentation model to obtain point masks and bounding box masks.
[0009] Furthermore, in step S3, a semantic entropy multi-mask weighted fusion mechanism is used to weight and fuse the initial prediction map obtained in step S1 and the mask obtained in step S2 to obtain pseudo-labels; the mask is divided into point mask and boundary mask: specifically: Step S31, the initial prediction map and dot mask respectively and bounding box mask The point semantic entropy and bounding box semantic entropy are calculated using the following formulas: ; in, Represents point semantic entropy, Represents dot product operation. Represents the logarithmic operation. Represents the smallest positive number; ; in, Represents the semantic entropy of the bounding box; Step S32, dot mask and bounding box mask Perform spatial domain decomposition to extract overlapping regions. = ⋂ Non-overlapping areas are designated as exclusive areas. = - And the bounding box indicates the exclusive area = - ; Step S33, based on the calculated point semantic entropy and bounding box semantic entropy The point mask weights and bounding box mask weights are obtained by assigning weights to non-overlapping regions based on semantic entropy; the formulas are as follows: ; ; in, Represents the point mask weights. Represents the weights of the bounding box mask; Step S34, for overlapping areas Non-overlapping areas are designated as exclusive areas. And the bounding box indicates the exclusive area According to the dot mask weight and bounding box mask weights The pseudo-labels are obtained by weighted fusion; the formula is as follows: ; in, This represents a pseudo-tag.
[0010] Furthermore, in step S4, a global-local perception module is constructed to obtain global-local interaction enhancement features; specifically: Step S41: Input the image into the backbone network, extract the initial features of the image, and construct a global-local perception module; Step S42: Based on the global-local perception module constructed in step S41; the global-local perception module includes a global perception module and a local perception module. The global perception module obtains global perception features, the local perception module obtains local perception features, and global-local interaction enhancement features are obtained based on the global perception features and the local perception features.
[0011] Furthermore, in step S42, the global perception module obtains the global perception features; specifically: Step S4201: The initial features of the image are processed by two parallel dilated convolutions with kernel sizes of 3×3 and 5×5 to extract spatial features under different receptive fields. The extracted spatial features are then added to the initial features of the image through a residual structure to obtain residual features. Step S4202: Input the residual features into an improved frequency domain self-attention mechanism to obtain the spatial domain residual features, and then transform the spatial domain residual features into frequency domain features through fast Fourier transform to obtain frequency domain query features, frequency domain matching features and frequency domain information features. Step S4203: Perform dimensional transformation on the obtained frequency domain query features, frequency domain matching features, and frequency domain information features to obtain dimensionally transformed frequency domain query features, frequency domain matching features, and frequency domain information features. Perform matrix multiplication on the dimensionally transformed frequency domain query features and frequency domain matching features to generate the transposed reference diagram A. f ; Step S4204, note the transpose diagram A. f Perform the calculation; the formula is as follows: ; in, Representative transpose attention diagram A f The real part of a complex number, conj represents the calculation of the conjugate value of the complex number; ; in, Representative transpose attention diagram A f The imaginary part of a complex number, where j represents the imaginary unit in complex number theory; Step S4205: Activate the real part and the imaginary part of the complex number of the transposed attention map using the flexible maximum activation function and then recombine them to obtain the intermediate attention map. Then, perform matrix multiplication on the intermediate attention map and the frequency domain information features after dimension transformation, and obtain the preliminary frequency enhancement features by inverse fast Fourier transform back to the spatial domain features. Step S4206: Frequency residual connection yields frequency enhancement features; the formula is as follows: ; in, Represents frequency enhancement characteristics, Represents an indicator function. Represents the inverse fast Fourier transform. Modify the activation function of the linear unit. Represents Fast Fourier Transform, Represents residual characteristics, Represents matrix multiplication; Step S4207: The initial frequency enhancement feature and the frequency enhancement feature are concatenated in the channel dimension, and the channel dimension is transformed by a 1×1 convolution to obtain the frequency domain self-attention enhancement feature; Step S4208: The initial features of the image and the frequency domain self-attention enhancement features are added through a residual connection, and the frequency domain enhancement features are obtained through a layer normalization process. The frequency domain enhancement features are then input into a depth-separable convolutional group with three parallel convolutional kernels of 3×3, 5×5 and 7×7 to obtain the orientation enhancement features. Step S4209: The frequency domain enhancement features and the direction enhancement features are combined through a residual structure to obtain the global perception features.
[0012] Furthermore, in step S42, the local perception module obtains local perception features; specifically: Step S4210: The initial features of the image are fed into the depthwise convolution branches with three parallel convolution kernels of 1×3, 3×1, and 3×3. The depthwise convolution branches extract local features in three different directions, namely, the local features in the horizontal direction. Vertical local features Local features of the same polarity as those in directions other than horizontal and vertical. ; Step S4211: The extracted local features from three different directions are concatenated along the channel dimension and activated using the Gaussian error linear unit activation function, and then combined with the initial features of the image. The directional characteristics are obtained by adding them together; the formula is as follows: ; in, Representing directional characteristics, Conv 1x1 Represents a regular 1x1 convolution. Represents the splicing of channel dimensions; Step S4212: The orientation features are processed by a layer normalization and a feedforward neural network, and the processed orientation features are added to a residual structure to obtain local perceptual features.
[0013] Furthermore, in step S42, global-local interaction enhancement features are obtained based on global perception features and local perception features; specifically: Step S4213: Divide the global perception features obtained in step S4209 and the local perception features obtained in step S4212 into four equal parts along the channel dimension to obtain four sets of global perception features and four sets of local perception features. The size of the four sets of global perception features and the four sets of local perception features is C / 4×H×W; where C represents the number of channels, H represents the height, and W represents the width. Step S4214: After compressing the spatial dimension of the four sets of global perception features by global average pooling, the four sets of global perception features are concatenated to obtain four new sets of global perception features. After compressing the space of the four sets of local perception features by global average pooling, the four sets of local perception features are concatenated to obtain four new sets of local perception features. Step S4215: The four new sets of global perception features and the four new sets of local perception features are respectively calculated by the lightweight fully connected network to obtain four sets of normalized fusion weights; wherein the lightweight fully connected network includes two fully connected layers and a modified linear unit activation function. Step S4216: Using four sets of normalized fusion weights, the four sets of global perception features and the four sets of local perception features are weighted and fused to obtain four sets of global-local features, and the four sets of global-local features are spliced together in the channel dimension. By fusing global and local perceptual features using residual connections, global-local enhanced features are obtained; the formula is as follows: ; Where i represents the level and m represents the group number. Represents the global-local enhancement features of the i-th layer. Represents the fusion weights of the m-th group in the i-th layer. This represents the global perception feature of the m-th group in the i-th layer. This represents the local perceptual features of the m-th group in the i-th layer. This represents the global perception features of the i-th layer. Represents the local perceptual features of the i-th layer; Step S4217: After the global-local enhancement features are processed by the gated convolution mechanism, they are fused with the global-local enhancement features through residual connections to obtain global-local interactive enhancement features. These global-local interactive enhancement features are divided into four layers: the initial global-local interactive enhancement features, the second adjacent layer global-local interactive enhancement features, the third adjacent layer global-local interactive enhancement features, and the deep global-local interactive enhancement features; the formula is as follows: ; in, Represents global-local interaction enhancement features Representatives criticized normalization. This represents a 3×3 gated convolution.
[0014] Furthermore, in step S5, the global-local interactive enhancement features obtained in step S4 are subjected to bidirectional cross-layer fusion to obtain fused features, and the obtained fused features are processed to obtain a prediction map; specifically: Step S51: Apply a 3x3 depthwise separable convolution to the deep global-local interaction enhancement features to obtain the separable features of the deep global-local interaction enhancement features; The separable features of the deep global-local interaction enhancement features are generated by passing a set of point convolution groups and a sigmoid activation function to produce a first feature weight map. The first feature weight map is then multiplied with the global-local interaction enhancement features of the third adjacent layer to obtain the first layer semantic enhancement features. Global flat pooling is performed on the global-local interaction enhancement features of the third adjacent layer. Then, a second feature weight map is generated through a set of 1x1 convolutions and sigmoid activation functions. The second feature weight map is multiplied with the separable features of the deep global-local interaction enhancement features to obtain the first layer of local spatial refinement features. The obtained first-layer semantic enhancement features and first-layer local spatial refinement features are aggregated to obtain the first-layer fusion features; Applying a 3x3 depthwise separable convolution to the first layer of fused features yields the separable features of the first layer of fused features; The separable features of the first layer fusion features are used to generate a third feature weight map through a set of point convolution groups and a sigmoid activation function. The third feature weight map is then multiplied with the global-local interaction enhancement features of the second adjacent layer to obtain the second layer semantic enhancement features. Global flat pooling is performed on the global-local interaction enhancement features of the second adjacent layer. Then, a fourth feature weight map is generated through a set of 1x1 convolutions and sigmoid activation functions. The fourth feature weight map is multiplied with the separable features of the first layer fusion features to obtain the second layer local spatial refinement features. The obtained second-layer semantic enhancement features and second-layer local spatial refinement features are aggregated to obtain the second-layer fusion features; Applying a 3x3 depthwise separable convolution to the second-layer fused features yields the separable features of the second-layer fused features; The separable features of the second layer of fusion features are used to generate the fifth feature weight map through a set of point convolution groups and a sigmoid activation function. The fifth feature weight map is then multiplied with the initial layer global-local interaction enhancement features to obtain the third layer of semantic enhancement features. Global flat pooling is performed on the initial global-local interaction enhancement features. Then, a sixth feature weight map is generated through a set of 1x1 convolutions and sigmoid activation functions. The sixth feature weight map is multiplied with the separable features of the second layer fusion features to obtain the third layer local spatial refinement features. The third-layer semantic enhancement features and the third-layer local spatial refinement features are aggregated to obtain the final layer of fused features; Step S52: The last layer of fused features is processed using 1x1 convolution, batch normalization, and sigmoid activation function to obtain the prediction map.
[0015] Furthermore, in step S6, a camouflage target detection model is constructed based on the pseudo-labels obtained in step S3 and the prediction map obtained in step S5. The binary cross-entropy loss function and the contrastive loss function are used as the final loss functions to train the constructed camouflage target detection model; specifically: Step S61: Calculate the partial binary cross-entropy loss and contrast loss of the prediction graph, using the following formulas: ; in, Represents the binary cross-entropy loss. Representative pseudo-label The true category of a pixel location in an image. The predicted category represents the pixel location in the image from which the predicted map is located; ; in, Represents the contrast loss, where d represents the number of pixels in the image. The predicted image representing the image. The image represents a predicted image obtained after Gaussian blurring, color jittering, and random rotation. Step S62: Train the camouflage target detection model based on the final loss function.
[0016] The beneficial effects of this invention are:
[0017] (1) Based on the advantages and disadvantages of the current weakly supervised camouflage target detection methods, bounding box annotations are obtained through point annotations. Combined with the Segmentation of Everything (SAM) model and the multi-mask weighted fusion mechanism based on semantic entropy, the quality of the supervision signal is improved while reducing the annotation cost.
[0018] (2) This invention not only considers the cross-layer fusion of global-local interaction enhancement features in most camouflage target detection methods, but also considers the global-local interaction enhancement features of global perception features and local perception features, which solves the problems of unstable response and blurred boundaries in detection and improves the accuracy of weakly supervised camouflage target detection. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the structure of the present invention. Detailed Implementation
[0020] like Figure 1 As shown, a pseudo-label optimization method for camouflaged target detection based on point supervision is proposed, the method comprising the following steps: Step S1: Input the image into the backbone network for preprocessing to obtain the corresponding initial prediction map; Step S2: Based on the image and the initial prediction map from Step S1, obtain multiple prompts, input the multiple prompts into the object segmentation model, and obtain the mask; Step S3: The initial prediction map obtained in step S1 and the mask obtained in step S2 are weighted and fused using the semantic entropy multi-mask weighted fusion mechanism to obtain pseudo-labels; Step S4: Construct a global-local perception module to obtain global-local interaction enhancement features; Step S5: Perform bidirectional cross-layer fusion on the global-local interactive enhancement features obtained in step S4 to obtain fused features, and process the obtained fused features to obtain a prediction map. Step S6: Construct a camouflaged target detection model based on the pseudo-labels obtained in step S3 and the prediction map obtained in step S5.
[0021] Furthermore, in step S1, the image is input into the backbone network for preprocessing to obtain the corresponding initial prediction map; specifically: Using the point annotations in the image as the center and r as the radius, a circular supervision signal is obtained. This circular supervision signal is then used to supervise the backbone network to obtain the corresponding initial prediction map.
[0022] Furthermore, in step 2, multiple prompts are obtained based on the image and the initial prediction map from step S1. These multiple prompts are then input into the object segmentation model to obtain a mask; specifically: Step 21: Based on the maximum bounding matrix in the initial prediction map of step S1, obtain the bounding box annotations; combine the point annotations of the image and the obtained bounding box annotations to form a multi-cue. Step 22: Input the multi-cue input into the object segmentation model to obtain point masks and bounding box masks.
[0023] Furthermore, in step S3, a semantic entropy multi-mask weighted fusion mechanism is used to weight and fuse the initial prediction map obtained in step S1 and the mask obtained in step S2 to obtain pseudo-labels; the mask is divided into point mask and boundary mask: specifically: Step S31, the initial prediction map and dot mask respectively and bounding box mask The point semantic entropy and bounding box semantic entropy are calculated using the following formulas: ; in, Represents point semantic entropy, Represents dot product operation. Represents the logarithmic operation. Represents the smallest positive number; ; in, Represents the semantic entropy of the bounding box; Step S32, dot mask and bounding box mask Perform spatial domain decomposition to extract overlapping regions. = ⋂ Non-overlapping areas are designated as exclusive areas. = - And the bounding box indicates the exclusive area = - ; Step S33, based on the calculated point semantic entropy and bounding box semantic entropy The point mask weights and bounding box mask weights are obtained by assigning weights to non-overlapping regions based on semantic entropy; the formulas are as follows: ; ; in, Represents the point mask weights. Represents the weights of the bounding box mask; Step S34, for overlapping areas Non-overlapping areas are designated as exclusive areas. And the bounding box indicates the exclusive area According to the dot mask weight and bounding box mask weights The pseudo-labels are obtained by weighted fusion; the formula is as follows: ; in, This represents a pseudo-tag.
[0024] Furthermore, in step S4, a global-local perception module is constructed to obtain global-local interaction enhancement features; specifically: Step S41: Input the image into the backbone network, extract the initial features of the image, and construct a global-local perception module; Step S42: Based on the global-local perception module constructed in step S41; the global-local perception module includes a global perception module and a local perception module. The global perception module obtains global perception features, the local perception module obtains local perception features, and global-local interaction enhancement features are obtained based on the global perception features and the local perception features.
[0025] Furthermore, in step S42, the global perception module obtains the global perception features; specifically: Step S4201: The initial features of the image are processed by two parallel dilated convolutions with kernel sizes of 3×3 and 5×5 to extract spatial features under different receptive fields. The extracted spatial features are then added to the initial features of the image through a residual structure to obtain residual features. Step S4202: Input the residual features into an improved frequency domain self-attention mechanism to obtain the spatial domain residual features, and then transform the spatial domain residual features into frequency domain features through fast Fourier transform to obtain frequency domain query features, frequency domain matching features and frequency domain information features. Step S4203: Perform dimensional transformation on the obtained frequency domain query features, frequency domain matching features, and frequency domain information features to obtain dimensionally transformed frequency domain query features, frequency domain matching features, and frequency domain information features. Perform matrix multiplication on the dimensionally transformed frequency domain query features and frequency domain matching features to generate the transposed reference diagram A. f ; Step S4204, note the transpose diagram A. f Perform the calculation; the formula is as follows: ; in, Representative transpose attention diagram A f The real part of a complex number, conj represents the calculation of the conjugate value of the complex number; ; in, Representative transpose attention diagram A f The imaginary part of a complex number, where j represents the imaginary unit in complex number theory; Step S4205: Activate the real part and the imaginary part of the complex number of the transposed attention map using the flexible maximum activation function and then recombine them to obtain the intermediate attention map. Then, perform matrix multiplication on the intermediate attention map and the frequency domain information features after dimension transformation, and obtain the preliminary frequency enhancement features by inverse fast Fourier transform back to the spatial domain features. Step S4206: Frequency residual connection yields frequency enhancement features; the formula is as follows: ; in, Represents frequency enhancement characteristics, Represents an indicator function. Represents the inverse fast Fourier transform. Modify the activation function of the linear unit. Represents Fast Fourier Transform, Represents residual characteristics, Represents matrix multiplication; Step S4207: The initial frequency enhancement feature and the frequency enhancement feature are concatenated in the channel dimension, and the channel dimension is transformed by a 1×1 convolution to obtain the frequency domain self-attention enhancement feature; Step S4208: The initial features of the image and the frequency domain self-attention enhancement features are added through a residual connection, and the frequency domain enhancement features are obtained through a layer normalization process. The frequency domain enhancement features are then input into a depth-separable convolutional group with three parallel convolutional kernels of 3×3, 5×5 and 7×7 to obtain the orientation enhancement features. Step S4209: The frequency domain enhancement features and the direction enhancement features are combined through a residual structure to obtain the global perception features.
[0026] Furthermore, in step S42, the local perception module obtains local perception features; specifically: Step S4210: The initial features of the image are fed into the depthwise convolution branches with three parallel convolution kernels of 1×3, 3×1, and 3×3. The depthwise convolution branches extract local features in three different directions, namely, the local features in the horizontal direction. Vertical local features Local features of the same polarity as those in directions other than horizontal and vertical. ; Step S4211: The extracted local features from three different directions are concatenated along the channel dimension and activated using the Gaussian error linear unit activation function, and then combined with the initial features of the image. The directional characteristics are obtained by adding them together; the formula is as follows: ; in, Representing directional characteristics, Conv 1x1 Represents a regular 1x1 convolution. Represents the splicing of channel dimensions; Step S4212: The orientation features are processed by a layer normalization and a feedforward neural network, and the processed orientation features are added to a residual structure to obtain local perceptual features.
[0027] Furthermore, in step S42, global-local interaction enhancement features are obtained based on global perception features and local perception features; specifically: Step S4213: Divide the global perception features obtained in step S4209 and the local perception features obtained in step S4212 into four equal parts along the channel dimension to obtain four sets of global perception features and four sets of local perception features. The size of the four sets of global perception features and the four sets of local perception features is C / 4×H×W; where C represents the number of channels, H represents the height, and W represents the width. Step S4214: After compressing the spatial dimension of the four sets of global perception features by global average pooling, the four sets of global perception features are concatenated to obtain four new sets of global perception features. After compressing the space of the four sets of local perception features by global average pooling, the four sets of local perception features are concatenated to obtain four new sets of local perception features. Step S4215: The four new sets of global perception features and the four new sets of local perception features are respectively calculated by the lightweight fully connected network to obtain four sets of normalized fusion weights; wherein the lightweight fully connected network includes two fully connected layers and a modified linear unit activation function. Step S4216: Using four sets of normalized fusion weights, the four sets of global perception features and the four sets of local perception features are weighted and fused to obtain four sets of global-local features, and the four sets of global-local features are spliced together in the channel dimension. By fusing global and local perceptual features using residual connections, global-local enhanced features are obtained; the formula is as follows: ; Where i represents the level and m represents the group number. Represents the global-local enhancement features of the i-th layer. Represents the fusion weights of the m-th group in the i-th layer. This represents the global perception feature of the m-th group in the i-th layer. This represents the local perceptual features of the m-th group in the i-th layer. This represents the global perception features of the i-th layer. Represents the local perceptual features of the i-th layer; Step S4217: After the global-local enhancement features are processed by the gated convolution mechanism, they are fused with the global-local enhancement features through residual connections to obtain global-local interactive enhancement features. These global-local interactive enhancement features are divided into four layers: the initial global-local interactive enhancement features, the second adjacent layer global-local interactive enhancement features, the third adjacent layer global-local interactive enhancement features, and the deep global-local interactive enhancement features; the formula is as follows: ; in, Represents global-local interaction enhancement features Representatives criticized normalization. This represents a 3×3 gated convolution.
[0028] Furthermore, in step S5, the global-local interactive enhancement features obtained in step S4 are subjected to bidirectional cross-layer fusion to obtain fused features, and the obtained fused features are processed to obtain a prediction map; specifically: Step S51: Apply a 3x3 depthwise separable convolution to the deep global-local interaction enhancement features to obtain the separable features of the deep global-local interaction enhancement features; The separable features of the deep global-local interaction enhancement features are generated by passing a set of point convolution groups and a sigmoid activation function to produce a first feature weight map. The first feature weight map is then multiplied with the global-local interaction enhancement features of the third adjacent layer to obtain the first layer semantic enhancement features. Global flat pooling is performed on the global-local interaction enhancement features of the third adjacent layer. Then, a second feature weight map is generated through a set of 1x1 convolutions and sigmoid activation functions. The second feature weight map is multiplied with the separable features of the deep global-local interaction enhancement features to obtain the first layer of local spatial refinement features. The obtained first-layer semantic enhancement features and first-layer local spatial refinement features are aggregated to obtain the first-layer fusion features; Applying a 3x3 depthwise separable convolution to the first layer of fused features yields the separable features of the first layer of fused features; The separable features of the first layer fusion features are used to generate a third feature weight map through a set of point convolution groups and a sigmoid activation function. The third feature weight map is then multiplied with the global-local interaction enhancement features of the second adjacent layer to obtain the second layer semantic enhancement features. Global flat pooling is performed on the global-local interaction enhancement features of the second adjacent layer. Then, a fourth feature weight map is generated through a set of 1x1 convolutions and sigmoid activation functions. The fourth feature weight map is multiplied with the separable features of the first layer fusion features to obtain the second layer local spatial refinement features. The obtained second-layer semantic enhancement features and second-layer local spatial refinement features are aggregated to obtain the second-layer fusion features; Applying a 3x3 depthwise separable convolution to the second-layer fused features yields the separable features of the second-layer fused features; The separable features of the second layer of fusion features are used to generate the fifth feature weight map through a set of point convolution groups and a sigmoid activation function. The fifth feature weight map is then multiplied with the initial layer global-local interaction enhancement features to obtain the third layer of semantic enhancement features. Global flat pooling is performed on the initial global-local interaction enhancement features. Then, a sixth feature weight map is generated through a set of 1x1 convolutions and sigmoid activation functions. The sixth feature weight map is multiplied with the separable features of the second layer fusion features to obtain the third layer local spatial refinement features. The third-layer semantic enhancement features and the third-layer local spatial refinement features are aggregated to obtain the final layer of fused features; Step S52: The last layer of fused features is processed using 1x1 convolution, batch normalization, and sigmoid activation function to obtain the prediction map.
[0029] Furthermore, in step S6, a camouflage target detection model is constructed based on the pseudo-labels obtained in step S3 and the prediction map obtained in step S5. The binary cross-entropy loss function and the contrastive loss function are used as the final loss functions to train the constructed camouflage target detection model; specifically: Step S61: Calculate the partial binary cross-entropy loss and contrast loss of the prediction graph, using the following formulas: ; in, Represents the binary cross-entropy loss. Representative pseudo-label The true category of a pixel location in an image. The predicted category represents the pixel location in the image from which the predicted map is located; ; in, Represents the contrast loss, where d represents the number of pixels in the image. The predicted image representing the image. The image represents a predicted image obtained after Gaussian blurring, color jittering, and random rotation. Step S62: Train the camouflage target detection model based on the final loss function.
[0030] The backbone network in step S1 is PVTV2-B4.
[0031] Among them, the object segmentation model in step S2 is SAM.
[0032] In step S4, the pseudo-label serves as a monitoring signal.
[0033] In step S4203, transpose attention diagram A f Since it is a complex number, the softmax activation function cannot be used directly for activation. The transpose needs to be considered (see Figure A). f Calculate the real and imaginary parts of a complex number separately.
[0034] In step S4217, a gated convolution mechanism is further introduced to suppress redundant information in the global-local interaction enhancement features and enhance the expression of key regions.
[0035] The method of this invention (a pseudo-label optimization method for camouflaged target detection based on point supervision) and other methods were compared and evaluated on the concealed target detection dataset CAMO, the camouflaged target detection dataset COD10K, and the camouflaged object detection dataset NC4K. Tables 1, 2, and 3 show the segmentation results of the method of this invention on the concealed target detection dataset CAMO, the camouflaged target detection dataset COD10K, and the camouflaged object detection dataset NC4K, respectively. Evaluation metrics include mean absolute error (MAE), structural metric Sm, enhancement matching metric Em, and segmentation accuracy Fw (weighted F metric).
[0036] Experimental results show that the segmentation accuracy (Fw) of the proposed method on the concealed target detection dataset CAMO is 0.057, 0.839, 0.907, and 0.79, respectively. On the camouflaged target detection dataset COD10K, the accuracy (Fw) is 0.028, 0.841, 0.914, and 0.752, respectively. On the camouflaged object detection dataset NC4K, the accuracy (Fw) is 0.04, 0.858, 0.921, and 0.811, respectively. This demonstrates the superior segmentation accuracy of the proposed method, fully validating its effectiveness in weakly supervised target localization tasks.
[0037] Table 1. Performance comparison of the proposed method and other methods on the concealed target detection dataset CAMO.
[0038] Table 2. Performance comparison of the proposed method and other methods on the camouflage target detection dataset COD10K.
[0039] Table 3. Performance comparison of the proposed method and other methods on the camouflage object detection dataset NC4K. .
Claims
1. A pseudo-label optimization method for camouflaged target detection based on point supervision, characterized in that: Includes the following steps: Step S1: Input the image into the backbone network for preprocessing to obtain the corresponding initial prediction map; Step S2: Based on the image and the initial prediction map from Step S1, obtain multiple prompts, input the multiple prompts into the object segmentation model, and obtain the mask; Step S3: The initial prediction map obtained in step S1 and the mask obtained in step S2 are weighted and fused using the semantic entropy multi-mask weighted fusion mechanism to obtain pseudo-labels. Step S4: Construct a global-local perception module to obtain global-local interaction enhancement features; Step S5: Perform bidirectional cross-layer fusion on the global-local interactive enhancement features obtained in step S4 to obtain fused features, and process the obtained fused features to obtain a prediction map. Step S6: Construct a camouflaged target detection model based on the pseudo-labels obtained in step S3 and the prediction map obtained in step S5.
2. The camouflaged target detection method based on point supervision and pseudo-label optimization according to claim 1, characterized in that: In step S1, the image is input into the backbone network for preprocessing to obtain the corresponding initial prediction map; specifically: Using the point annotations in the image as the center and r as the radius, a circular supervision signal is obtained. This circular supervision signal is then used to supervise the backbone network to obtain the corresponding initial prediction map.
3. The camouflaged target detection method based on point supervision and pseudo-label optimization according to claim 2, characterized in that: In step 2, multiple prompts are obtained based on the image and the initial prediction map from step S1. These multiple prompts are then input into the object segmentation model to obtain a mask; specifically: Step 21: Based on the maximum bounding matrix in the initial prediction map of step S1, obtain the bounding box annotations; combine the point annotations of the image and the obtained bounding box annotations to form a multi-cue. Step 22: Input the multi-cue input into the object segmentation model to obtain point masks and bounding box masks.
4. The camouflaged target detection method based on point supervision and pseudo-label optimization according to claim 3, characterized in that: In step S3, a semantic entropy multi-mask weighted fusion mechanism is used to weight and fuse the initial prediction map obtained in step S1 and the mask obtained in step S2 to obtain pseudo-labels; the mask is divided into point mask and boundary mask: specifically: Step S31, the initial prediction map and dot mask respectively and bounding box mask The point semantic entropy and bounding box semantic entropy are calculated using the following formulas: ; in, Represents point semantic entropy, Represents dot product operation. Represents the logarithmic operation. Represents the smallest positive number; ; in, Represents the semantic entropy of the bounding box; Step S32, dot mask and bounding box mask Perform spatial domain decomposition to extract overlapping regions. = ∩ Non-overlapping areas are designated as exclusive areas. = - And the bounding box indicates the exclusive area = - ; Step S33, based on the calculated point semantic entropy and bounding box semantic entropy The point mask weights and bounding box mask weights are obtained by assigning weights to non-overlapping regions based on semantic entropy; the formulas are as follows: ; ; in, Represents the point mask weights. Represents the weights of the bounding box mask; Step S34, for overlapping areas Non-overlapping areas are designated as exclusive areas. And the bounding box indicates the exclusive area According to the dot mask weight and bounding box mask weights The pseudo-labels are obtained by weighted fusion; the formula is as follows: ; in, This represents a pseudo-tag.
5. The camouflaged target detection method based on point supervision and pseudo-label optimization according to claim 4, characterized in that: In step S4, a global-local perception module is constructed to obtain global-local interaction enhancement features; Specifically: Step S41: Input the image into the backbone network, extract the initial features of the image, and construct a global-local perception module; Step S42: Based on the global-local perception module constructed in step S41; the global-local perception module includes a global perception module and a local perception module. The global perception module obtains global perception features, the local perception module obtains local perception features, and global-local interaction enhancement features are obtained based on the global perception features and the local perception features.
6. The camouflaged target detection method based on point supervision and pseudo-label optimization according to claim 5, characterized in that: In step S42, the global perception module obtains the global perception features; Specifically: Step S4201: The initial features of the image are processed by two parallel dilated convolutions with kernel sizes of 3×3 and 5×5 to extract spatial features under different receptive fields. The extracted spatial features are then added to the initial features of the image through a residual structure to obtain residual features. Step S4202: Input the residual features into an improved frequency domain self-attention mechanism to obtain the spatial domain residual features, and then transform the spatial domain residual features into frequency domain features through fast Fourier transform to obtain frequency domain query features, frequency domain matching features and frequency domain information features. Step S4203: Perform dimensional transformation on the obtained frequency domain query features, frequency domain matching features, and frequency domain information features to obtain dimensionally transformed frequency domain query features, frequency domain matching features, and frequency domain information features. Perform matrix multiplication on the dimensionally transformed frequency domain query features and frequency domain matching features to generate the transposed reference diagram A. f ; Step S4204, note the transpose diagram A. f Perform the calculation; the formula is as follows: ; in, Representative transpose attention diagram A f The real part of a complex number, conj represents the calculation of the conjugate value of the complex number; ; in, Representative transpose attention diagram A f The imaginary part of a complex number, where j represents the imaginary unit in complex number theory; Step S4205: Activate the real part and the imaginary part of the complex number of the transposed attention map using the flexible maximum activation function and then recombine them to obtain the intermediate attention map. Then, perform matrix multiplication on the intermediate attention map and the frequency domain information features after dimension transformation, and obtain the preliminary frequency enhancement features by inverse fast Fourier transform back to the spatial domain features. Step S4206: Frequency residual connection yields frequency enhancement features; the formula is as follows: ; in, Represents frequency enhancement characteristics, Represents an indicator function. Represents the Fast Inverse Fourier Transform. Modify the linear unit activation function. Represents Fast Fourier Transform, Represents residual characteristics, Represents matrix multiplication; Step S4207: The initial frequency enhancement feature and the frequency enhancement feature are concatenated in the channel dimension, and the channel dimension is transformed by a 1×1 convolution to obtain the frequency domain self-attention enhancement feature; Step S4208: The initial features of the image and the frequency domain self-attention enhancement features are added through a residual connection, and the frequency domain enhancement features are obtained through a layer normalization process. The frequency domain enhancement features are then input into a depth-separable convolutional group with three parallel convolutional kernels of 3×3, 5×5 and 7×7 to obtain the orientation enhancement features. Step S4209: The frequency domain enhancement features and the direction enhancement features are combined through a residual structure to obtain the global perception features.
7. The camouflaged target detection method based on point supervision and pseudo-label optimization according to claim 6, characterized in that: In step S42, the local sensing module obtains local sensing features; Specifically: Step S4210: The initial features of the image are fed into the depthwise convolution branches with three parallel convolution kernels of 1×3, 3×1, and 3×3. The depthwise convolution branches extract local features in three different directions, namely, the local features in the horizontal direction. Vertical local features Local features of the same polarity as those in directions other than horizontal and vertical. ; Step S4211: The extracted local features from three different directions are concatenated along the channel dimension and activated using the Gaussian error linear unit activation function, and then combined with the initial features of the image. The directional characteristics are obtained by adding them together; the formula is as follows: ; in, Representing directional characteristics, Conv 1x1 Represents a regular 1x1 convolution. Represents the splicing of channel dimensions; Step S4212: The orientation features are processed by a layer normalization and a feedforward neural network, and the processed orientation features are added to a residual structure to obtain local perceptual features.
8. The camouflaged target detection method based on point supervision and pseudo-label optimization according to claim 7, characterized in that: In step S42, global-local interaction enhancement features are obtained based on global perception features and local perception features; Specifically: Step S4213: Divide the global perception features obtained in step S4209 and the local perception features obtained in step S4212 into four equal parts along the channel dimension to obtain four sets of global perception features and four sets of local perception features. The size of the four sets of global perception features and the four sets of local perception features is C / 4×H×W; where C represents the number of channels, H represents the height, and W represents the width. Step S4214: After compressing the spatial dimension of the four sets of global perception features by global average pooling, the four sets of global perception features are concatenated to obtain four new sets of global perception features. After compressing the space of the four sets of local perception features by global average pooling, the four sets of local perception features are concatenated to obtain four new sets of local perception features. Step S4215: The four new sets of global perception features and the four new sets of local perception features are respectively calculated by the lightweight fully connected network to obtain four sets of normalized fusion weights; wherein the lightweight fully connected network includes two fully connected layers and a modified linear unit activation function. Step S4216: Using four sets of normalized fusion weights, the four sets of global perception features and the four sets of local perception features are weighted and fused to obtain four sets of global-local features, and the four sets of global-local features are spliced together in the channel dimension. By fusing global and local perceptual features using residual connections, global-local enhanced features are obtained; the formula is as follows: ; Where i represents the level and m represents the group number. Represents the global-local enhancement features of the i-th layer. Represents the fusion weights of the m-th group in the i-th layer. This represents the global perception feature of the m-th group in the i-th layer. This represents the local perceptual features of the m-th group in the i-th layer. This represents the global perception features of the i-th layer. Represents the local perceptual features of the i-th layer; Step S4217: After the global-local enhancement features are processed by the gated convolution mechanism, they are fused with the global-local enhancement features through residual connections to obtain global-local interactive enhancement features. These global-local interactive enhancement features are divided into four layers: the initial global-local interactive enhancement features, the second adjacent layer global-local interactive enhancement features, the third adjacent layer global-local interactive enhancement features, and the deep global-local interactive enhancement features; the formula is as follows: ; in, Represents global-local interaction enhancement features Representatives criticized normalization. This represents a 3×3 gated convolution.
9. The camouflaged target detection method based on point supervision and pseudo-label optimization according to claim 8, characterized in that: In step S5, the global-local interactive enhancement features obtained in step S4 are subjected to bidirectional cross-layer fusion to obtain fused features. The obtained fused features are then processed to obtain the prediction map; specifically: Step S51: Apply a 3x3 depthwise separable convolution to the deep global-local interaction enhancement features to obtain the separable features of the deep global-local interaction enhancement features; The separable features of the deep global-local interaction enhancement features are generated by passing a set of point convolution groups and a sigmoid activation function to produce a first feature weight map. The first feature weight map is then multiplied with the global-local interaction enhancement features of the third adjacent layer to obtain the first layer semantic enhancement features. Global flat pooling is performed on the global-local interaction enhancement features of the third adjacent layer. Then, a second feature weight map is generated through a set of 1x1 convolutions and sigmoid activation functions. The second feature weight map is multiplied with the separable features of the deep global-local interaction enhancement features to obtain the first layer of local spatial refinement features. The obtained first-layer semantic enhancement features and first-layer local spatial refinement features are aggregated to obtain the first-layer fusion features; Applying a 3x3 depthwise separable convolution to the first layer of fused features yields the separable features of the first layer of fused features; The separable features of the first layer fusion features are used to generate a third feature weight map through a set of point convolution groups and a sigmoid activation function. The third feature weight map is then multiplied with the global-local interaction enhancement features of the second adjacent layer to obtain the second layer semantic enhancement features. Global flat pooling is performed on the global-local interaction enhancement features of the second adjacent layer. Then, a fourth feature weight map is generated through a set of 1x1 convolutions and sigmoid activation functions. The fourth feature weight map is multiplied with the separable features of the first layer fusion features to obtain the second layer local spatial refinement features. The obtained second-layer semantic enhancement features and second-layer local spatial refinement features are aggregated to obtain the second-layer fusion features; Applying a 3x3 depthwise separable convolution to the second-layer fused features yields the separable features of the second-layer fused features; The separable features of the second layer of fusion features are used to generate the fifth feature weight map through a set of point convolution groups and a sigmoid activation function. The fifth feature weight map is then multiplied with the initial layer global-local interaction enhancement features to obtain the third layer of semantic enhancement features. Global flat pooling is performed on the initial global-local interaction enhancement features. Then, a sixth feature weight map is generated through a set of 1x1 convolutions and sigmoid activation functions. The sixth feature weight map is multiplied with the separable features of the second layer fusion features to obtain the third layer local spatial refinement features. The third-layer semantic enhancement features and the third-layer local spatial refinement features are aggregated to obtain the final layer of fused features; Step S52: The last layer of fused features is processed using 1x1 convolution, batch normalization, and sigmoid activation function to obtain the prediction map.
10. The camouflaged target detection method based on point supervision and pseudo-label optimization according to claim 9, characterized in that: In step S6, a camouflage target detection model is constructed based on the pseudo-labels obtained in step S3 and the prediction map obtained in step S5. The binary cross-entropy loss function and the contrastive loss function are used as the final loss functions to train the constructed camouflage target detection model; specifically: Step S61: Calculate the partial binary cross-entropy loss and contrast loss of the prediction graph, using the following formulas: ; in, Represents the binary cross-entropy loss. Representative pseudo-label The true category of a pixel location in an image. The predicted category represents the pixel location in the image from which the predicted map is located; ; in, Represents the contrast loss, where d represents the number of pixels in the image. The predicted image representing the image. The image represents a predicted image obtained after Gaussian blurring, color jittering, and random rotation. Step S62: Train the camouflage target detection model based on the final loss function.
Citation Information
Patent Citations
Camouflage target detection method based on attention mechanism and convolutional neural network
CN116228702A
Remote sensing image semantic segmentation method and device based on point annotation expansion network
CN118968046A
Single-point supervision infrared small target detection method based on mixed pseudo label generation
CN119516160A
Cited By
Structural integrity enhanced camouflage target detection method based on graffiti annotation
CN121661337A
Weak supervision data set labeling method and device, labeling equipment and storage medium
CN122244622A