Camouflage target detection method and system based on cross-scale global and local feature combination

By adopting a cross-scale global local feature combination method in camouflage object detection, using the multi-scale feature fusion and decoder subscale fusion, the problem of low detection accuracy of camouflage object in complex background is solved, and the detection effect of high precision and boundary integrity is achieved.

CN120125812AActive Publication Date: 2025-06-10HUNAN NORMAL UNIVERSITY

Patent Information

Application Number
CN202510623152.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-10
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Existing camouflage object detection methods are difficult to accurately detect camouflage objects in complex backgrounds, especially when the target is highly similar to the background, has no obvious edges or has a large amount of occlusion, the detection accuracy is low.

Method used

A camouflage object detection method based on cross-scale global local feature combination is adopted, and multi-scale features are extracted through pre-trained encoder, and feature enhancement and scale fusion are combined with multi-branch convolution blocks and decoder to achieve high-precision detection of camouflage targets.

Benefits of technology

Effectively extract important clues in images in complex environments, improve the accuracy and boundary integrity of camouflage target detection, and overcome the problem of low detection accuracy caused by variable or small target size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125812A_ABST
    Figure CN120125812A_ABST
Patent Text Reader

Abstract

The invention discloses a camouflage target detection method and system based on cross-scale global and local feature combination, and the method comprises the steps: extracting feature maps of M scales from an input image # imgabs0 # through an encoder, inputting a multi-branch convolution block to obtain enhanced features, inputting the enhanced features into a pre-trained decoder, and carrying out the multi-scale fusion to obtain final fusion features, a decoding result obtained through up-sampling is classified to obtain a camouflage target detection result; wherein the decoder is provided with a # imgabs1 # layer, any # imgabs2 # layer comprises # imgabs3 # decoder nodes, all the decoder nodes are subjected to up-sampling layer by layer and step by step, fusion is carried out through double convolution operation, and then the output features of the decoder nodes are obtained through multi-stage combined scanning feature fusion blocks. The invention aims to solve the detection difficulty when the camouflage target is highly similar to the background, the edge is not obvious or a large amount of shielding exists under the complex background, and overcome the problem that the camouflage target detection precision is low due to the fuzzy edge and the small target size.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a camouflaged target detection method and system based on the joint of cross-scale global and local features. Background Art

[0002] Camouflaged target detection aims to identify and segment imperceptible targets in a visual scene, especially those camouflaged objects that are highly similar to the background or occluded by the background. Its main purpose is to quickly and accurately detect these targets, which are widely used in multiple fields such as medical image analysis and ecological protection. In existing camouflaged target detection methods, the detection task of camouflaged targets is usually completed based on the use of local features or edge features of camouflaged targets. For example, the F2-EDNet model enhances multi-scale context features through a feature enhancement module, and combines an edge prediction branch guided by cross-layer features to extract edge features, thereby improving the accuracy of camouflaged target detection. However, in the face of complex backgrounds, too small camouflaged object sizes, or low contrast with the background edge, relying on edge features may lead to missed detections. The Vim model uses a bidirectional state space model (SSM) for data-dependent global visual context modeling, but ignores the preservation of local two-dimensional dependencies. Therefore, the camouflaged target detection method based on the joint scanning and fusion of cross-scale global and local features has become an important approach. To sum up, how to effectively utilize the local features and global features of images at different scales, accurately capture the dependencies at different distances of images, and achieve high-precision detection of camouflaged targets of different sizes in complex environments has become a key technical problem to be solved urgently. Summary of the Invention

[0003] The technical problem to be solved by the present invention: Aiming at the above problems of the prior art, a camouflaged target detection method and system based on the joint of cross-scale global and local features are provided. The present invention aims to solve the detection difficulties of camouflaged targets that are highly similar to the background, have unclear edges or a large number of occlusions in complex backgrounds, and overcome the problem of low accuracy of camouflaged target detection due to blurred edges and small target sizes.

[0004] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A camouflaged target detection method based on the joint of cross-scale global and local features, comprising the following steps: Input an image Use a pre-trained encoder to extract feature maps of M scales; Input the feature maps of M scales into a multi-branch convolutional block MDCS to obtain enhanced features; Input the enhanced features of M scales into a pre-trained decoder for multi-scale fusion to obtain the final fusion feature; Upsample the final fusion feature to obtain a decoded result, and classify the decoded result to obtain a camouflaged target detection result; Among them, the decoder has a total of layers, and any layer includes a total of decoder nodes. The first decoder node of any layer inputs enhanced features of the th scale. The remaining any decoder nodes then upsample the output features of the th decoder node of the th layer, and then splice them with the output features of the 1st to th decoder nodes of the th layer, and fuse them through a double convolution operation, and then obtain the output features of the decoder node through a multi-level joint scanning feature fusion block LVBlock. The processing of the input features by the joint scanning feature fusion block LVBlock includes: for the input features, probability is used to select the four directions with the highest preference probability from a preset variety of scanning methods. After the input features are tiled in sequence, they are input into a selective spatial state model SSM to obtain one-dimensional scanning sequences generated by scanning in four directions. The one-dimensional scanning sequences in the four directions are fused through an attention feature fusion block with channel attention and spatial attention to globally and locally fuse the features, and then spliced to obtain the output features of the decoder node.

[0005] Optionally, before using the pre-trained encoder to extract feature maps of M scales from the input image, it further includes the step of preprocessing the original image to obtain the input image: converting the original image into an RGB three-channel matrix to obtain an image input ; reorganizing the shape or size of the image input to a specified size to obtain the input image .

[0006] Optionally, the multi-branch convolution block MDCS includes three parallel convolution branches with different receptive fields, a feature fusion convolution layer, a residual connection convolution layer, and incorporates channel attention and spatial attention. The first branch includes a convolution layer for channel adjustment; the second branch includes convolution layer, convolution layer, convolution layer, and a deformable convolution; the third branch includes convolution layer, convolution layer, a convolutional layer and a deformable convolutional layer; the output features of two parallel convolutional branches containing the deformable convolutional layer are concatenated along the channel dimension through a feature fusion convolutional layer and then undergo feature fusion and channel adjustment through a 3×3 convolution, and then are input into a residual connection convolutional layer to be added to the output features of the convolutional branch without the dilated convolutional layer through residual connection and then pass through activation function processing, and an enhanced feature with the same spatial size as the input feature map and a compressed number of channels is obtained by combining channel attention and spatial attention; the output features of the three branches have different receptive fields, these features are concatenated along the channel dimension, and then passed through a convolutional module with a convolutional kernel size of for feature fusion and channel adjustment, the features after feature fusion and channel adjustment are added to the input feature map through residual connection, and then pass through activation function processing, and by combining channel attention and spatial attention, an enhanced feature with the same spatial size as the input feature map and 64 channels is finally obtained .

[0007] Optionally, the encoder is composed of a cascaded M-level backbone network, and each level of the backbone network is used to extract a feature map of one scale. An N-level adapter is inserted at the input end of the backbone network, and the input features of the backbone network are concatenated with the original input features of the backbone network after passing through the N-level adapter and then used as the input of the subsequent network of the backbone network. The adapter consists of a linear layer for downsampling, an activation function, a dropout layer, a linear layer for upsampling, and an activation function in sequence.

[0008] Optionally, the backbone network is a Hiera backbone network. The Hiera backbone network includes a normalization module, an attention module, a splicing module, a normalization module, a multi-layer perceptron MLP, and a splicing module connected in sequence. The first splicing module splices and outputs the output features of an attention module and the input features of the Hiera backbone network, and the second splicing module is used to splice the output features of the first splicing module and the output features of the multi-layer perceptron MLP to be used as the output features of the Hiera backbone network.

[0009] Optionally, the double convolution operation in the decoder includes two convolution operations, and after each convolution operation, batch normalization and activation function processing are performed in sequence; when using probability to select the four directions with the highest preference probability, the calculation function expression of the probability is: , where is the Regarding the scanning method in the hierarchical joint scanning feature fusion block LVBlock of probability, is the weight of the scanning method in the hierarchical joint scanning feature fusion block LVBlock ; is the weight of the scanning method in the hierarchical joint scanning feature fusion block LVBlock ; is a set of scanning methods composed of a plurality of preset scanning methods.

[0010] Optionally, the plurality of preset scanning methods include: progressive scanning: scanning the image row by row from left to right, reverse progressive scanning: scanning the image row by row from right to left, progressive column scanning: scanning the image column by column from top to bottom, reverse column scanning: scanning the image column by column from bottom to top, 2×2 local scanning: performing local scanning with a 2×2 window size, reverse 2×2 local scanning: performing reverse local scanning with a 2×2 window size, 7×7 local scanning: performing local scanning with a 7×7 window size, reverse 7×7 local scanning: performing reverse local scanning with a 7×7 window size.

[0011] Optionally, it further includes training an end-to-end image segmentation model composed of an encoder, a multi-branch convolution block MDCS, a decoder, an upsampling module for upsampling the final fused features to obtain a decoded result, and a classifier for classifying the decoded result, supervised by a true mask G, and the functional expression of the training loss function used during training is: , wherein, is the total loss function, is the weighted loss of the layer decoder node, is the binary cross-entropy loss of the layer decoder node, and there is: , , , wherein, and are the height and width of the input image respectively, represents the output image of the last decoder node of the layer, is the output image in the true mask G of the true value of the pixel at the position, is the output image in the predicted value of the pixel at the position, is the adjustment parameter, is the output image in the weight of the pixel at the position; is the category, , being 0 represents the background, being 1 represents the foreground; are the model parameters of the network model, represents given model parameters when the logarithm probability that the pixel at the position belongs to the category , is the set of pixels around the position; is the output image in the true mask G of the true value of the pixel at the position; is the indicator function, when the input is True, the value of the indicator function is 1, and when the input is False, the value of the indicator function is 0.

[0012] In addition, the present invention also provides a camouflaged target detection system based on the joint of cross-scale global and local features, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the camouflaged target detection method based on the joint of cross-scale global and local features.

[0013] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the camouflaged target detection method based on the joint of cross-scale global and local features through a processor.

[0014] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: Considering that the camouflage target is highly similar to the background, has unclear edges or there is a lot of occlusion, the detection of camouflaged objects in complex environments is more complex and challenging than traditional prominent object detection. The present invention introduces multi-scale information, provides representations of the original image at different resolutions, and fuses multi-scale features at the same resolution and multi-scale features at different resolutions through tight skip connections between decoder nodes, enabling the analysis and processing of images at different scales, and being able to fully extract important clues in the image in complex environments, effectively overcoming the problem of low detection accuracy of camouflage targets caused by variable or small sizes; The present invention fuses global information and local information, realizes the fusion of global features and local features under different scanning methods, can effectively discover the correlation and semantic information between different distances in the image, can more accurately locate the target, and improve the boundary integrity of the detection result; The present invention is applicable to the detection of concealed targets in complex environments, can effectively overcome the fuzzy edges and size changes of the target, and provides the possibility for the accurate recognition and positioning of concealed targets. The present invention can solve the problem of detecting camouflage targets in images where the edges of camouflage targets are unclear or the sizes are uncertain due to complex environments, realize the accurate detection of small-sized objects, and overcome the problem of low detection accuracy of camouflage targets caused by boundary and size changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic diagram of the basic process of the method according to an embodiment of the present invention.

[0016] Figure 2 It is a schematic diagram of the network structure of the encoder in an embodiment of the present invention.

[0017] Figure 3 It is a schematic diagram of the network structure of the adapter in an embodiment of the present invention.

[0018] Figure 4 It is a schematic diagram of the network structure of the decoder in an embodiment of the present invention.

[0019] Figure 5 It is a schematic diagram of the network structure of the decoder node in an embodiment of the present invention.

[0020] Figure 6 It is a schematic diagram of the network structure of the attention feature fusion block in an embodiment of the present invention.

[0021] Figure 7 It is a schematic diagram of the network structure of the multi-branch convolution block MDCS in an embodiment of the present invention.

[0022] Figure 8 It is a schematic diagram of the process of image enhancement in an embodiment of the present invention.

[0023] Figure 9 This is the example test effect of the present invention in the CAMO dataset. Among them, (a1)-(a8) are the target images on the CAMO test set respectively, and (b1)-(b8) are the corresponding mask images GT of (a1)-(a8); (c1)-(c8) are the segmentation results of the method of this embodiment for the corresponding target images of (a1)-(a8). Detailed implementation manners

[0024] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0025] As Figure 1 shown, the camouflaged target detection method based on the joint cross-scale global-local features in this embodiment includes the following steps: S1, using a pre-trained encoder to extract feature maps of M scales from the input image ; S2, inputting the feature maps of M scales into a multi-branch convolutional block MDCS to obtain enhanced features ; S3, inputting the enhanced features of M scales into a pre-trained decoder for multi-scale fusion to obtain the final fused feature

[0026] In this embodiment, before using a pre-trained encoder to extract feature maps of M scales from the input image, it further includes the step of preprocessing the original image to obtain the input image: converting the original image into an RGB three-channel matrix to obtain an image input ; re-organizing the shape or size to a specified size to obtain the input image . For example, in this embodiment, the size of the input image is scaled by bilinear interpolation, and the input label is scaled by bicubic interpolation to avoid the loss of edge information in the label. It should be noted that interpolation is a well-known data processing method, and the required interpolation algorithm can be adopted according to needs. For example, the bilinear interpolation algorithm adopted in this embodiment is an interpolation method based on the weighted average of known pixel values in a local area. It takes into account the changes in pixel values in the horizontal and vertical directions, so as to better retain the details and quality of the image. Specifically, bilinear interpolation will find the four pixel points closest to the target pixel position in the original image, and then perform a weighted average according to the pixel values of these four pixel points to obtain the value of the target pixel. First, interpolation is implemented on the X-axis. Along , find two points with the same abscissa as the insertion point and : , , wherein and are respectively the interpolation calculation results in the X-axis direction of points and . The coordinates of point are , The coordinates of are is the coordinate of the point to be inserted. ~ are the four points near the insertion point, and the coordinates correspond to . Then use the , obtained by the previous interpolation step to perform linear interpolation in the vertical axis direction and then interpolation in the Y-axis direction: , wherein is the value of the insertion point finally obtained by interpolation. For example, in this embodiment, the input image will be scaled to the specified size accordingly indicates that the number of channels is 3, the image size is , and the input image is obtained.

[0027] In this embodiment, the input image is represented by the feature maps of M scales extracted by the pre-trained encoder as . As an alternative embodiment, in this embodiment takes the value of 4. In step S2, the preprocessed image is input into the trained encoder to be encoded into features of 4 scales , and there is: , where the subscript represents the th scale, represents the number of channels of the th scale, and respectively represent the height and width of the feature map, and both take the value of 352 For example Figure 2As shown, the encoder in this embodiment is composed of a backbone network. N-level adapters are inserted at the input end of the backbone network to achieve efficient parameter fine-tuning. After passing through the N-level adapters, the input features of the backbone network are concatenated with the original input features of the backbone network and used as the input for the subsequent network of the backbone network, such as Figure 3 shown, the adapter consists of a linear layer for downsampling, an activation function, a dropout layer, a linear layer for upsampling, and an activation function connected in sequence. By wrapping each multi-scale block of the encoder into an adapter, the feature extraction ability of the encoder can be enhanced, and layer-scale features are output .

[0028] The backbone network in this embodiment can select the backbone network that can extract feature maps of M scales as needed. For example, as an optional implementation, the backbone network in this embodiment is the Hiera backbone network pre-trained by SAM2. The Hiera backbone network includes a normalization module, an attention module, a splicing module, a normalization module, a multi-layer perceptron MLP, and a splicing module connected in sequence. The first splicing module splices and outputs the output features of an attention module and the input features of the Hiera backbone network. The second splicing module is used to splice the output features of the first splicing module and the output features of the multi-layer perceptron MLP as the output features of the Hiera backbone network. Hiera is a new type of hierarchical vision transformer (Vit), which constructs the model only using standard Vit blocks and learns spatial biases through pre-training tasks (such as masked autoencoder MAE), rather than relying on complex architecture designs. Hiera adopts an efficient multi-scale design, capturing the details and global information of the image by using feature maps of different resolutions at different stages. In addition, other encoders can also be used.

[0029] As Figure 4 shown, the decoder in this embodiment has a total of layers, and any layer includes a total of decoder nodes. The first decoder node of any layer inputs the enhanced features of the th scale, as Figure 5 shown. For the rest, any decoder node upsamples the output features of the th decoder node of the th layer and then concatenates them with the output features of the 1st to th layer of the The output features of the decoder nodes are concatenated and then fused through a double convolution operation, and then passed through a multi-level joint scanning feature fusion block LVBlock to obtain the output features of the decoder nodes. The processing of the input features by the joint scanning feature fusion block LVBlock includes: for the input features, select the four directions with the highest preference probability from a preset variety of scanning methods with probability, tile the input features in sequence and input them into the selective spatio-temporal model SSM to obtain one-dimensional scanning sequences generated by scanning in four directions, and fuse the global and local features of the one-dimensional scanning sequences in four directions through an attention feature fusion block with channel attention and spatial attention, and then splice them to obtain the output features of the decoder nodes. In this embodiment, the image features are input into the decoder for upsampling layer by layer and feature fusion to obtain the fused features of each decoder node and the final output image features of the decoder ; there are layers in the decoder. The th layer of the decoder includes decoder nodes in total. Each decoder node is represented as , and the output features corresponding to the decoder nodes are denoted as , and the features are respectively input into the decoder nodes . For the multi-layer input features , they are upsampled layer by layer and feature fusion is performed to obtain the fused features of each decoder node, where represents the downsampling layer along the encoder, and represents the convolutional layer along the fully connected block of the skip connection. As Figure 3 shown, there are four layers in the decoder in this embodiment. The features , , , are respectively input into the decoder nodes , , , as the output features , , , of the corresponding decoder nodes. As Figure 5 shown, the non-first-layer decoder nodes all perform upsampling, feature connection, and joint scanning, including: first inputting the multi-layer features into the first node of the corresponding th layer of the decoder respectively as the output features of the node; inside the decoder, for the features Perform bilinear interpolation upsampling to double the width and height of the image, obtaining the upsampled features ; (In this embodiment, the output feature map sizes corresponding to the decoder nodes of layers 1 to 4 are ). The upsampled features are fused with the features of layer to obtain the intermediate features of the encoder node . The fusion process includes a double convolution operation, followed by batch normalization and activation function processing after each convolution, reducing the number of channels to 64; among them, the batch normalization function expression is: , In the above formula, is each pixel of the input feature map, is the input mean, is the input variance, is a small constant used to prevent the denominator from being zero. The activation function is expressed as: , In the above formula, is each pixel of the input feature map. The intermediate features are input into two-layer joint scanning feature fusion block LVBlock to obtain the node features at .

[0030] As Figure 6 shown, the attention feature fusion block includes a global branch and a local branch. The global branch refers to performing global average pooling on the entire input feature sequence, and then performing linear dimensionality reduction, so that the value of each channel represents the information of the entire feature map. The processing steps in the global branch include: Step 1, perform global average pooling on the input feature map to obtain a feature vector with a shape of . Among them, the function expression of global average pooling is: , In the above formula, is the feature vector of the input feature map at position , is the feature vector after global average pooling.

[0031] Step 2, pass through a fully connected layer to output a shape of , where is a scaling ratio, which is defaulted to 0.125 in this embodiment to reduce the number of parameters. The shape of the decoder node feature map. Then, through activation function activation processing; Step 3: The feature vector after activation by the activation function is linearly dimension-reduced through a fully connected layer, and the feature vector is mapped back to the shape of, and when using the activation function to generate channel attention weights.

[0032] The local branch refers to linearly dimension-reducing each token in the feature sequence to retain local features. The processing steps in the local branch include: Step 1: The input feature map passes through a fully connected layer to obtain a feature map with the shape of , followed by the activation function; Step 2: For each token in the local features output by the activation function in the local branch, the global features output in Step 2 of the global branch are concatenated with the local features output in Step 2 of the local branch, so that each token in the concatenated global features has both global and local features; Step 3: Calculate the spatial attention weights from the concatenated global features.

[0033] Finally, multiply the channel attention and the spatial attention to obtain the final attention weights. Apply the final attention weights to the input feature map to obtain the output feature with the shape of .

[0034] In this embodiment, the double convolution operation in the decoder includes two convolution operations, and after each convolution operation, batch normalization and activation function processing are performed in sequence.

[0035] The joint scanning feature fusion block LVBlock is a selective spatial state model with four scanning directions and spatial and channel attention feature fusion blocks. In this embodiment, when the joint scanning feature fusion block LVBlock uses probability to select the four directions with the highest preference probability, the calculation function expression of the probability is: , where is the probability of the scanning method in the th-level joint scanning feature fusion block LVBlock, is the The weights regarding the scanning method in the hierarchical joint scanning feature fusion block LVBlock are the weights regarding the scanning method in the th hierarchical joint scanning feature fusion block LVBlock are a set of scanning methods composed of multiple preset scanning methods.

[0036] In this embodiment, the multiple preset scanning methods include: progressive scanning: scanning the image row by row from left to right, reverse progressive scanning: scanning the image row by row from right to left, column scanning: scanning the image column by column from top to bottom, reverse column scanning: scanning the image column by column from bottom to top, 2×2 local scanning: performing local scanning with a 2×2 window size, reverse 2×2 local scanning: performing reverse local scanning with a 2×2 window size, 7×7 local scanning: performing local scanning with a 7×7 window size, reverse 7×7 local scanning: performing reverse local scanning with a 7×7 window size. Through the 8 scanning methods in the preset set of scanning methods, local features and global features are respectively captured, and the four directions with the highest preference probability are selected from the 8 scanning methods at this layer, thereby constructing a search space , where represents the number of LVBlocks. A differentiable search mechanism is set up, and probabilities are used to represent the selection preference for each direction, and finally the four directions with the highest probability are selected as the scanning directions at this layer to generate corresponding one-dimensional scanning sequences. The feature sequences in the four directions are respectively input into the spatial and channel attention feature fusion block for feature merging to obtain the output feature . Feature merging can be expressed as: , where represents the th layer of the joint scanning feature fusion block LVBlock, represents the scanning result obtained by weighted summation of the scanning results in the four directions of the th layer, represents the probability of the th layer and the th selected scanning direction, represents the output obtained by scanning the input feature according to the specific scanning direction . For the feature merging in the joint scanning feature fusion block LVBlock, a combination of channel attention and spatial attention is adopted. Taking the output feature of the th layer as the input feature of the th layer, repeating the layer operations of the joint scanning feature fusion block LVBlock, the final node output feature can be finally obtained Note that the space state model (SSM) used in this embodiment is specifically selected as the Selective State Space Model. The Selective State Space Model introduces an input-dependent parameterization and a selective mechanism compared with the traditional space state model, solving the limitations of the traditional model in dealing with complex dynamic systems, so that the model can more flexibly adapt to the changes of input data. Finally, for the node features at Repeat the decoding operations of each decoder node, and finally the features of each decoder node can be obtained; the decoder finally outputs the fused feature , denoted as the output feature . In this embodiment, the decoder output feature is input into a two-dimensional convolutional layer to adjust the number of channels to obtain the decoding result and classify to obtain the detection result of the camouflage target.

[0037] In step S3 of this embodiment, M image features of multiple scales are respectively input into the multi-branch convolutional block MDCS to obtain image features ; in this embodiment, M = 4. As Figure 7 shown, the multi-branch convolutional block MDCS includes three parallel convolutional branches with different receptive fields, a feature fusion convolutional layer, a residual connection convolutional layer, and integrates channel attention and spatial attention. The first branch contains a convolutional layer for channel adjustment; the second branch contains convolutional layer, convolutional layer, convolutional layer, and a deformable convolution; the third branch contains convolutional layer, convolutional layer, convolutional layer, and a deformable convolutional layer; the output features of the two parallel convolutional branches containing the deformable convolutional layer are concatenated along the channel dimension through the feature fusion convolutional layer and then fused and channel-adjusted through a 3×3 convolution, and then input into the residual connection convolutional layer to be added to the output feature of the convolutional branch without the dilated convolutional layer through the residual connection and processed by the activation function, and an enhanced feature with the same spatial size as the input feature map and compressed channel number is obtained by combining channel attention and spatial attention. The output features of the three branches have different receptive fields, these features are concatenated along the channel dimension, and then passed through a convolutional module with a convolutional kernel size of for feature fusion and channel adjustment. The fused feature is added to the input feature map through the residual connection and passed through Activation function processing, by combining channel attention and spatial attention, finally obtains enhanced features with the same spatial size as the input feature map and 64 channels. .

[0038] This embodiment also includes training an end-to-end image segmentation model composed of an encoder, a multi-branch convolutional block MDCS, a decoder, an upsampling module for upsampling the final fused features to obtain a decoded result, and a classifier for classifying the decoded result, supervised by a true mask G, and the functional expression of the training loss function used during training is: , where, is the total loss function, is the weighted loss of the layer decoder node, is the binary cross-entropy loss of the layer decoder node, and there is: , , , where, and are the height and width of the input image respectively, represents the output image of the last decoder node of the layer, is the true value of the pixel at the position in the true mask G of the output image , is the predicted value of the pixel at the position in the output image , is the adjustment parameter, is the weight of the pixel at the position in the output image ; is the category, , being 0 represents the background, being 1 represents the foreground; are the model parameters of the network model, is the logarithmic probability that the pixel at the position belongs to the category when the given model parameters are , is A set of pixel points around the position; For the output image In the true mask G of The true value of the pixel point at the position; Is an indicator function. When the input is True, the value of the indicator function is 1, and when the input is False, the value of the indicator function is 0. In the training phase, the segmentation output of the model Adopts a deep supervision mechanism, and inputs the segmentation output of each stage Into the loss function and is supervised by the true mask G. The loss function combines weighted Loss and binary cross-entropy Loss as the training objective. In the numerator calculation of, the weighted intersection is calculated, that is, the weighted sum of the pixel points correctly predicted as the positive class by the model, and the denominator calculates the weighted union, that is, the weighted sum of the pixel points predicted as the positive class or with the true label as the positive class by the model; in this embodiment, The value of is 4, , , Respectively correspond to the decoder nodes , , Of the output image.

[0039] It should be noted that, as Figure 8 Shown, in the training phase, this embodiment also includes samples for the input image To perform enhancement processing to obtain new samples of the input image , including: S101, set a random probability , if Then input the image into the sample of the image Perform a mirror symmetry transformation along the horizontal axis of the image; if Then do not flip the image and directly return the original data; for example, in this embodiment , perform a mirror symmetry transformation on the image to obtain the enhanced image ; S102, set a random probability , if Then perform a vertical flip transformation on the image along the horizontal axis of the image; if Then do not flip the image and directly return the original data; for example, in this embodiment , perform a mirror symmetry transformation on the image to obtain the enhanced image ; S103, convert the enhanced image To a tensor , and the pixel value range from Normalize to , the data format from Convert to ; for example, in this embodiment, an image of size is converted into a tensor of size ; S104, standardize each channel of the image using predefined mean and standard deviation. For example, in this embodiment, the mean is set to and the standard deviation is set to . Obtain the input feature image

[0040] To verify the effectiveness of the camouflage target detection method based on the joint cross-scale global-local features in this embodiment, in this embodiment, the optimal image segmentation model is obtained through model training and testing on the well-known CAMO dataset. The CAMO dataset is designed specifically for the camouflaged object segmentation task. The camouflaged object images consist of 1250 images (1000 images in the training set and 250 images in the test set). The existing F2-EDNet method is used as a comparison for the method in this embodiment, and the structural similarity metric ( ), enhancement metric ( ), harmonic mean of precision and recall ( ), and mean absolute error ( ) are used to measure the recognition accuracy. The higher the structural similarity metric, enhancement metric, and adaptive weight index, the better the segmentation effect, while the smaller the mean absolute error value indicates the smaller the difference between the predicted result and the true result; the structural similarity metric mainly evaluates the performance by comparing the structural similarity between the predicted segmentation result and the true segmentation; the enhancement metric is an index used to evaluate the performance of image segmentation algorithms, especially for the measurement of boundary accuracy and region similarity. It combines concepts such as structural similarity, region similarity, and mutual information in information theory, and can more comprehensively evaluate the accuracy of image segmentation results; the harmonic mean of precision and recall takes into account the imbalance of sample categories; the mean absolute error is the average of the absolute values of the differences between the predicted values and the true values. In the camouflage target detection task, it measures the average distance between the predicted boundary position and the true boundary position. The final results are shown in Table 1.

[0041] Table 1 Comparison table of recognition results of the method in this embodiment and the F2-EDNet method on the CAMO dataset

[0042] As can be seen from Table 1, for the camouflaged target segmentation of the method in this embodiment, the structural similarity metric ( ), enhancement metric ( ), harmonic mean of precision and recall ( ), and mean absolute error ( ) are all higher than the existing F2-EDNet method. Figure 9 This is the example test effect on the CAMO dataset in this embodiment. Among them, (a1)-(a8) are the target images on the CAMO test set respectively, and (b1)-(b8) are the corresponding mask images GT of (a1)-(a8); (c1)-(c8) are the segmentation results of the method of this embodiment for the target images corresponding to (a1)-(a8) respectively. Among them, the mask image GT is a mask image corresponding to the image to be segmented, where the salient region is marked white or with a high brightness value, while the non-salient region is marked black or with a low brightness value. This mask image provides the correct annotation of the salient region for the model, which is used to train and evaluate the performance of the model. It can be seen that the method of this embodiment has good detection effects in the cases where the target size is small, the light is blurred, the target is in a complex environment and hidden in the environment. It fully explores the correlation between local features, extracts important semantic information in the input, can locate the target more accurately, and improves the boundary integrity of the segmentation result.

[0043] In summary, in the camouflaged target detection task, the changes in target categories and the diversity of complex scenes make the detection of target objects more challenging compared to traditional salient object detection or other segmentation tasks. The camouflaged target detection method based on the joint of cross-scale global and local features in this embodiment considers the different expressions of the scene at different scales, and adopts a multi-branch scanning method to capture the local dependence relationship of the image, which helps to improve the understanding and judgment of target objects. First, the multi-scale information of the image is extracted through a hierarchical decoder, and then a multi-branch convolutional block is used to enhance the feature representation. Next, in the hierarchical cross-scale decoder, the cross-scale semantics are further mined through the strategies of grouping, mixing and fusion. The camouflaged target detection method based on the joint scanning and fusion of cross-scale global and local features in this embodiment simultaneously utilizes the correlation between different scales, extracts semantic information according to the correlation between channels and the selective scanning mechanism, and provides a precise and effective solution for camouflaged target detection in complex environments.

[0044] In addition, this embodiment also provides a camouflaged target detection system based on the joint of cross-scale global and local features, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the camouflaged target detection method based on the joint of cross-scale global and local features.

[0045] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the camouflaged target detection method based on the joint of cross-scale global and local features through a processor.

[0046] In addition, this embodiment also provides a computer program product, including a computer program or instruction, which is programmed or configured to execute the camouflaged target detection method based on the joint of cross-scale global and local features through a processor.

[0047] Those skilled in the art should understand that the technical solution provided by the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0048] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A camouflaged target detection method based on cross-scale global and local feature combination, characterized in that: The steps include: The input image Use the pre-trained encoder to extract feature maps of M scales; The feature maps of M scales are input into the multi-branch convolution block MDCS to obtain enhanced features; The enhanced features of M scales are input into the pre-trained decoder for scale fusion to obtain the final fusion features; The final fusion feature is upsampled to obtain the decoding result, and the decoding result is classified to obtain the camouflaged target detection result; Among them, the decoder has Layer and any The layers include decoder nodes, any The first decoder node of the layer inputs The enhanced features of the first scale and the rest of the The decoder node will Layer The output features of the decoder nodes are upsampled and then combined with the Layer 1~ The output features of the decoder nodes are concatenated and fused through a double convolution operation, and then through a multi-level joint scanning feature fusion block LVBlock to obtain the output features of the decoder node. The processing of the input features by the joint scanning feature fusion block LVBlock includes: using a preset multiple scanning methods for the input features The four directions with the highest probability of probability selection are selected. The input features are flattened in sequence and then input into the selective spatial state model SSM to obtain a one-dimensional scanning sequence generated by scanning in four directions. The one-dimensional scanning sequences in four directions are fused with global and local features through an attention feature fusion block containing channel attention and spatial attention, and then spliced ​​to obtain the output features of the decoder node.

2. The camouflaged target detection method based on cross-scale global and local feature combination according to claim 1 is characterized in that: Before extracting the feature maps of M scales from the input image using the pre-trained encoder, the step of pre-processing the original image to obtain the input image is also included: converting the original image into an RGB three-channel matrix to obtain an image input ; Input the image Reshape or resize the input image to the specified size .

3. The camouflaged target detection method based on cross-scale global and local feature combination according to claim 1 is characterized in that: The multi-branch convolutional block MDCS includes three parallel convolutional branches with different receptive fields, a feature fusion convolutional layer, a residual connection convolutional layer and a fusion of channel attention and spatial attention. The first branch contains a The convolutional layer is used for channel adjustment; the second branch contains The convolutional layer, The convolutional layer, The convolutional layer and a deformable convolution; the third branch contains The convolutional layer, The convolutional layer, The output features of the two parallel convolution branches including the deformable convolution layer are concatenated along the channel dimension through the feature fusion convolution layer, and then fused and adjusted through 3×3 convolution. The output features of the convolution branch without the hole convolution layer are then input into the residual connection convolution layer and added through the residual connection. Activation function processing, by combining channel attention and spatial attention, obtains enhanced features with the same spatial size as the input feature map and compressed channel number; the output features of the three branches have different receptive fields, these features are spliced ​​along the channel dimension, and then passed through a convolution kernel size of The convolution module performs feature fusion and channel adjustment. The features after feature fusion and channel adjustment are added to the input feature map through residual connection and then Activation function processing, by combining channel attention and spatial attention, finally obtains an enhanced feature with the same spatial size as the input feature map and 64 channels ,in is the scale quantity.

4. The camouflaged target detection method based on cross-scale global and local feature combination according to claim 1 is characterized in that: The encoder is composed of a cascaded M-level backbone network, and each level of the backbone network is used to extract a feature map of a scale. The input end of the backbone network is inserted with an N-level adapter, and the input features of the backbone network are spliced ​​with the original input features of the backbone network after passing through the N-level adapter as the input of the subsequent network of the backbone network. The adapter is composed of a linear layer for downsampling, a activation function, a random dropout layer, a linear layer for upsampling, and a Activation function composition.

5. The method for detecting camouflaged targets based on cross-scale global and local feature combination according to claim 4, characterized in that: The backbone network is a Hiera backbone network, which includes a normalization module, an attention module, a splicing module, a normalization module, a multi-layer perceptron MLP and a splicing module connected in sequence. The first splicing module splices and outputs the output features of an attention module and the input features of the Hiera backbone network, and the second splicing module is used to splice the output features of the first splicing module and the output features of the multi-layer perceptron MLP as the output features of the Hiera backbone network.

6. The camouflaged target detection method based on cross-scale global and local feature combination according to claim 1 is characterized in that: The double convolution operation in the decoder includes two convolution operations, and each convolution operation includes batch normalization and Activation function processing; the use of When the probability choice prefers the four directions with the highest probability, The calculation function expression of probability is: , in, For the Scanning method in LVBlock of Probability, For the Scanning method in LVBlock The weight of For the Scanning method in LVBlock The weight of A scanning mode collection consisting of multiple preset scanning modes.

7. The camouflaged target detection method based on cross-scale global and local feature combination according to claim 1 is characterized in that: The preset multiple scanning modes include: line-by-line scanning: scanning the image line by line from left to right, reverse line-by-line scanning: scanning the image line by line from right to left, column-by-column scanning: scanning the image column by column from top to bottom, reverse column-by-column scanning: scanning the image column by column from bottom to top, 2×2 local scanning: local scanning with a 2×2 window size, reverse 2×2 local scanning: reverse local scanning with a 2×2 window size, 7×7 local scanning: local scanning with a 7×7 window size, reverse 7×7 local scanning: reverse local scanning with a 7×7 window size.

8. The method for detecting camouflaged targets based on cross-scale global and local feature combination according to claim 1, characterized in that: The end-to-end image segmentation model including an encoder, a multi-branch convolution block MDCS, a decoder, an upsampling module for upsampling the final fusion feature to obtain a decoding result, and a classifier for classifying the decoding result is trained by supervision by a real mask G, and the function expression of the training loss function used in the training is: , in, is the total loss function, is the scale quantity, For the Weights of layer decoder nodes loss, For the Binary cross entropy of layer decoder nodes Loss, and there are: , , , in, and The input images are The height and width of Indicates The output image of the last decoder node of the layer, For output image The real mask G The true value of the pixel at the position, For output image middle The predicted value of the pixel at the position, To adjust the parameters, For output image middle The pixel weight of the position; For categories, , 0 means background, 1 indicates foreground; are the model parameters of the network model, To represent the given model parameters hour The pixel at the position belongs to the category The logarithmic probability of for The set of pixels around the location; For output image The real mask G The true value of the pixel at the position; It is an indicator function. When the input is True, the value of the indicator function is 1. When the input is False, the value of the indicator function is 0.

9. A camouflaged target detection system based on cross-scale global and local feature combination, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the camouflaged target detection method based on cross-scale global-local feature combination as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the camouflaged target detection method based on cross-scale global and local feature combination as described in any one of claims 1 to 8 through a processor.

Citation Information

Patent Citations

  • Remote sensing image saliency target detection method

    CN118015332A

  • Cross-modal multi-scale unmanned aerial vehicle camouflage target detection method and system and medium

    CN119169004A

  • Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment

    WO2024230038A1

Cited By

  • Electron density map generation method and system, electronic equipment and storage medium

    CN120833398A

  • Camouflage object semantic segmentation method and system based on two-stage edge guidance, and medium

    CN121544883A

  • A two-stage edge-guided semantic segmentation method, system, and medium for camouflaged objects.

    CN121544883B

  • Multi-target adaptive optimization camouflage target segmentation lightweight method and system

    CN122090072A

  • A lightweight method and system for multi-objective adaptive optimization of camouflaged target segmentation

    CN122090072B