Method and device for detecting camouflaged objects
By using the encoder structure and refinement perception module processing formed by Transformer blocks, the problem of poor local details and edge features of the existing camouflage object detection model is solved, and the detection efficiency is improved through the camouflage object detection model.
Patent Information
- Application Number
- CN202411339656.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing camouflage object detection model is based on convolutional neural networks, resulting in poor performance in capturing local details and edge features of camouflage objects, detection accuracy needs to be improved, and training data is scarce and difficult to obtain.
The encoder structure composed of Transformer blocks is used for feature extraction and fusion, combined with the processing of the refinement perception module, and the training data is generated through a camouflage transformation strategy, and the model is optimized using binary cross entropy loss and cross-match loss function.
It improves the accuracy and efficiency of camouflage object detection, reduces information loss during feature extraction, and enhances the ability to capture local details and edge features of camouflage object.
Smart Images

Figure CN119273981B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of disguised object detection, and in particular to a disguised object detection method and device. Background Art
[0002] The purpose of camouflaged object detection technology is to identify objects in an image that are difficult to distinguish due to their high degree of integration with the background. In the related art, camouflaged objects in an image are often detected through a preset camouflaged object detection model. However, most of the camouflaged object detection models in the related art are based on convolutional neural networks, which leads to poor performance of the related art in capturing local details and edge features of camouflaged objects, and the detection accuracy of camouflaged objects in the related art needs to be improved. Summary of the invention
[0003] The embodiments of the present application provide a disguised object detection method and device for improving the detection accuracy of a disguised object.
[0004] On the one hand, an embodiment of the present application provides a method for detecting a disguised object, comprising the following steps:
[0005] Acquire a test image containing a camouflaged object;
[0006] Performing disguised object detection on the image to be tested by using a disguised object detection model to obtain a contour image of the disguised object;
[0007] The disguised object detection model is obtained by training a preset image sample set, and the disguised object detection model includes:
[0008] An input layer, used for preprocessing the image to be tested to obtain an input sequence;
[0009] An encoder structure, the encoder structure comprising a plurality of encoders connected in sequence, each of the encoders comprising at least one Transformer block, the encoder structure being used to extract features from the input sequence through the plurality of Transformer blocks to obtain a plurality of encoding feature sequences;
[0010] A coding fusion module, used for fusing a plurality of coding feature sequences to obtain a fused feature sequence;
[0011] A decoder structure, used for decoding the fused feature sequence through at least one Transformer block to obtain a decoded feature sequence;
[0012] The refinement perception module is used to perform refinement perception processing on the decoded feature sequence and the encoding feature sequence of the last encoder in the encoder structure to obtain a contour map of the camouflaged object.
[0013] On the other hand, an embodiment of the present application provides a disguised object detection device, including:
[0014] An acquisition module, used for acquiring an image to be tested containing a camouflaged object;
[0015] a first processing module, equipped with a disguised object detection model, configured to perform disguised object detection on the image to be tested by using the disguised object detection model to obtain a contour image of the disguised object;
[0016] The disguised object detection model is obtained by training a preset disguised object image sample set, and the disguised object detection model includes:
[0017] An input layer, used for preprocessing the image to be tested to obtain an input sequence;
[0018] An encoder structure, the encoder structure comprising a plurality of encoders connected in sequence, each of the encoders comprising at least one Transformer block, the encoder structure being used to extract features from the input sequence through the plurality of Transformer blocks to obtain a plurality of encoding feature sequences;
[0019] A coding fusion module, used for fusing a plurality of coding feature sequences to obtain a fused feature sequence;
[0020] A decoder structure, used for decoding the fused feature sequence through at least one Transformer block to obtain a decoded feature sequence;
[0021] The refinement perception module is used to perform refinement perception processing on the decoded feature sequence and the encoding feature sequence of the last encoder in the encoder structure to obtain a contour map of the camouflaged object.
[0022] The beneficial effects of the present application are: providing a camouflaged object detection method and device, first, obtaining a test image containing a camouflaged object; then, performing camouflaged object detection on the test image through a camouflaged object detection model to obtain a contour map of the camouflaged object; in the camouflaged object detection model, the input layer preprocesses the test image to obtain an input sequence, the encoder structure extracts features from the input sequence through multiple transformers to obtain multiple coding feature sequences, the coding fusion module fuses the multiple coding feature sequences to obtain a fused feature sequence, the decoder structure decodes the fused feature sequence through at least one transformer to obtain a decoded feature sequence, and the refined perception module refines the decoded feature sequence and the coding feature sequence of the last encoder in the encoder structure to obtain a contour map of the camouflaged object. The present application can effectively improve the detection accuracy of camouflaged objects.
[0023] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flow chart of the camouflaged object detection method provided by the present application;
[0025] Figure 2 is a structural diagram of the camouflaged object detection model provided by this application;
[0026] Figure 3 is a schematic diagram of the overlapping process provided by the present application;
[0027] Figure 4 It is a structural diagram of the coding fusion module provided by this application;
[0028] Figure 5 is a structural diagram of the refined perception module provided by this application;
[0029] Figure 6 It is a schematic diagram of the discrete wavelet transform provided by this application;
[0030] Figure 7 It is a schematic diagram of the camouflage transformation strategy provided by this application;
[0031] Figure 8 It is a structural diagram of the camouflaged object detection device provided in this application. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0033] The present application is further described below in conjunction with the accompanying drawings and specific embodiments. The described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.
[0034] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0036] Camouflaged Object Detection (COD) technology aims to identify objects in an image that are difficult to distinguish due to their high degree of integration with the background. These objects are often called camouflaged objects. They usually hide their own features by imitating the surrounding environment, and camouflaged object detection technology can overcome this visual confusion. Due to its potential applications in various fields, camouflaged object detection technology is considered a key processing step in Computer Vision (CV) tasks, including industrial defect detection and agricultural pest detection.
[0037] In the task of camouflaged object detection, camouflaged objects are often small and difficult to detect. In response to this, related technologies use multi-scale feature fusion to achieve more accurate camouflaged object detection, which aims to capture easily overlooked camouflaged object targets by integrating feature information of different scales. For example, F3Net introduces a cross-feature module to fuse camouflaged object features of different scales to extract the shared parts between different features, suppress each other's background noise while supplementing each other's missing parts, thereby improving the detection accuracy of camouflaged objects. For another example, PFRNet proposes an adaptive feature fusion module that directly extracts and fuses the guidance information for camouflaged object detection from image features, thereby achieving accurate camouflaged object detection. For another example, FDNet designs a feature grafting module based on cross-attention, grafting the features extracted from the transformer branch to the convolutional neural network (CNN) branch, and then aggregates different features in the feature fusion module, using the aggregated features to achieve accurate camouflaged object detection.
[0038] However, most of the camouflaged object detection models in the related technologies are based on convolutional neural networks, which have limited receptive fields. This leads to poor performance of the related technologies in capturing local details and edge features of camouflaged objects, and the detection accuracy of camouflaged objects in the related technologies needs to be improved.
[0039] In addition, the training data for disguised object detection is often scarce and difficult to obtain. Due to the lack of training data for disguised object detection, some related technologies increase the training data for disguised object detection by manually annotating the disguised objects in the images, but this method increases the data preparation cost and time consumption; another part of the related technology uses the existing large model to generate a large number of images containing disguised objects to increase the training data for disguised object detection, but generating these images containing disguised objects requires a lot of time and generates a lot of calculations, which increases the computing load of the terminal, and the quality of the generated images is usually not enough to serve as training data for disguised object detection.
[0040] In view of this, embodiments of the present application provide a disguised object detection method and apparatus, aiming to improve the detection accuracy of disguised objects and achieve effective expansion of training data for disguised object detection.
[0041] First, the implementation steps of the camouflaged object detection method provided by the embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0042] The camouflaged object detection method provided in the embodiment of the present application can be applied to a terminal, a server, or software running in a terminal or a server. The terminal can be a tablet computer, a laptop computer, a desktop computer, etc., but is not limited to this. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. In addition, the server can also be a node server in a blockchain network, but is not limited to this. Among them, blockchain is a new application model of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, and encryption algorithm.
[0043] Reference Figure 1 , Figure 1 is a flow chart of the camouflaged object detection method provided by the present application, Figure 2 : is a structural diagram of a disguised object detection model provided by the present application. The disguised object detection method may include the following steps S101-S102:
[0044] S101, obtaining an image to be tested containing a camouflaged object;
[0045] S102, performing disguised object detection on the image to be tested by using a disguised object detection model to obtain a contour map of the disguised object;
[0046] The disguised object detection model is obtained by training a preset image sample set, and the disguised object detection model includes:
[0047] The input layer is used to preprocess the image to be tested to obtain an input sequence;
[0048] An encoder structure, the encoder structure includes a plurality of encoders connected in sequence, each encoder includes at least one Transformer block, and the encoder structure is used to extract features from an input sequence through the plurality of Transformer blocks to obtain a plurality of encoding feature sequences;
[0049] A coding fusion module is used to fuse multiple coding feature sequences to obtain a fused feature sequence;
[0050] A decoder structure, used for decoding the fused feature sequence through at least one Transformer block to obtain a decoded feature sequence;
[0051] The refinement perception module is used to perform refinement perception processing on the decoded feature sequence and the encoded feature sequence of the last encoder in the encoder structure to obtain a contour map of the camouflaged object.
[0052] In the embodiment of the present application, first, an image to be tested is obtained, and the image to be tested refers to an image containing a camouflaged object to be tested; then, a camouflaged object detection model is used to perform camouflaged object detection on the image to be tested, and a contour map of the camouflaged object is obtained, thereby realizing detection of the camouflaged object. The camouflaged object detection model is obtained by training a preset image sample set, and the camouflaged object detection model may include an input layer, an encoder structure, a coding fusion module, a decoder structure, and a refinement perception module.
[0053] Specifically, in the camouflaged object detection model, when the image to be tested is input into the camouflaged object detection model, the image to be tested is first preprocessed through the input layer to convert the image to be tested into an input sequence, so as to reduce the computational complexity of the camouflaged object detection model and reduce redundant information in the image data. Secondly, the input sequence is feature extracted through multiple Transformer blocks of the encoder structure to extract the encoded feature sequence representing different feature dimensions. Then, the encoded feature sequences of different feature dimensions are fused through the encoding fusion module to obtain a fused feature sequence, which is intended to accurately capture the local detail information and global detail information of the camouflaged object and reduce the loss of local detail information. Then, the fused feature sequence is decoded through at least one Transformer block of the decoder structure to obtain a decoded feature sequence. Finally, the decoded feature sequence and the encoded feature sequence of the last encoder in the encoder structure are refined and perceived through the refinement perception module to accurately capture the edge features of the camouflaged object, thereby obtaining a contour map of the camouflaged object. It can be seen that compared with the camouflaged object detection model based on convolutional neural network in the related art, the embodiment of the present application can not only capture the local detail information and global detail information of the camouflaged object more comprehensively and accurately, reduce the risk of losing local detail information in the feature extraction process, but also accurately capture the edge features of the camouflaged object, thereby improving the detection accuracy of the camouflaged object and having high availability.
[0054] In the above step S101, an image to be tested is obtained as an input of the disguised object detection model. The image to be tested refers to an image containing a disguised object to be tested, and the disguised object is a detection object in the embodiment of the present application.
[0055] The above-mentioned camouflage objects can be set according to actual conditions, and the embodiments of the present application do not make specific limitations on this.
[0056] In the above step S102, after obtaining the image to be tested, the image to be tested is input into the disguised object detection model, and the disguised object detection model performs disguised object detection on the image to be tested to obtain a contour map of the disguised object, thereby realizing detection of the disguised object.
[0057] The above-mentioned camouflaged object detection model refers to a neural network model trained by a preset image sample set.
[0058] The above-mentioned camouflaged object detection model may include but is not limited to an input layer, an encoder structure, a coding fusion module, a decoder structure and a refined perception module. The input layer is mainly used to preprocess the image to be tested to obtain an input sequence; the encoder structure is mainly used to extract features from the input sequence through at least one Transformer block to obtain at least one coding feature sequence; the coding fusion module is mainly used to fuse at least one coding feature sequence to obtain a fused feature sequence; the decoder structure is mainly used to decode the fused feature sequence through at least one Transformer block to obtain a decoded feature sequence; the refined perception module is mainly used to refine the decoding feature sequence and at least one coding feature sequence to obtain a contour map of the camouflaged object.
[0059] The implementation and principle of the above-mentioned camouflaged object detection model will be further explained below.
[0060] In some implementations, in the input layer, the preprocessing of the image to be tested to obtain the input sequence may include:
[0061] Perform patch processing on the image to be tested to obtain multiple pixel blocks;
[0062] Vectorize multiple pixel blocks to obtain a pixel vector sequence;
[0063] The pixel vector sequence and the preset position code are added to obtain an input sequence.
[0064] In this embodiment, first, the image to be tested is patched to divide the image to be tested into multiple pixel blocks (Patch); then, the multiple pixel blocks are vectorized to flatten each pixel block into a corresponding vector (Token) and concatenate all vectors into a pixel vector sequence; finally, the pixel vector sequence and the preset position code are added to obtain an input sequence. In this way, by converting the image to be tested into a vector expression of natural language, the computational complexity of the camouflage object detection model can be effectively reduced and the redundant information in the image data can be reduced.
[0065] The above pixel blocks are all two-dimensional pixel blocks, and the sizes of all pixel blocks are the same.
[0066] The size of the pixel block can be set according to actual conditions, and this embodiment does not impose any specific limitation on this.
[0067] For example, the size of the pixel block may be 16×16, but is not limited thereto.
[0068] The patching process of the image to be tested to obtain a plurality of pixel blocks may include patching the image to be tested through a preset convolution layer to obtain a plurality of pixel blocks, but is not limited thereto.
[0069] The above preset convolutional layer can be set according to actual conditions, and this embodiment does not make any specific limitation to this.
[0070] For example, the above-mentioned preset convolution layer can be a convolution layer with 3 input channels, 768 output channels, a convolution kernel size of 16×16, and a stride of 16, but is not limited thereto.
[0071] The above pixel vector sequence is a one-dimensional vector sequence, which may include vectors corresponding to each pixel block.
[0072] The length of the vector corresponding to the above pixel block can be set according to actual conditions, and this embodiment does not specifically limit this.
[0073] For example, the vector corresponding to the above pixel block may be a vector with a length of 768, but is not limited thereto.
[0074] The size of the above pixel vector sequence can be set according to actual conditions, and this embodiment does not impose any specific limitation on this.
[0075] For example, the above pixel vector sequence may be a vector of size N×M, where N represents the number of remaining pixel blocks and M represents the length of the vector corresponding to the pixel block, but is not limited thereto.
[0076] The above position encoding may be a preset learnable vector.
[0077] The size of the position code can be set according to actual conditions, and this embodiment does not impose any specific limitation on this.
[0078] For example, the position encoding may be a learnable vector of the same size as the pixel vector sequence, but is not limited thereto.
[0079] The above-mentioned addition processing of the pixel vector sequence and the preset position code to obtain the input sequence may include randomly deleting a preset percentage of data in the pixel vector sequence, and randomly deleting a preset percentage of data in the position code, and then adding the randomly deleted pixel vector sequence and the randomly deleted position code to obtain the input sequence.
[0080] Alternatively, the above-mentioned addition processing of the pixel vector sequence and the preset position code to obtain the input sequence may include adding the pixel vector sequence and the preset position code as an initial sequence, and randomly deleting a preset percentage of data in the initial sequence to obtain the input sequence, but is not limited to this.
[0081] The above preset percentage can be set according to actual conditions, and this embodiment does not make any specific limitation to this.
[0082] For example, the preset percentage may be 5%, but is not limited thereto.
[0083] In one example, when a test image of size 384×384 is input to the input layer, the test image is patched through a convolution layer with 3 input channels, 768 output channels, a convolution kernel size of 16×16, and a stride of 16 to obtain 24×24 pixel blocks of size 16×16, and then each pixel block is flattened into a one-dimensional vector of length 768, so that a pixel vector sequence of size 574×768 can be obtained. Finally, a learnable vector of size 574×768 is generated as a position code, and the position code of size 574×768 and the pixel vector sequence of size 574×768 are added to obtain an initial sequence, and 5% of the data in the initial sequence is randomly deleted to obtain an input sequence of size 548×768.
[0084] In some embodiments, the encoder structure may include multiple encoders connected in sequence, each encoder may include at least one Transformer block, the input of the first encoder is the input sequence, and the inputs of other encoders except the first encoder are the outputs of the previous encoder; each encoder is used to extract features from the encoder input through at least one Transformer block to obtain the encoder's encoding feature sequence.
[0085] In this embodiment, in the disguised object detection model, the encoder structure may include multiple encoders, and all encoders are connected in sequence, which means that the input of other encoders except the first encoder is the output of the previous encoder, that is, the first encoder. The input of the encoder is The output of the first encoder is the input sequence. Each encoder can include at least one Transformer block. encoder, through the At least one Transformer block of the encoder is The input of the encoder is used for feature extraction to obtain the It can be seen that by mining more abundant texture information and color information of camouflaged objects through multiple encoders connected in sequence, a coded feature sequence representing different feature dimensions is obtained, which can effectively improve the perception ability of the camouflaged object detection model for camouflaged objects, ensure the feature extraction ability of the camouflaged object detection model, and thus improve the detection accuracy of camouflaged objects.
[0086] The number of the above encoders can be set according to actual conditions, and this embodiment does not specifically limit this.
[0087] For example, the above encoder structure may include four encoders connected in sequence, but is not limited thereto.
[0088] The essence of the above encoding feature sequence is a vector (Token) sequence.
[0089] The number of Transformer blocks included in the above encoder can be set according to actual conditions, and this embodiment does not specifically limit this.
[0090] For example, each of the above encoders may include 3 sequentially connected Transformer blocks, that is, the above encoder structure may include 12 sequentially connected Transformer blocks, but is not limited thereto.
[0091] The number of self-attention heads of the multi-head self-attention mechanism of the above Transformer block can be set according to actual conditions, and this embodiment does not make any specific limitation on this.
[0092] For example, the multi-head self-attention mechanism of the above Transformer block can have 12 self-attention heads, but is not limited to this.
[0093] The structure of the multilayer perceptron of the above Transformer block can be set according to actual conditions, and this embodiment does not specifically limit this.
[0094] For example, the multilayer perceptron of the above Transformer block can be sequentially provided with two linear layers and one activation layer, wherein the two linear layers are used to perform a dimensionality increase operation on the input of the multilayer perceptron and a dimensionality reduction operation on the output of the dimensionality increase operation, and the activation layer is used to perform a delinearization process on the output of the last linear layer, thereby obtaining the output of the multilayer perceptron, but is not limited thereto.
[0095] The activation function of the activation layer can be set according to actual conditions, and this embodiment does not specifically limit this.
[0096] For example, the activation function of the activation layer may be a GELU activation function, but is not limited thereto.
[0097] In one example, referring to Figure 2 The above encoder structure is a typical Vision Transformer (ViT), which includes 12 sequentially connected Transformer blocks. Every 3 Transformer blocks constitute an encoder, that is, the encoder structure includes 4 sequentially connected encoders. encoders, , through 3 Transformer blocks to The input of the encoder is used for feature extraction to obtain the The encoded feature sequence of the encoder .
[0098] In each Transformer block, it is mainly composed of a multi-head self-attention layer, a multi-layer perceptron and two additive normalization layers; the multi-head self-attention layer has 12 self-attention layers, and the self-attention layer is mainly used to perform attention mechanism operations on the input of the self-attention layer, that is, multiplying the input of the self-attention layer with the key matrix, value matrix and query matrix respectively to obtain the key vector, value vector and query vector, and then performing attention mechanism operations on the key vector, value vector and query vector to obtain the output of the self-attention layer. In the multi-head self-attention layer, the outputs of all self-attention layers are processed to obtain the output of the multi-head self-attention layer as the input of the first additive normalization layer; in the first additive normalization layer, the multi-head The output of the self-attention layer is added to the input of the Transformer block, and then normalized to obtain the output of the first addition normalization layer as the input of the multilayer perceptron, thereby completing the residual connection; the multilayer perceptron is sequentially provided with two linear layers and a GELU activation layer, and the input of the multilayer perceptron is dimensionally increased and then dimensionally reduced through the two linear layers, and then the output of the dimension reduction operation is delinearized to obtain the output of the multilayer perceptron as the input of the second addition normalization layer; in the second addition normalization layer, its processing flow is the same as that of the first addition normalization layer, which will not be repeated here, and the output of the Transformer block is obtained after processing by the second addition normalization layer, thereby completing the feature extraction.
[0099] The normalization operation in the above-mentioned addition normalization layer can be set according to actual conditions, and this embodiment does not make any specific limitation to this.
[0100] For example, the layer normalization operation is adopted in the above-mentioned additive normalization layer, but it is not limited thereto.
[0101] In some embodiments, the above-mentioned encoding fusion module may include a first splicing layer and multiple parallel fusion branches, the multiple fusion branches correspond one-to-one to the multiple encoders in the encoder structure, and the input of the fusion branch is the encoding feature sequence of the encoder corresponding to the fusion branch; each fusion branch is used to de-patch, overlap and linearize the input of the fusion branch to obtain the overlapping feature sequence of the fusion branch; the first splicing layer is used to splice the overlapping feature sequences of multiple fusion branches and the encoding feature sequence of the last encoder to obtain a fusion feature sequence.
[0102] In this embodiment, in the disguised object detection model, the coding fusion module is mainly used to fuse the coding feature sequences of different feature dimensions. Specifically, the coding fusion module may include a first splicing layer and a plurality of parallel fusion branches, and the plurality of fusion branches correspond to the plurality of encoders one by one, that is, the first The encoder corresponds to the fusion branch, and the input of the fusion branch is the encoding feature sequence of the encoder corresponding to the fusion branch, that is, The input of the fusion branch is The encoded feature sequence output by the encoder. A fusion branch, which is mainly used for The input of the first fusion branch is de-patched, overlapped and linearized to obtain the The overlapping feature sequences of the fused branches are output. When all the fused branches have corresponding overlapping feature sequences, the overlapping feature sequences of the multiple fused branches and the encoded feature sequence of the last encoder are spliced together through the first splicing layer to obtain a fused feature sequence. It can be seen that by re-encoding the encoded feature sequences of different dimensions to obtain overlapping feature sequences of different dimensions, and by fusing the overlapping feature sequences of different dimensions to obtain the fused feature sequence, the camouflaged object detection model can focus on the local features of the camouflaged object. In this way, the local detail information and global detail information of the camouflaged object can be accurately captured, the loss of local detail information can be reduced, and the feature extraction capability of the camouflaged object detection model can be improved, thereby improving the detection accuracy of the camouflaged object.
[0103] The essence of the above fused feature sequence is a vector (Token) sequence.
[0104] The above-mentioned fusion branch may include a depatching layer, an overlapping layer and two linear layers, which are connected in sequence, wherein the depatching layer is used to perform depatching processing on the input of the fusion branch, that is, converting the coded feature sequence into a corresponding feature map; the overlapping layer is used to perform overlapping processing on the feature map obtained by the depatching processing, and converting the feature map corresponding to the coded feature sequence into a corresponding vector sequence, and the vector sequence is a vector sequence in an overlapping form, such as Figure 3 As shown in the figure, two linearization layers are used to linearize the vector sequence, and the last linearization layer outputs the overlapping feature sequence of the fusion branch.
[0105] In one example, referring to Figure 4 The encoder structure includes four encoders connected in sequence. The encoding fusion module includes the first concatenation layer and four parallel fusion branches. The input of the first fusion branch is the encoding feature sequence of the first encoder. , the input of the second fusion branch is the encoded feature sequence of the second encoder , the input of the third fusion branch is the encoded feature sequence of the third encoder , the input of the fourth fusion branch is the encoded feature sequence of the fourth encoder ; in the In the fusion branch, , first The encoded feature sequence of the encoder Convert to feature map, and then The feature map is converted to A sequence of overlapping vectors, and finally The overlapping vector sequence is linearized twice to obtain The overlapping feature sequences of the fusion branches are obtained; in the first concatenation layer, the overlapping feature sequences output by the four fusion branches and the encoded feature sequence of the fourth encoder are concatenated to obtain a fusion feature sequence .
[0106] In another example, after the fused feature sequence is obtained, a dimensionality reduction process is performed on the fused feature sequence to reduce the dimension of data input into the decoder structure.
[0107] In some implementations, the decoder structure includes at least one Transformer block, and the decoder structure is used to decode the fused feature sequence through the at least one Transformer block to obtain a decoded feature sequence.
[0108] In this embodiment, in the disguised object detection model, the decoder structure includes at least one Transformer block. The decoder structure is mainly used to decode the fused feature sequence through at least one Transformer block to obtain a decoded feature sequence, which is beneficial to improving the detection accuracy of the disguised object.
[0109] The essence of the above decoding feature sequence is a vector (Token) sequence.
[0110] The number of Transformer blocks included in the above decoder structure can be set according to actual conditions, and this embodiment does not specifically limit this.
[0111] For example, the above decoder structure may include 8 sequentially connected Transformer blocks, but is not limited thereto.
[0112] The number of self-attention heads of the multi-head self-attention mechanism of the above Transformer block can be set according to actual conditions, and this embodiment does not make any specific limitation on this.
[0113] For example, the multi-head self-attention mechanism of the above Transformer block can have 16 self-attention heads, but is not limited to this.
[0114] The structure of the multilayer perceptron of the above Transformer block can be set according to actual conditions, and this embodiment does not specifically limit this.
[0115] For example, the multilayer perceptron of the above Transformer block can be sequentially provided with two linear layers and one activation layer, wherein the two linear layers are used to perform a dimensionality increase operation on the input of the multilayer perceptron and a dimensionality reduction operation on the output of the dimensionality increase operation, and the activation layer is used to perform a delinearization process on the output of the last linear layer, thereby obtaining the output of the multilayer perceptron, but is not limited thereto.
[0116] The activation function of the activation layer can be set according to actual conditions, and this embodiment does not specifically limit this.
[0117] For example, the activation function of the activation layer may be a GELU activation function, but is not limited thereto.
[0118] In one example, referring to Figure 2 The above decoder structure is also a typical Vision Transformer (ViT) decoder, but different from the encoder structure, the decoder structure includes 8 sequentially connected Transformer blocks, and the multi-head self-attention mechanism of the Transformer block has 16 self-attention heads. The fusion feature sequence is decoded through these 8 sequentially connected Transformer blocks to obtain a decoded feature sequence. Among them, the implementation of the Transformer block in the decoder structure is the same as that of the Transformer block in the encoder structure, which will not be repeated.
[0119] In some embodiments, reference Figure 5 , is the encoded feature sequence of the fourth encoder, It is a decoding feature sequence. The above-mentioned refined perception module may include a first processing branch, a second processing branch, a first convolution splicing layer and a second convolution splicing layer; the first processing branch is used to perform discrete wavelet transform, convolution processing and upsampling on the decoding feature sequence to obtain a first feature map; the second processing branch is used to perform discrete wavelet transform, convolution processing and upsampling on the encoding feature sequence of the last encoder in the encoder structure to obtain a second feature map; the first convolution splicing layer is used to splice and convolve the first feature map and the second feature map to obtain a third feature map; the second convolution splicing layer is used to splice and convolve the first feature map, the second feature map and the third feature map to obtain a contour map of the camouflaged object.
[0120] In this embodiment, in the camouflaged object detection model, the refined perception module may include a first processing branch, a second processing branch, a first convolutional splicing layer and a second convolutional splicing layer. The refined perception module is mainly used to fuse the high-level detail information of the encoder structure and the low-level detail information of the decoder structure, so as to accurately refine the edge information of the camouflaged object.
[0121] Specifically, first, the high-level detail information of the encoder structure and the low-level detail information of the decoder structure are subjected to discrete wavelet transform, convolution processing and upsampling respectively through the first processing branch and the second processing branch, aiming to perceive the frequency characteristics associated with the camouflaged object in the detail information, which is conducive to further refining the foreground information and background information in the image to be tested, and at the same time, the edge features are upgraded to the corresponding feature map to complete the image restoration.
[0122] Then, the first feature map and the second feature map are fused through the first convolutional splicing layer and a convolution operation is performed to obtain the third feature map, which aims to transfer the deep high-level detail information of the encoder structure to the shallow low-level detail information of the decoder structure, and combine the deep high-level detail information with the shallow low-level detail information to further refine the foreground information and background information in the image to be tested, thereby realizing feature refinement and feature combination.
[0123] After that, all feature maps are fused and convolution is performed through the second convolutional splicing layer, aiming to pass the deep high-level detail information of the encoder structure and the shallow low-level detail information of the decoder structure to the refined feature map, i.e., the third feature map, and combine the refined feature map with the deep high-level detail information and the shallow low-level detail information, thereby realizing the fusion processing of multi-dimensional features. After the processing of the second convolutional splicing layer, the contour map of the camouflaged object is obtained, thereby realizing the detection of the camouflaged object.
[0124] In this way, this embodiment can perceive the frequency characteristics associated with the camouflaged object and refine the foreground information and background information in the image to be tested by refining the processing branches of the perception module, and restore the image. Through the convolutional splicing layers, the fusion processing of multi-dimensional features can be realized, thereby accurately segmenting the edges of the camouflaged object and improving the detection accuracy of the camouflaged object.
[0125] The above-mentioned processing branch may include a first discrete wavelet transform layer, a first convolution layer and a first upsampling layer connected in sequence, the first discrete wavelet transform layer is used to perform discrete wavelet transform on the input of the processing branch, the output of the first convolution layer is used to perform convolution processing on the output of the first discrete wavelet transform layer, and the first upsampling layer is used to upsample the output of the first convolution layer, thereby obtaining a feature map as the output of the processing branch.
[0126] The convolution kernel size of the first convolutional layer can be set according to actual conditions, and this embodiment does not specifically limit this.
[0127] For example, the first convolution layer may be a convolution layer with a convolution kernel of 1×1.
[0128] The above-mentioned first convolutional splicing layer may include a second splicing layer and a second convolutional layer connected sequentially, the second splicing layer is used to splice the first feature map and the second feature map into a fourth feature map, and the second convolutional layer is used to perform convolution processing on the fourth feature map to obtain a third feature map as the output of the first convolutional splicing layer.
[0129] The convolution kernel size of the second convolutional layer can be set according to actual conditions, and this embodiment does not specifically limit this.
[0130] For example, the second convolution layer may be a convolution layer with a convolution kernel of 1×1.
[0131] The above-mentioned third convolutional splicing layer may include a third splicing layer, a third convolutional layer and a fourth convolutional layer connected in sequence, the third splicing layer is used to splice the first feature map, the second feature map and the third feature map into a fifth feature map, the third convolutional layer is used to perform convolution processing on the fifth feature map to obtain the convolved fifth feature map, and the fourth convolutional layer is the output layer, which performs a convolution operation on the convolved fifth feature map to obtain a contour map of the camouflaged object as the output of the camouflaged object detection model.
[0132] The convolution kernel size of the third convolution layer and the convolution kernel size of the fourth convolution layer can be set according to actual conditions, and this embodiment does not specifically limit this.
[0133] For example, the third convolution layer and the fourth convolution layer may both be convolution layers with a convolution kernel of 1×1.
[0134] In some embodiments, reference Figure 6In the first processing branch, the discrete wavelet transform, convolution processing and upsampling of the decoded feature sequence to obtain the first feature map may include:
[0135] Performing low-frequency transformation on the decoding feature sequence to obtain low-frequency features of the decoding feature sequence;
[0136] Performing horizontal transformation on the decoding feature sequence to obtain horizontal features of the decoding feature sequence;
[0137] Performing vertical transformation on the decoding feature sequence to obtain vertical features of the decoding feature sequence;
[0138] Performing a diagonal transformation on the decoding feature sequence to obtain a diagonal feature of the decoding feature sequence;
[0139] The low-frequency features, horizontal features, vertical features and diagonal features of the decoding feature sequence are concatenated to obtain the frequency features of the decoding feature sequence;
[0140] Performing convolution processing on the frequency features of the decoding feature sequence to obtain the convolution features of the decoding feature sequence;
[0141] The convolutional features of the decoded feature sequence are upsampled to obtain the first feature map.
[0142] In this embodiment, first, the decoding feature sequence is subjected to parallel low-frequency transformation, horizontal transformation, vertical transformation and diagonal transformation, so as to obtain the low-frequency features, horizontal features, vertical features and diagonal features of the decoding feature sequence; then, the low-frequency features, horizontal features, vertical features and diagonal features of the decoding feature sequence are spliced into the frequency features of the decoding feature sequence; thereafter, the frequency features of the decoding feature sequence are subjected to convolution processing to obtain the convolution features of the decoding feature sequence, and the convolution operation processing is implemented to ensure the feature accuracy; finally, the convolution features of the decoding feature sequence are upsampled to restore the convolution features to the image form, so as to obtain the first feature map. In this way, the frequency features associated with the camouflaged object in the shallow low-level detail information are fully perceived, which is conducive to further refining the foreground information and background information in the image to be tested, and at the same time, the edge features are upgraded to the corresponding feature map to complete the image restoration.
[0143] In some embodiments, reference Figure 6 In the second processing branch, the discrete wavelet transform, convolution processing and upsampling of the encoding feature sequence of the last encoder in the encoder structure to obtain the first feature map may include:
[0144] Performing a low-frequency transformation on the coding feature sequence of the last encoder to obtain a low-frequency feature of the coding feature sequence of the last encoder;
[0145] Performing a horizontal transformation on the coding feature sequence of the last encoder to obtain the horizontal features of the coding feature sequence of the last encoder;
[0146] Performing vertical transformation on the coding feature sequence of the last encoder to obtain vertical features of the coding feature sequence of the last encoder;
[0147] Performing a diagonal transformation on the coding feature sequence of the last encoder to obtain a diagonal feature of the coding feature sequence of the last encoder;
[0148] The low-frequency features, horizontal features, vertical features, and diagonal features of the coding feature sequence of the last encoder are concatenated to obtain the frequency features of the coding feature sequence of the last encoder;
[0149] Performing convolution processing on the frequency features of the coding feature sequence of the last encoder to obtain the convolution features of the coding feature sequence of the last encoder;
[0150] The convolutional features of the encoded feature sequence of the last encoder are upsampled to obtain the second feature map.
[0151] In this embodiment, first, the coding feature sequence of the last encoder is subjected to parallel low-frequency transformation, horizontal transformation, vertical transformation and diagonal transformation, so as to obtain the low-frequency features, horizontal features, vertical features and diagonal features of the coding feature sequence of the last encoder; then, the low-frequency features, horizontal features, vertical features and diagonal features of the coding feature sequence of the last encoder are spliced into the frequency features of the coding feature sequence of the last encoder; thereafter, the frequency features of the coding feature sequence of the last encoder are subjected to convolution processing to obtain the convolution features of the coding feature sequence of the last encoder, and the convolution operation processing is implemented to ensure the feature accuracy; finally, the convolution features of the coding feature sequence of the last encoder are upsampled to restore the convolution features to the image form, so as to obtain the second feature map. In this way, the frequency features associated with the camouflaged object in the deep high-level detail information are fully perceived, which is conducive to further refining the foreground information and background information in the image to be tested, and at the same time, the edge features are upgraded to the corresponding feature map to complete the image restoration.
[0152] In some embodiments, reference Figure 7 The step of acquiring the image sample set for training the above-mentioned disguised object detection model may include:
[0153] According to the preset mask, a background pixel set is obtained, where the background pixel set includes at least one zero-value pixel point, where the zero-value pixel point is a pixel point whose value in the preset mask is zero;
[0154] Randomly select a number of zero-value pixels from the background pixel set as target pixels;
[0155] Calculate the mean color value of all target pixels as the average background color value;
[0156] Acquire a number of salient sample images and binary images corresponding to each salient sample image, wherein the binary image is a binary image of a salient object contained in the salient sample image, and the total number of salient sample images is equal to the total number of target pixels;
[0157] For each salient sample image, the color value of the pixel of the salient object in the salient sample image is replaced with the average background color value to obtain a camouflaged sample image corresponding to the salient sample image;
[0158] An image sample set is obtained through the disguised sample images and binary images corresponding to each salient sample image.
[0159] In this embodiment, salient object detection (SOD) is a visual task that is completely different from disguised object detection. The purpose of disguised object detection is to identify objects that blend into the surrounding environment, while salient object detection focuses on identifying and highlighting the most prominent or visually impactful objects in the image. Related technologies have been exploring the joint learning of these two opposing tasks to exploit their complementarity. However, related technologies simply explore the properties between the two tasks without in-depth research on their interaction and mutual transformation. At present, there is still a lack of a method that can achieve mutual transformation between salient objects and disguised objects.
[0160] In this regard, the present embodiment provides a disguise transformation strategy, which can transform an image containing a salient object into an image containing a disguised object, that is, realize mutual transformation between a salient object and a disguised object.
[0161] In the camouflage transformation strategy, first, according to the preset mask, a background pixel set is obtained, and the background pixel set may include at least one zero-value pixel point, and the zero-value pixel point refers to a pixel point with a value of 0 in the preset mask. Specifically, if the value of a certain pixel point on the preset mask is 0, the pixel point is classified into the background pixel set, thereby obtaining a background pixel set. Then, N zero-value pixels are randomly selected from the background pixel set as target pixels, and the average of the color values of all target pixels is calculated as the average background color value, and the average background color value is used to modify the region of interest in the salient sample image, that is, the salient object. After that, N salient sample images and binary images corresponding to each salient sample image are obtained, the binary image is the label image of the salient sample image, and the binary image refers to the binary image of the salient object contained in the salient sample image, and the color value of the pixel point of the salient object in each salient sample image is replaced with the average background color value, thereby obtaining a camouflage sample image corresponding to each salient sample image; finally, the image sample set is obtained through the camouflage sample image and the binary image corresponding to each salient sample image. In this way, through the mutual conversion between salient objects and camouflaged objects, the training data for camouflaged object detection can be effectively expanded, which can reduce the data preparation cost and time cost while improving the image quality of the training data, thereby ensuring the accuracy of the camouflaged object detection model.
[0162] In the above binary image, the black part is the background environment and the white part is the camouflage object.
[0163] In some implementations, the step of acquiring the image sample set for training the camouflaged object detection model may further include at least one of the following:
[0164] Scaling the disguised sample images and binary images in the image sample set;
[0165] Data augmentation is performed on disguised sample images and binary images in the image sample set.
[0166] In this embodiment, after obtaining the image sample set, the camouflaged sample images and binary images in the image sample set are scaled so that the size of the camouflaged sample images and the size of the binary images are the same; and / or, the camouflaged sample images and binary images in the image sample set are data amplified to increase the amount of training data, thereby improving the image quality and quantity of the training data, thereby ensuring the accuracy of the camouflaged object detection model.
[0167] The implementation of the above-mentioned scaling process, the size of the above-mentioned disguised sample image and the size of the above-mentioned binary image can be set according to actual conditions, and this embodiment does not make any specific limitation on this.
[0168] For example, the above-mentioned scaling process may be to scale the image using bilinear interpolation, but is not limited thereto.
[0169] For another example, the size of the disguised sample image and the size of the binary image may both be 384×384, but are not limited thereto.
[0170] The implementation of the above data expansion can be set according to actual conditions, and this embodiment does not make any specific limitation to this.
[0171] For example, the data augmentation may be to randomly horizontally flip all images with a probability of 50%, but is not limited thereto.
[0172] In some embodiments, the loss function of the camouflaged object detection model may include at least one of a binary cross entropy loss function or an intersection-over-union loss function.
[0173] In this embodiment, the loss function defining the disguised object detection model may include at least one of a binary cross entropy loss function (BCE-Loss) or an intersection-over-union loss function (IOU-Loss). In this way, the difference between the output of the disguised object detection model and the ground truth is measured by the binary cross entropy loss function and the intersection-over-union loss function, which is conducive to optimizing the performance of the disguised object detection model based on the difference and improving the accuracy of the disguised object detection model.
[0174] In one example, during the training process of the disguised object detection model, the binary cross entropy loss function of the output of the disguised object detection model and the true value is calculated, and the intersection-over-union loss function of the output of the disguised object detection model and the true value is calculated at the same time, and then the binary cross entropy loss function and the intersection-over-union loss function are weighted to obtain the loss function value of the disguised object detection model.
[0175] In order to facilitate the understanding of the above-mentioned disguised object detection method of the present application, the actual application scenario of the above-mentioned disguised object detection method of the present application is used as an example. Figures 2 to 7 The overall process of implementing camouflaged object detection in this application is as follows: steps S201-S203.
[0176] S201, constructing an image sample set.
[0177] Specifically, if the value of a certain pixel on the preset mask is 0, the pixel is classified into the background pixel set, thereby obtaining the background pixel set, which may include at least one zero-value pixel, and the zero-value pixel refers to a pixel in the preset mask with a value of 0. Then, N zero-value pixels are randomly selected from the background pixel set as target pixels, and the average of the color values of all target pixels is calculated as the average background color value. Afterwards, N salient sample images and the binary images corresponding to each salient sample image are obtained. The salient sample image is a color image of three primary color channels, and the binary image is a label image of the salient sample image. The binary image refers to the binary image of the salient object contained in the salient sample image. The object in the binary image is represented by white and the background is represented by black. Then, the color value of the pixel point of the salient object in each salient sample image is replaced with the average background color value, so as to obtain the camouflaged sample image corresponding to each salient sample image; finally, bilinear interpolation is performed on the camouflaged sample image and the binary image corresponding to each salient sample image, and then the camouflaged sample image and the binary image corresponding to each salient sample image are randomly horizontally flipped with a probability of 50%, so as to construct an image sample set.
[0178] Among them, in the image sample set, the size of all disguised sample images and all binary images is 384×384.
[0179] S202: Train the disguised object detection model using the image sample set to obtain a trained disguised object detection model.
[0180] Among them, during the training process, the binary cross entropy loss function of the output of the camouflaged object detection model and the true value is calculated, and the intersection-over-union loss function of the output of the camouflaged object detection model and the true value is calculated at the same time. The binary cross entropy loss function and the intersection-over-union loss function are weighted to obtain the loss function value of the camouflaged object detection model. The loss function value of the camouflaged object detection model is used to tune the camouflaged object detection model.
[0181] S203 , obtaining a test image with a size of 384×384 and containing the test disguised object as an input of the trained disguised object detection model, and outputting a contour map of the test disguised object in the test image through the trained disguised object detection model.
[0182] Specifically, the disguised object detection model includes the following parts:
[0183] (1) Input layer. When a 384×384 image is input to the input layer, it is patched by a convolutional layer with 3 input channels, 768 output channels, 16×16 kernel size, and 16 stride. Each pixel block is flattened into a one-dimensional vector of length 768. This results in a 574×768 pixel vector sequence. Finally, a 574×768 learnable vector is generated as the position code. The 574×768 position code and the 574×768 pixel vector sequence are added to obtain the initial sequence. 5% of the data in the initial sequence is randomly deleted to obtain an input sequence of 548×768 as the input of the encoder structure.
[0184] (2) The encoder structure includes 12 sequentially connected Transformer blocks, where every 3 Transformer blocks form an encoder, i.e., the encoder structure includes 4 sequentially connected encoders. The multi-head self-attention mechanism of each Transformer block has 12 self-attention heads. encoders, , through 3 Transformer blocks to The input of the encoder is used for feature extraction to obtain the The encoded feature sequences of the encoders are used as the input of the encoding fusion module.
[0185] (3) A coding fusion module, which includes a first concatenation layer and four parallel fusion branches, wherein the input of the first fusion branch is the coding feature sequence of the first encoder, the input of the second fusion branch is the coding feature sequence of the second encoder, the input of the third fusion branch is the coding feature sequence of the third encoder, and the input of the fourth fusion branch is the coding feature sequence of the fourth encoder; In the fusion branch, , first The encoded feature sequence of the encoder is converted into the feature map, and then The feature map is converted to A sequence of overlapping vectors, and finally The overlapping vector sequence is linearized twice to obtain In the first concatenation layer, the overlapping feature sequences output by the four fusion branches and the encoded feature sequence of the fourth encoder are concatenated to obtain a fused feature sequence, and the fused feature sequence is subjected to dimensionality reduction so that the dimension of the fused feature sequence is reduced from 768 dimensions to 512 dimensions.
[0186] (4) The decoder structure includes 8 sequentially connected Transformer blocks, where the multi-head self-attention mechanism of the Transformer block has 16 self-attention heads. The fused feature sequence is decoded by these 8 sequentially connected Transformer blocks to obtain a decoded feature sequence.
[0187] (5) A refinement perception module, comprising a first processing branch, a second processing branch, a first convolutional splicing layer, and a second convolutional splicing layer, wherein:
[0188] The first processing branch is used to perform discrete wavelet transform, convolution processing and upsampling on the decoded feature sequence in sequence to obtain a first feature map; the second processing branch is used to perform discrete wavelet transform, convolution processing and upsampling on the encoded feature sequence of the fourth encoder in sequence to obtain a second feature map; the first convolution splicing layer includes a second splicing layer connected in sequence and a second convolution layer with a convolution kernel of 1×1, the second splicing layer is used to splice the first feature map and the second feature map into a fourth feature map, and the second convolution layer is used to perform convolution processing on the fourth feature map to obtain a third feature map. The third convolutional splicing layer may include a third splicing layer, a third convolutional layer with a convolution kernel of 1×1, and a fourth convolutional layer with a convolution kernel of 1×1 connected in sequence. The third splicing layer is used to splice the first feature map, the second feature map, and the third feature map into a fifth feature map. The third convolutional layer is used to perform convolution on the fifth feature map to obtain the convolved fifth feature map. The fourth convolutional layer is the output layer, which performs a convolution operation on the convolutional fifth feature map to obtain a contour map of the camouflaged object as the output of the camouflaged object detection model.
[0189] Among them, for the discrete wavelet transform, the implementation process is: perform parallel low-frequency transform, horizontal transform, vertical transform and diagonal transform on the input of the processing branch to obtain the corresponding low-frequency features, horizontal features, vertical features and diagonal features, and then splice the low-frequency features, horizontal features, vertical features and diagonal features into frequency features, thereby realizing the processing of discrete wavelet transform.
[0190] In addition, refer to Figure 8 , the embodiment of the present application further provides a disguised object detection device, which may include:
[0191] An acquisition module 301 is used to acquire an image to be tested containing a camouflaged object;
[0192] The first processing module 302 is equipped with a disguised object detection model, and is used to perform disguised object detection on the image to be tested through the disguised object detection model to obtain a contour map of the disguised object;
[0193] The disguised object detection model is obtained by training a preset disguised object image sample set, and the disguised object detection model includes:
[0194] The input layer is used to preprocess the image to be tested to obtain an input sequence;
[0195] An encoder structure, the encoder structure includes a plurality of encoders connected in sequence, each encoder includes at least one Transformer block, and the encoder structure is used to extract features from an input sequence through the plurality of Transformer blocks to obtain a plurality of encoding feature sequences;
[0196] A coding fusion module is used to fuse multiple coding feature sequences to obtain a fused feature sequence;
[0197] A decoder structure, used for decoding the fused feature sequence through at least one Transformer block to obtain a decoded feature sequence;
[0198] The refinement perception module is used to perform refinement perception processing on the decoded feature sequence and the encoded feature sequence of the last encoder in the encoder structure to obtain a contour map of the camouflaged object.
[0199] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.
[0200] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0201] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several programs to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., various media that can store program codes.
[0202] The logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable programs for implementing logical functions, and may be embodied in any computer-readable medium for use by a program execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch and execute a program from a program execution system, device or apparatus), or in conjunction with such program execution system, device or apparatus. For purposes of this specification, a "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by a program execution system, device or apparatus, or in conjunction with such program execution system, device or apparatus.
[0203] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0204] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0205] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0206] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
[0207] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.
Claims
1. A method for detecting a disguised object, characterized in that: include: Acquire a test image containing a camouflaged object; Performing disguised object detection on the image to be tested by using a disguised object detection model to obtain a contour image of the disguised object; The disguised object detection model is obtained by training a preset image sample set, and the disguised object detection model includes: An input layer, used for preprocessing the image to be tested to obtain an input sequence; An encoder structure, the encoder structure comprising a plurality of encoders connected in sequence, each of the encoders comprising at least one Transformer block, the encoder structure being used to extract features from the input sequence through the plurality of Transformer blocks to obtain a plurality of encoding feature sequences; A coding fusion module, used for fusing a plurality of coding feature sequences to obtain a fused feature sequence; A decoder structure, used for decoding the fused feature sequence through at least one Transformer block to obtain a decoded feature sequence; A refinement perception module, used for performing refinement perception processing on the decoded feature sequence and the encoding feature sequence of the last encoder in the encoder structure to obtain a contour image of the camouflaged object; Wherein, the refined perception module includes: A first processing branch is used to perform discrete wavelet transform, convolution processing and up-sampling on the decoded feature sequence to obtain a first feature map; A second processing branch is used to perform discrete wavelet transform, convolution processing and up-sampling on the encoding feature sequence of the last encoder in the encoder structure to obtain a second feature map; A first convolutional concatenation layer, used for concatenating and convolving the first feature map and the second feature map to obtain a third feature map; The second convolutional splicing layer is used to perform splicing and convolution processing on the first feature map, the second feature map and the third feature map to obtain a contour map of the camouflaged object.
2. The method for detecting a disguised object according to claim 1, characterized in that: The preprocessing of the image to be tested to obtain an input sequence includes: Performing patch processing on the image to be tested to obtain a plurality of pixel blocks; Vectorizing the plurality of pixel blocks to obtain a pixel vector sequence; The pixel vector sequence and the preset position code are added to obtain the input sequence.
3. The method for detecting a disguised object according to claim 1, wherein: The encoding fusion module includes a first splicing layer and a plurality of fusion branches connected in parallel, the plurality of fusion branches correspond one-to-one to a plurality of encoders in the encoder structure, and the input of the fusion branch is a coding feature sequence of the encoder corresponding to the fusion branch; Each of the fusion branches is used to perform a de-patch process, an overlap process and a linear process on the input of the fusion branch to obtain an overlapping feature sequence of the fusion branch; The first concatenation layer is used to concatenate the overlapping feature sequences of the multiple fusion branches and the encoding feature sequence of the last encoder in the encoder structure to obtain the fusion feature sequence.
4. The method for detecting a disguised object according to claim 1, wherein: The step of performing discrete wavelet transform, convolution processing and upsampling on the decoded feature sequence to obtain a first feature map comprises: Performing a low-frequency transformation on the decoding feature sequence to obtain a low-frequency feature of the decoding feature sequence; Performing a horizontal transformation on the decoding feature sequence to obtain horizontal features of the decoding feature sequence; Performing vertical transformation on the decoding feature sequence to obtain vertical features of the decoding feature sequence; Performing a diagonal transformation on the decoding feature sequence to obtain a diagonal feature of the decoding feature sequence; The low-frequency features, horizontal features, vertical features and diagonal features of the decoding feature sequence are concatenated to obtain the frequency features of the decoding feature sequence; Performing convolution processing on the frequency characteristics of the decoding feature sequence to obtain the convolution characteristics of the decoding feature sequence; The convolutional features of the decoded feature sequence are upsampled to obtain the first feature map.
5. The method for detecting a disguised object according to claim 1, wherein: The step of performing discrete wavelet transform, convolution processing and upsampling on the encoding feature sequence of the last encoder in the encoder structure to obtain a second feature map comprises: Performing a low-frequency transformation on the coding feature sequence of the last encoder to obtain a low-frequency feature of the coding feature sequence of the last encoder; Performing a horizontal transformation on the coding feature sequence of the last encoder to obtain a horizontal feature of the coding feature sequence of the last encoder; Performing a vertical transformation on the coding feature sequence of the last encoder to obtain a vertical feature of the coding feature sequence of the last encoder; Performing a diagonal transformation on the coding feature sequence of the last encoder to obtain a diagonal feature of the coding feature sequence of the last encoder; The low-frequency features, horizontal features, vertical features and diagonal features of the coding feature sequence of the last encoder are concatenated to obtain the frequency features of the coding feature sequence of the last encoder; Performing convolution processing on the frequency characteristics of the coding feature sequence of the last encoder to obtain the convolution characteristics of the coding feature sequence of the last encoder; The convolutional features of the last encoded feature sequence of the encoder are upsampled to obtain the second feature map.
6. The method for detecting a disguised object according to claim 1, wherein: The step of acquiring the image sample set comprises: According to the preset mask, a background pixel set is obtained, wherein the background pixel set includes at least one zero-value pixel point, and the zero-value pixel point is a pixel point whose value in the preset mask is zero; Randomly select a number of zero-valued pixels from the background pixel set as target pixels; Calculate the average color value of all target pixels as the average background color value; Acquire a plurality of salient sample images and binary images corresponding to the salient sample images, wherein the binary images are binary images of salient objects contained in the salient sample images, and the total number of the salient sample images is equal to the total number of the target pixels; For each of the salient sample images, replacing the color value of the pixel of the salient object in the salient sample image with the average background color value to obtain a camouflaged sample image corresponding to the salient sample image; The image sample set is obtained through the disguised sample images and the binary images corresponding to the salient sample images.
7. The method for detecting a disguised object according to claim 1, characterized in that: The loss function of the disguised object detection model includes at least one of a binary cross entropy loss function or an intersection-over-union loss function.
8. A disguised object detection device, characterized in that: include: An acquisition module, used for acquiring an image to be tested containing a camouflaged object; a first processing module, equipped with a disguised object detection model, configured to perform disguised object detection on the image to be tested by using the disguised object detection model to obtain a contour image of the disguised object; The disguised object detection model is obtained by training a preset disguised object image sample set, and the disguised object detection model includes: An input layer, used for preprocessing the image to be tested to obtain an input sequence; An encoder structure, the encoder structure comprising a plurality of encoders connected in sequence, each of the encoders comprising at least one Transformer block, the encoder structure being used to extract features from the input sequence through the plurality of Transformer blocks to obtain a plurality of encoding feature sequences; A coding fusion module, used for fusing a plurality of coding feature sequences to obtain a fused feature sequence; A decoder structure, used for decoding the fused feature sequence through at least one Transformer block to obtain a decoded feature sequence; A refinement perception module, used for performing refinement perception processing on the decoded feature sequence and the encoding feature sequence of the last encoder in the encoder structure to obtain a contour image of the camouflaged object; Wherein, the refined perception module includes: A first processing branch is used to perform discrete wavelet transform, convolution processing and up-sampling on the decoded feature sequence to obtain a first feature map; A second processing branch is used to perform discrete wavelet transform, convolution processing and up-sampling on the encoding feature sequence of the last encoder in the encoder structure to obtain a second feature map; A first convolutional concatenation layer, used for concatenating and convolving the first feature map and the second feature map to obtain a third feature map; The second convolutional splicing layer is used to perform splicing and convolution processing on the first feature map, the second feature map and the third feature map to obtain a contour map of the camouflaged object.
Citation Information
Patent Citations
Camouflage object segmentation method based on feature fusion Transform
CN118038061A
Camouflage target detection method and system based on pyramid type visual Transform
CN118172540A