Diffusion model-based camouflage target generation method and related device
Through the method of camouflage target generation based on the diffusion model, the attention fusion mechanism and loss guidance technology are used to solve the problem of camouflage target generation in the existing technology without additional background information, and high-quality camouflage target generation is achieved.
Patent Information
- Application Number
- CN202510101298.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art is difficult to generate high-quality camouflage targets without additional background information, and the existing camouflage target data set is not rich enough, making it difficult to meet the camouflage target detection needs in complex backgrounds.
The camouflage target generation method based on the diffusion model is adopted, and the pre-trained diffusion model is used to infer, and the attention fusion mechanism and content-based loss guidance are used to generate pictures containing the camouflage target.
Generating high-quality camouflage targets without providing additional background knowledge overcomes the problem of poor generation of prior art when background information is insufficient, and improves the realism and consistency of generated images.
Smart Images

Figure CN119919522A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image generation in computer vision, and in particular relates to a camouflage target generation method based on a diffusion model and a related device. Background Art
[0002] Camouflaged targets refer to targets whose texture and features are very similar to their surroundings, so that the human eye cannot distinguish them. Since camouflaged targets are difficult to identify and segment, existing conventional target detection methods are ineffective. Research on camouflaged targets has always been a relatively complex problem in the field of computer vision.
[0003] Camouflaged target detection is a research direction in the field of computer vision, which aims to detect and identify camouflaged targets in complex backgrounds. One of the major difficulties faced in the field of camouflaged target detection is that it is very difficult to construct a dataset. Explanation: there are not many images of camouflaged targets, and it is very difficult to collect a large number of images of camouflaged targets. In addition, most of the existing camouflaged target datasets are based on natural images and camouflaged targets mainly of wild animals. It is difficult to obtain relevant camouflaged targets to be detected, resulting in the low richness of existing camouflaged target datasets. In summary, it is particularly important to use methods such as image generation and style transfer to achieve the generation of camouflaged targets.
[0004] Camouflage target generation refers to the use of technical means to create a target that is highly similar to the background or can blend into the background to achieve the purpose of hiding and confusing. Camouflage target generation is different from general target generation. Camouflage target generation requires that the features of the target and the background are highly consistent and can provide diversity and richness. In addition, since camouflage target generation serves camouflage target detection, the target and the environment cannot be completely integrated. Instead, the target needs to be as similar as possible to the environmental information while retaining its own characteristics, so that the target detection network can be trained more effectively.
[0005] At present, there are two main branches of existing camouflaged target generation technology solutions. One branch is to use the image fusion method to fuse the foreground and background images together after the foreground and complete background images are given. This method is mainly based on the style loss and content loss proposed in the field of style transfer. Since the background information of the target position has been specified, the fusion method between the target and the background can be used to achieve the fusion of the target and the background at the corresponding position. However, in actual camouflaged target generation applications, it is often impossible to obtain the background information corresponding to the target area of the corresponding camouflaged target, which makes this type of method unapplicable. The other branch is to use the Inpainting model to match the target with the appropriate Because the characteristics of the camouflaged target and the background knowledge are very close, the background information is also contained in the target, and the target information is also contained in the background. In camouflaged target detection, it is not desired that the camouflaged target generate some unreal artifacts or undesirable areas due to changes in the target, and the process of generating the background is always more difficult than the process of generating the target. Therefore, the "background knowledge retrieval" method is used to match the target with appropriate background information. However, this method also has defects, that is, the richness of the generated image depends on the diversity of the target, that is, when constructing the data set, it is necessary to provide targets that have a certain degree of camouflage and are quite diverse. Such requirements are often difficult to achieve in practical applications. Summary of the invention
[0006] The purpose of the present invention is to provide a camouflage target generation method and related device based on a diffusion model to solve one or more of the above-mentioned technical problems. The technical solution disclosed in the present invention can convert the target in an image into a camouflage target only through one image without giving additional prior knowledge (such as background knowledge, etc.), thus overcoming the problem that the existing generation method is limited in practical application.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a method for generating a camouflaged target based on a diffusion model, comprising the following steps:
[0009] Acquire a picture containing a target to be detected, and segment the picture containing the target to be detected to obtain a detection target image and a background image;
[0010] Based on the picture containing the target to be detected, the detection target image and the background image, a pre-trained diffusion model is used for reasoning to obtain a picture containing the camouflaged target;
[0011] Wherein, the diffusion model includes:
[0012] An encoder module, used for inputting the picture containing the target to be detected, the detection target image and the background image, encoding them, and outputting an initial state latent variable;
[0013] An image inversion module, used for inputting the initial state latent variables and performing image inversion processing, and outputting the latent variables after image inversion;
[0014] A Unet module is used to input the hidden variables after the image is inverted and diffuse multiple times based on a scheduler, and output the diffused hidden variables; wherein the scheduler is used to predict the hidden variables of the next stage according to the time step and noise; the Unet module includes an encoding layer and a decoding layer, the encoding layer and the decoding layer include the same number of Unet processing units, each of which includes a residual connected ResNet module, an attention module and a cross attention module; the attention module in each of the Unet processing units in the encoding layer adopts a self-attention mechanism; the attention module in each of the Unet processing units in the decoding layer adopts a self-attention mechanism before a predetermined number of time steps, and then adopts one or both of the attention direct fusion calculation method and the attention indirect injection method to inject the feature information of the hidden variables of the background image into the target to be detected in the picture containing the target to be detected;
[0015] The decoder module is used to input the diffused latent variable and decode it, and output a picture containing the disguised target.
[0016] A further improvement of the camouflage target generation method of the present invention is that:
[0017] The step of injecting the characteristic information of the hidden variable of the background image into the target to be detected in the picture containing the target to be detected by using one or both of the direct attention fusion calculation method and the indirect attention injection method, the step of injecting the characteristic information of the hidden variable of the background image into the target to be detected in the picture containing the target to be detected by using the direct attention fusion calculation method specifically includes:
[0018] Introduce an attention fusion term in the self-attention calculation of the image containing the target to be detected, use the query Query of the image containing the target to be detected, the key of the target image and the value of the background image to perform attention calculation, and add the attention calculation result to the reconstruction process of the original image containing the target to be detected;
[0019] A mapping relationship between a picture containing a target to be detected and a target area in the detection target image is established, where the mapping relationship is an attention weight matrix established by a query Query of the picture containing the target to be detected and a key Key of the detection target image; using the mapping relationship, the value Value of the background image is directly injected into the picture containing the target to be detected in the form of self-attention calculation.
[0020] A further improvement of the camouflage target generation method of the present invention is that:
[0021] In the step of injecting the feature information of the hidden variable of the background image into the target to be detected in the picture containing the target to be detected by using one or both of the direct attention fusion calculation method and the indirect attention injection method,
[0022] The direct attention fusion calculation method is to directly use the query containing the image of the target to be detected, the key of the target image and the value of the background image to perform self-attention calculation;
[0023] The indirect attention injection method is to inject the key and value of the background image into the detection target image;
[0024] The complete reasoning process for the final three images is as follows:
[0025]
[0026] In the formula, q comp , k comp 、v comp Respectively represent the query, key, and value of the image containing the target to be detected; q for , k for 、v for Respectively represent the query, key and value of the target image to be detected; q bg , k bg 、v bg denote the query, key, and value of the background image, respectively; Attention denotes the scaled dot product attention mechanism; Softmax refers to the normalized exponential function; d k represents scaling, with a dimension of size K; w1 and w2 represent the weights of indirect injection and direct injection, respectively.
[0027] A further improvement of the camouflage target generation method of the present invention is that:
[0028] In the step of performing multiple diffusions based on the scheduler, in each time step after a predetermined number of time steps, after the latent variables of the picture containing the target to be detected are updated, content-based loss guidance is performed;
[0029] The steps of content-based loss guidance are: set the mode to allow gradient propagation, re-set the hidden variable z of the image containing the target to be detected t Input into the Unet module and get the result z t-1 , using the current z t-1 Predict the hidden variable z0 at time t=0; input the hidden variable z0 into the decoder module to obtain the image reconstructed at the current time step Using the reconstructed image and the image I containing the target to be detected comp Calculate the content loss L of the background area mse , using the calculated loss to optimize z t-1 ;
[0030] The loss calculation expression is:
[0031]
[0032] In the formula, x t represents the latent variable of the diffusion model at step t; σ is a predefined hyperparameter when training the diffusion model; ε θ is the parameterized network for predicting noise; t represents the number of steps of the current diffusion model inference; z t represents I at step t comp The corresponding hidden variable; η is the learning rate of gradient descent; downsample(·) represents the downsampling operation; z refers to the hidden variable containing the target image to be detected; L mse Refers to pixel-level difference loss.
[0033] A further improvement of the camouflage target generation method of the present invention is that:
[0034] The Unet module is provided with 16 layers of Unet processing units in total. As the depth of the layer in Unet increases, the resolution of each layer of Unet processing units first decreases and then increases, and the feature dimension of each pixel first increases and then decreases; wherein, the decoding layer of the Unet module is the part where the resolution increases and the number of feature dimensions of the pixels decreases.
[0035] A further improvement of the camouflage target generation method of the present invention is that:
[0036] The attention module in each of the Unet processing units of the decoding layer adopts a self-attention mechanism before a predetermined time step, and then adopts one or both of an attention direct fusion calculation method and an attention indirect injection method to inject feature information of latent variables of the background image into the target to be detected in the picture containing the target to be detected,
[0037] In terms of time, in the steps after time step t=30, one or both of the direct attention fusion calculation method and the indirect attention injection method are used to inject the feature information of the latent variable of the background image into the target to be detected in the picture containing the target to be detected;
[0038] Spatially, in the Unet processing units after the 12th layer of the Unet module, one or both of the attention direct fusion calculation method and the attention indirect injection method are used to inject the feature information of the latent variables of the background image into the target to be detected in the picture containing the target to be detected.
[0039] A further improvement of the camouflage target generation method of the present invention is that in the image inversion module, when performing image inversion, the hidden variables corresponding to the background image, the detection target image and the image containing the target to be detected are inverted to the hidden variables corresponding to the time t=30, and then in the subsequent reasoning process, the hidden variables at the time t=30 are used as the initial state to perform the reverse denoising process.
[0040] In a second aspect, the present invention provides a camouflage target generation system based on a diffusion model, comprising:
[0041] A data acquisition module is used to acquire a picture containing a target to be detected, and to segment the picture containing the target to be detected to obtain a detection target image and a background image;
[0042] A disguised target generation module, used to perform reasoning using a pre-trained diffusion model based on the picture containing the target to be detected, the detection target image and the background image, to obtain a picture containing the disguised target;
[0043] Wherein, the diffusion model includes:
[0044] An encoder module, used for inputting the picture containing the target to be detected, the detection target image and the background image, encoding them, and outputting an initial state latent variable;
[0045] An image inversion module, used for inputting the initial state latent variables and performing image inversion processing, and outputting the latent variables after image inversion;
[0046] A Unet module is used to input the hidden variables after the image is inverted and diffuse multiple times based on a scheduler, and output the diffused hidden variables; wherein the scheduler is used to predict the hidden variables of the next stage according to the time step and noise; the Unet module includes an encoding layer and a decoding layer, the encoding layer and the decoding layer include the same number of Unet processing units, each of which includes a residual connected ResNet module, an attention module and a cross attention module; the attention module in each of the Unet processing units in the encoding layer adopts a self-attention mechanism; the attention module in each of the Unet processing units in the decoding layer adopts a self-attention mechanism before a predetermined number of time steps, and then adopts one or both of the attention direct fusion calculation method and the attention indirect injection method to inject the feature information of the hidden variables of the background image into the target to be detected in the picture containing the target to be detected;
[0047] The decoder module is used to input the diffused latent variable and decode it, and output a picture containing the disguised target.
[0048] In a third aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a method for generating a camouflaged target based on a diffusion model as described in any one of the first aspect of the present invention is implemented.
[0049] According to a fourth aspect of the present invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for generating a camouflaged target based on a diffusion model as described in any one of the first aspects of the present invention is implemented.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] The present invention provides a camouflage target generation method based on a diffusion model, specifically a camouflage target generation scheme based on an attention fusion mechanism of a diffusion model. By constructing a data set of a detection target image, a background image and an image to be camouflaged, a pre-trained diffusion model is used for reasoning, and the implementation mechanism of attention is modified during the reasoning process, and finally the purpose of generating a camouflage target without providing additional background information is achieved, solving the limitation problem of the prior art scheme when generating a camouflage target without additional background information. Specifically, the diffusion model generation network disclosed in the present invention includes three stages. In the first stage, the image is reversed into a hidden variable of the corresponding number of steps under a given number of steps by using the image inversion technology; in the second stage, the diffusion model is inferred by using the hidden variable, and the attention module of the decoding layer of Unet includes an attention fusion and attention injection mechanism. In summary, in view of the problem that the existing camouflage target generation schemes require the corresponding background knowledge of the target area and the need for training, the method of the present invention does not require additional background knowledge and training, and can realize the generation of camouflage targets; based on the diffusion model framework, the present invention effectively realizes the generation of camouflage targets by adjusting the attention mechanism during model reasoning, and the module is simple, plug-and-play, without excessive dependence, and has strong applicability.
[0052] In the preferred embodiment of the present invention, the attention fusion mechanism establishes a mapping relationship between the generated image and the target image by calculating the attention weight matrix, and then uses the mapping relationship to inject the feature information of the background into the generated image; the attention injection mechanism directly injects the feature information of the background into the target image, and then uses the reconstruction item of the generated image to inject the feature into the generated image. Furthermore, the direct attention fusion calculation and indirect attention injection methods provided by the present invention can inject the feature information of the background into the target, so that the target has a camouflaged effect.
[0053] In the preferred embodiment of the present invention, after completing the reasoning of Unet, a content-based loss-guiding method is used to optimize latent variables, guide the consistency of the background of the generated image and the background of the original image, and reduce the content loss of the generated image and the original image. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below; obviously, the drawings described below are some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0055] Figure 1 is a flow chart of a method for generating a camouflaged target based on a diffusion model in an embodiment of the present invention;
[0056] Figure 2 is a schematic diagram of the principle of a method for generating a camouflaged target based on a diffusion model in an embodiment of the present invention;
[0057] Figure 3 It is a schematic diagram of a camouflage target generation system based on a diffusion model in an embodiment of the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention; it is obvious that the described embodiments and technical solutions are only part of the embodiments of the present invention, not all of the embodiments.
[0059] All other embodiments obtained by those of ordinary skill in the art without creative work based on the technical solutions disclosed in the embodiments of the present invention belong to the scope of protection of the present invention. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0060] See also Figure 1 and Figure 2 The embodiment of the present invention provides a method for generating a camouflaged target based on a diffusion model, comprising the following steps:
[0061] Step 1, obtaining a picture containing a target to be detected, and segmenting the picture containing the target to be detected to obtain a detection target image and a background image;
[0062] Step 2, based on the picture containing the target to be detected, the detection target image and the background image, a pre-trained diffusion model is used for inference to obtain a picture containing the camouflaged target;
[0063] Wherein, the diffusion model includes:
[0064] An encoder module, used for inputting the picture containing the target to be detected, the detection target image and the background image, encoding them, and outputting an initial state latent variable;
[0065] An image inversion module, used for inputting the initial state latent variables and performing image inversion processing, and outputting the latent variables after image inversion;
[0066] A Unet module is used to input the hidden variables after the image is inverted and diffuse multiple times based on a scheduler, and output the diffused hidden variables; wherein the scheduler is used to predict the hidden variables of the next stage according to the time step and noise; the Unet module includes an encoding layer and a decoding layer, the encoding layer and the decoding layer include the same number of Unet processing units, each of which includes a residual connected ResNet module, an attention module and a cross attention module; the attention module in each of the Unet processing units in the encoding layer adopts a self-attention mechanism; the attention module in each of the Unet processing units in the decoding layer adopts a self-attention mechanism before a predetermined number of time steps, and then adopts one or both of the attention direct fusion calculation method and the attention indirect injection method to inject the feature information of the hidden variables of the background image into the target to be detected in the picture containing the target to be detected;
[0067] The decoder module is used to input the diffused latent variable and decode it, and output a picture containing the disguised target.
[0068] The embodiment of the present invention provides a camouflage target generation method. The technical solution encodes the input image through the encoder module, extracts key features, and decodes the hidden variables through the decoder module to generate a camouflage target picture. This flexible encoding and decoding mechanism enables the model to adapt to different input and output requirements, and improves the controllability and diversity of the generation effect. In addition, the Unet module is the core part of the technical solution. It adopts the ResNet module, attention module and cross attention module with residual connection, improves the expression ability and generation quality of the model, and predicts the hidden variables of the next stage according to the time step and noise through the scheduler, so as to achieve fine control of the generation process. Furthermore, in the attention module of the decoding layer, the attention direct fusion calculation method and the attention indirect injection method are adopted to inject the hidden variable feature information of the background image into the target to be detected in the picture containing the target to be detected, so as to achieve the natural fusion of the background and the camouflage target, and improve the realism and consistency of the generated picture. The technical solution of the embodiment of the present invention can not only be applied to the generation of camouflage targets, but also can be extended to other image generation and editing tasks, such as image restoration, style migration, etc., and has potential application value in the fields of military camouflage and artistic creation.
[0069] Further explanation, compared with traditional methods such as generative adversarial networks (GANs), the diffusion model avoids problems such as mode collapse, the training process is more stable, and the generation quality is higher. By introducing the Unet module and the innovative background information fusion method, the technical solution further improves the controllability and diversity of the generation effect. Through the sophisticated noise addition and removal process and the advanced Unet module design, the embodiment method of the present invention can generate highly realistic and detailed camouflaged target images; at the same time, through the innovative background information fusion method, the natural fusion of the background and the camouflaged target is achieved, which improves the realism and consistency of the generated images.
[0070] In summary, the technical solution of the embodiment of the present invention provides a camouflaged target generation method based on the diffusion model, which has the advantages and significant progress of high-quality generation effect, flexible encoding and decoding mechanism, advanced Unet module design, innovative background information fusion method and broad application prospects.
[0071] In a specific embodiment of the present invention, the process of generating a disguised target may specifically include the following steps:
[0072] Step 1: Construct a soldier dataset for camouflage target generation; illustratively, the collected images meet the RGB three channels and contain manual annotation results, and the dataset size is 2048×2048; it is worth noting that the dataset does not have a camouflage effect, and the collected images are transmitted to the computer that executes the algorithm;
[0073] Step 2: Construct a diffusion model capable of generating camouflaged targets; wherein the diffusion model is divided into three stages, an image inversion stage, an attention injection stage, and a loss guidance stage; further, the diffusion model includes an encoder module for converting an image into a latent variable, a decoder module for converting a latent variable into an image, a Unet module for predicting noise, and a scheduler for predicting the latent variable of the next stage according to the time step and noise.
[0074] Specifically, in the image inversion stage, the image I containing the target to be detected is comp , detect target image I for and background image I bg As a batch input to the encoder module of the diffusion model, the representation of the latent variable of the image is obtained; then the latent variable is processed by image inversion and used as the input of the Unet module of the diffusion model to predict the noise and perform noise prediction and latent variable update according to the reverse DDIM (Denoising Diffusion Implicit Models) planner. The total number of steps is set to T, and the update steps are executed repeatedly until t = T to t = 0, and I is obtained. bg ,Ifor ,I comp Noisy representation of latent variables at the initial time.
[0075] Among them, the denoising process of reverse DDIM is expressed as follows:
[0076]
[0077] In the formula, z t represents the hidden variable at time t; σ t is a hyperparameter of the diffusion model; ε θ Represents the output result of Unet.
[0078] In the preferred technical solution, since the error will increase with the increase of the number of inversion steps during the image inversion process, and it is noted that the subsequent attention reorganization mechanism and the guidance based on content loss are used in the steps after time step t=30, so here when performing image inversion, the hidden variables at step t=30 are selected for inversion. In the subsequent reasoning process, the hidden variables at step t=30 are used as the initial state to start the reverse denoising process, which can significantly reduce the distortion of the image and the reduction of content quality, and improve the reconstruction quality.
[0079] In the embodiment of the present invention, in the denoising reasoning step of the diffusion model, an attention reorganization mechanism is introduced in the reasoning process of the Unet module, and content-based loss guidance is performed after the latent variable is updated; in a specific exemplary preferred scheme, the attention reorganization mechanism and content-based loss guidance are introduced in each step after time step t=30. Explanatory example, firstly, the latent variable is used as the input of the Unet module of the diffusion model. A layer of Unet processing units in the Unet module includes a Resnet module, a Self-Attention module, and a Cross-Attention module. These three modules are linearly connected in sequence, and the output of each module is added to the input in the form of a residual, and then injected into the next module as input. In a specific exemplary technical solution, there are a total of 16 layers in the Unet module. As the depth of each layer in the Unet increases, the resolution of the latent variable first decreases and then increases, and the feature dimension of each pixel first increases and then decreases; in the decoding layer part of the Unet module, that is, the part where the resolution increases and the number of feature dimensions decreases, the latent variable has a clearer feature representation capability and contains more complete structural information. In the Self-Attention module of some layers of the decoding layer of the Unet module, an attention injection mechanism is used, and the attention injection mechanism can include two parts.
[0080] The first part is in I compAdd an attention fusion term to the self-attention calculation, that is, use I comp Query, I for Key and I bg The value of I is used for attention calculation, and the result is compared with the original I comp This part is achieved by establishing I comp and I for The mapping relationship of the target area in I comp Query and I for The attention weight matrix established by the key Key is then used to directly convert I bg The value Value is injected into I in the form of self-attention calculation comp It is worth noting that I comp The reconstruction process is set as follows, that is, from I for Rebuild the target from I bg After this is achieved, the background is reconstructed. for The changes made will be reflected in I comp superior.
[0081] The second part is to use indirect attention injection, that is, first inject the background I bg The key value Key, Value is injected into the target I for Because I comp The reconstruction of the target part in the reconstruction item is to replace I for The key value is injected into I comp Among them, I for Since the feature information of the injected background is more similar to the background, it can indirectly bg The characteristic information of I is first injected for and then injected into I comp Specifically, use I for Query and I bg Key and I bg The value of Value is used for self-attention calculation.
[0082] The complete reasoning process for the final three images is as follows:
[0083]
[0084] In the formula, q comp ,k comp ,v comp The query, key and value, q, represents the image containing the object to be detected. for ,k for ,vfor Represents the query, key and value for detecting the target image, q bg ,k bg ,v bg represents the query, key and value of the background image. Attention refers to the scaled dot product attention mechanism, and softmax refers to the normalized exponential function; d k represents scaling, with a dimension of size K; w1 and w2 represent the weights of indirect injection and direct injection, respectively.
[0085] In the preferred technical solution of the embodiment of the present invention, I bg Background features are injected into I in the above two ways comp The goal is to make I comp The texture information of the target in is similar to that of the background. Furthermore, in terms of time, the attention reorganization mechanism is used in the steps after time step t=30, and in terms of space, the attention reorganization mechanism is used in the layers after the 12th layer of Unet.
[0086] In the preferred technical solution of the embodiment of the present invention, in order to solve the problem that the content of the disguised image and the original image may be greatly different during the injection process, a content-based loss guidance is proposed. Specifically, after the update is completed in each time step, the model is set to a mode that allows gradient propagation, and then the hidden variable z is re-entered t To Unet, get the result z t-1 , and then use the current z t-1 Predict the hidden variable z0 at time t=0; input z0 into the decoder to obtain the image reconstructed at the current time step Then use the reconstructed image and truth value I comp Calculate the mse loss L of the background area mse , using this loss to optimize z t-1 ; The formula is as follows:
[0087]
[0088] In the formula, x t represents the latent variable of the diffusion model at step t; σ is a predefined hyperparameter when training the diffusion model; ε θ is the parameterized network for predicting noise; t represents the number of steps of the current diffusion model inference; z t represents I at step t comp The corresponding hidden variable; η is the learning rate of gradient descent; downsample(·) represents the downsampling operation, z refers to the hidden variable containing the target image to be detected, L mse Refers to pixel-level difference loss.
[0089] After completing the set number of steps, the hidden variable corresponding to the final result is obtained, and the result is decoded using the diffusion model to obtain the final generated result.
[0090] In the technical solution of the embodiment of the present invention, the image generation technology based on the diffusion model is applied to the field of camouflage target style transfer, an attention reorganization mechanism is proposed to inject background features into the target, and a content-based loss guidance technical means is designed to improve the consistency of the background content of the camouflage image. According to the above two technical means, an improved diffusion model for camouflage target generation is designed. Further explanatory, the attention reorganization disclosed in the embodiment of the present invention is based on the calculation operation of self-attention, and the calculation method of attention is reconstructed during the reasoning process, and the background features are introduced into the target in the form of associated attention calculation. No training is required, and only the pre-trained diffusion model needs to be used for reasoning. The present invention can solve the problem of insufficient data sets in the field of camouflage target cognition, and the problem that the field of camouflage target generation requires additional background knowledge corresponding to the target area for calculation; in addition, compared with the traditional camouflage target generation method, the complexity of generation is reduced and the quality of generation is improved.
[0091] In a specific embodiment of the present invention, the target dataset used contains a total of 1,300 RGB color images of soldiers taken in real scenes, with an image size of 2048×2048 pixels; the region of the target in each image is annotated using an artificial dataset, and the annotation result is represented by a binary image. Image I comp In the preprocessing stage, the image is first scaled to 512×512 pixels, and then the image is additionally cropped to an image containing only the target according to the annotation results. for and an image containing only the background I bg .
[0092] The method of initializing the network parameters of the diffusion model in the embodiment of the present invention is to use the pre-trained StableDiffusion to initialize the parameters. The operating environment is a computer with a framework such as Pytorch, which can read a given picture and complete the reasoning calculation process of the model of this method. It takes about one minute for the embodiment of the present invention to infer an image on a L20 (48GB) GPU. The specific implementation steps include: obtaining the dependency library information required by the diffusion model, and loading the pre-trained diffusion model, including Unet, encoder, decoder, text encoder and scheduler, etc.; initializing the attention fusion module and the mse-based loss guidance module, and setting the proportion of attention injection and the learning rate of loss; then registering the attention fusion module to the diffusion model, specifically, traversing all modules of the diffusion model, finding the self-attention module therein, and then replacing it with the attention fusion module. It is worth noting that there is a counter inside the attention fusion module to record the current number of layers and time steps to determine the attention mechanism that should be used. Set the total number of inference steps of the diffusion model to 50, the number of steps for starting the attention reorganization mechanism and the MSE-based loss guidance mechanism to 30, and the number of starting layers to 12. Load the preprocessed image and convert I bg , I for , I comp Set it as a batch as input, and then calculate the state z of the hidden variable corresponding to the 30th step bg , z for , z comp , the hidden variables are used as inputs to the Unet of the diffusion model. The Unet with attention reorganization mechanism calculates the noise output, then updates the hidden variables, and then sets the diffusion model to a mode that allows gradient propagation, recalculates the noise output, and then calculates the gradient and updates the hidden variables. After completing all the time steps, the decoder of the diffusion model is used to obtain the final generated result.
[0093] The embodiment of the present invention discloses a camouflaged target generation method based on a diffusion model, which uses the attention mechanism of the diffusion model to inject background features into the target, and uses a loss-guided method to ensure that the background content of the image remains unchanged. The method converts a non-camouflaged target in an image into a camouflaged target in a training-free manner without providing a background corresponding to the target area, and each module is simple, does not require excessive reliance, and has strong applicability.
[0094] In a specific embodiment of the present invention, a target image I to be detected that needs style transfer is obtained. comp , the size is (c, h, w), c represents the number of channels, h and w represent the height and width of the image respectively; then the data annotation method is used to obtain the segmented image I of the target maskThe segmented image is binarized. Among its pixels, the area with a value of 1 represents the target, and the area with a value of zero represents the background. The target image is preprocessed using the segmented image to obtain an image containing only the target. for and the background-only image I bg , where all images have the same size of (c, h, w).
[0095] I for =I comp I mask
[0096]
[0097] Here, the · operator represents positional multiplication.
[0098] Three representations of a camouflaged image are obtained. After processing all the data sets, each picture has an image containing only the target and an image containing only the background.
[0099] After obtaining the image, to apply the real image to the diffusion model, it is necessary to use image inversion technology. The so-called image inversion technology refers to using the reverse diffusion model to predict the noise at each step and add noise to the image until it becomes pure z comp ,z for ,z bg noise.
[0100]
[0101] In the formula, x t represents the latent variable of the diffusion model at step t; σ is a predefined hyperparameter when training the diffusion model. For different time steps, the added The value of σ is different; ε θ is a parameterized network for predicting noise, implemented as a Unet in the diffusion model; I comp ,I for ,I bg As a batch input to the diffusion model, each image is preprocessed into a (3, 512, 512) image, and the image is first encoded into the latent space using Encoder.
[0102] z comp ,z for ,z bg =D(I comp ),D(I for ),D(I bg )
[0103] Among them, z comp ,zfor ,z bg The dimension is (4, 64, 64).
[0104] Then the latent variable after the encoder is encoded is inverted, and after setting the number of steps to 50, pure noise that can be reconstructed into the corresponding image is obtained until the inference becomes
[0105]
[0106] After obtaining the noise of the corresponding image, the noise corresponding to the three images is input into the diffusion model as a batch to start the inference process of the diffusion model. The standard denoising process of the diffusion model using the DDIM sampler is as follows:
[0107]
[0108] Among them, β is a predefined hyperparameter,
[0109] In the early stage of denoising, the image is very blurry and the structural information is not obvious. Adding noise at this time will lead to completely different results. Therefore, the embodiment of the present invention adds an attention injection mechanism in step 30.
[0110] In the embodiment of the present invention, the noise predictor ε of the diffusion model θ (z t ,t) is a Unet structure. There are sixteen modules in a Unet, each of which contains a cross-attention mechanism module, a self-attention mechanism module and a ResNet. Some studies have shown that the self-attention module can control the generation of image feature information.
[0111] A standard self-attention process is as follows,
[0112]
[0113] Among them, Q stands for query; K stands for key; V stands for value; d k Represents scaling, whose size is the dimension of K.
[0114] In the reasoning process of the diffusion model, the implicit state of self-attention in Unet is expressed as (h comp ,h for ,h bg). In the diffusion model, the implicit state is first converted to the corresponding query Query, key Key, and value Value through the weight matrix to_q, to_k, and to_v; the shape of the implicit state is (batch_size, seq_length, input_dim), batch_size refers to the input batch, seq_length refers to the length of the input data after serialization, and input_dim refers to the dimension of the input.
[0115] q comp ,k comp ,v comp = to_q(h comp ),to_k(h comp ),to_v(h comp )
[0116] q for ,k for ,v for = to_q(h for ),to_k(h for ),to_v(h for )
[0117] q bg ,k bg ,v bg = to_q(h bg ),to_k(h bg ),to_v(h bg )
[0118] The reasoning process of the original diffusion model is to calculate the attention on the implicit representation of each image one by one.
[0119]
[0120] After the final inference is completed, the noise of one inference is obtained; the noise is brought into the scheduler corresponding to the diffusion model for calculation, and the final result is obtained after the number of loop steps is set. This process is called the reconstruction process. Because no new information is introduced at this time, the noise obtained by the image inversion process is just reconstructed into the original image.
[0121] Among the multiple modules of Unet, the dimensions of the modules in Unet are from large to small, and then from small to large. The process of modules from large to small is called the encoding layer of Unet, and the process of modules from small to large is called the decoding layer of Unet. In the decoder part of Unet, the image has more structural information and feature information. Therefore, in the spatial dimension, an improved mechanism related to attention is added in the decoder part.
[0122] In the reasoning steps, we found that the structural information and feature information of the image were not obvious from step 0 to step 25, but the feature information of the image gradually began to emerge from step 25 to step 50. Therefore, in the time dimension, we first added the attention-related improvement mechanism from step 25. We found that using I for , I bg Attention to reorganize I comp The attention calculation process can also realize I comp The specific method of reconstruction is as follows:
[0123]
[0124] Since the dot product is performed between the query and the key in the attention calculation process, the tensor shape of the query is (batch_size, seq_length, dk), and the tensor shape of the key is (batch_size, seq_length, dk), we can get the attention weight matrix Sim, the tensor shape of Sim is (batch_size, seq_length, seq_length), which represents the attention weight corresponding to the key for each query. In this calculation process, the query and the key can be calculated even if they are inferred from different images, so as to obtain the attention weight matrix Sim, so we can use I comp Query and I for The key and value of the key, plus I comp Query and I bg The key Key and value Value. Reconstruct I comp .
[0125] In the technical solution of the embodiment of the present invention, the attention reorganization mechanism is used to achieve the approximation of target features and background features. comp Query and I for The key Key establishes an attention weight matrix. If we use the value Value here for the final attention calculation, the target will be rebuilt. At this time, the value Value represents the value of the matrix when the key is established from I comp To I for After the attention weight matrix, the information injected; therefore, I is used here bg The value Value is used as the value of the weight matrix.
[0126]
[0127] In the decoder part of Unet, from step 25 to step 50 of the diffusion model inference, I comp The reasoning process is set up as follows:
[0128]
[0129] Among them, w2 represents the weight of the direct fusion attention calculation item, and the default value is 1. Increasing this weight can achieve better camouflage effect, and reducing this weight can retain more target information.
[0130] In the technical solution of the embodiment of the present invention, an attention fusion mechanism is used to inject the environmental features of the background into the target. for and I bg In the reasoning process, I bg The corresponding key Key and value Value are injected into I for Among them, use I for Query and I bg The key and value are used for attention calculation. Then the result is multiplied by a weight and added to I for The reasoning part of attention.
[0131] The formula is as follows:
[0132]
[0133] Among them, w1 represents the weight of indirect attention injection;
[0134] The complete reasoning process for the final three images is as follows:
[0135]
[0136] In the formula, q comp , k comp 、v comp Respectively represent the query, key, and value of the image containing the target to be detected; q for , k for 、v for Respectively represent the query, key and value of the target image to be detected; q bg , k bg 、v bg denote the query, key, and value of the background image, respectively; Attention denotes the scaled dot product attention mechanism; Softmax refers to the normalized exponential function; d k represents scaling, with a dimension of size K; w1 and w2 represent the weights of indirect injection and direct injection, respectively.
[0137] It is worth noting that in the process of reasoning, the target and the background are obtained by intercepting the mask, and the spatial positions of the target and the background do not have any intersection. In other words, the general method based on traditional neural networks such as VGG network extraction with spatial relationship cannot inject the feature information of the background into the target. In the calculation of self-attention, we use I for I M and To indicate I for , I bg The corresponding target area and background area, I for , I bg We use 0 to represent the pixels in other areas of the . In the calculation, only the element values in the weight matrix obtained by the attention calculation of the non-zero area are high. This is because the area with a pixel value of 0 approaches 0 after being mapped to the query matrix. In addition, the attention mechanism calculates the relationship between each query Query and each key Key. In the process of calculating the attention weight matrix, I bg I M This calculation method allows us to inject attention when the spatial positions between the two images do not overlap.
[0138] In the technical solution provided by the embodiment of the present invention, a method guided by an MSE-based loss is used to achieve consistency in the content of the style transfer image with the original image. The only content that changes is the target, and the direction of change is only the direction of camouflage. In the reconstruction process of the diffusion model, we chose the DDIM Inversion method, which can invert the image into pure noise, and then re-infer it into an image. However, in the inversion process, we cannot accurately restore the image, which will greatly damage the performance of our method. At the same time, with the reasoning of the diffusion model, after using the above-mentioned attention injection mechanism, the image information and the reconstructed information will also deviate. Therefore, we proposed a content loss based on MSE to ensure the consistency of image information.
[0139] Specifically, the original image is referred to as I ref , the reconstructed image is called I rec An intuitive idea is that we can establish the MSE (Mean Squared Error) Loss of the two as a guide to ensure that the generated image can be close to the original image content. And after we implement the attention injection mechanism, we can achieve a good camouflage effect while the image background content is basically the same.
[0140] Given an image and The calculation method of MSE is as follows:
[0141]
[0142] However, there are two problems here. First, in the reasoning process of the diffusion model, the denoised result x at each step t Both are noisy versions. If you compare the MSE with the noisy version, the result will be inaccurate. Second, when performing selfattention injection, we hope that the texture information of the target will change, that is, the target information will change in the direction of camouflage. If a simple MSE loss is used, it will eventually affect the camouflage properties, causing the entire image to be consistent with the original image.
[0143] To solve the first problem, let's review the sampling mechanism of DDIM. First, the modeling process of the diffusion model is as follows: For the forward process, first, according to the Markov theorem, the state at the current moment is only related to the state at the previous moment, which is converted into the noise relationship between each moment and the previous moment.
[0144]
[0145] For the reverse process, given the state at each moment, we need to predict the state at the previous moment, model the mean and variance, and use neural network prediction (that is, Unet) to get the results we need.
[0146]
[0147] Since the forward process is fixed, given x0, we can directly calculate x t result.
[0148]
[0149] In the process of reverse reasoning, when we know the current state, we can use the above formula to derive the result of x0 based on the current prediction. ref By performing calculations, we can eliminate the problem that the MSE loss cannot be calculated well due to the noise in the intermediate process of the diffusion model.
[0150]
[0151] Among them, ε θ (z t ,t) is a parameterized neural network, namely Unet, which is used to predict the noise at each step.
[0152] The final calculation process is as follows:
[0153]
[0154] Where D refers to the decoder part in the diffusion model, and the predicted After decoding, the predicted The with I ref Calculate the MSE and use this loss as a guide. After the 30th step of the diffusion model, the input noise z t After a Unet, the noise is predicted and updated to z t-1 Later, the model is set up to allow gradient propagation and reuse z in the process t Perform inference, use Unet to perform inference, predict noise, and get z t-1 Later, predict x0 and calculate the MSE loss, and finally use this loss to calculate the relative t gradient.
[0155]
[0156] The second problem is that when we use the improved mechanism of attention, the target area will change because its features will be more similar to the background, and this change will also be reflected in L mse In the gradient relative to z, that is, when we calculate the MSE loss and update z, the target that has been camouflaged will also move in the direction of the target that has not been camouflaged. Therefore, we propose to use a mask to limit the background part to the area where gradient descent can be performed.
[0157] Specifically, given I rec ,I mask And a downsampling factor downsample, when we calculate the gradient, we limit the calculation of gradient propagation to the background area.
[0158]
[0159] The final noise update process for one time step is as follows:
[0160]
[0161] Here, η refers to the learning rate of the gradient descent.
[0162] In a preferred embodiment of the present invention, we found that when using image inversion, the more steps of inversion, the greater the error accumulated in each step, and the worse the quality of the final reconstruction. Looking back at all the above techniques, we can find that all improvements are based on time steps after step 25, so when we perform image inversion, we only invert to step 30, get the expression of the latent variable of step 25, and then use the latent variable of step 25 to perform the reverse denoising process. The content retention of the reconstructed image is significantly better than inverting to step 0 and then starting reconstruction.
[0163] The following are device embodiments of the present invention, which can be used to implement the method embodiments of the present invention. For details not disclosed in the device embodiments, please refer to the method embodiments of the present invention.
[0164] See also Figure 3 In an embodiment of the present invention, a camouflage target generation system based on a diffusion model is provided, comprising:
[0165] A data acquisition module is used to acquire a picture containing a target to be detected, and to segment the picture containing the target to be detected to obtain a detection target image and a background image;
[0166] A disguised target generation module, used to perform reasoning using a pre-trained diffusion model based on the picture containing the target to be detected, the detection target image and the background image, to obtain a picture containing the disguised target;
[0167] Wherein, the diffusion model includes:
[0168] An encoder module, used for inputting the picture containing the target to be detected, the detection target image and the background image, encoding them, and outputting an initial state latent variable;
[0169] An image inversion module, used for inputting the initial state latent variables and performing image inversion processing, and outputting the latent variables after image inversion;
[0170] A Unet module is used to input the hidden variables after the image is inverted and diffuse multiple times based on a scheduler, and output the diffused hidden variables; wherein the scheduler is used to predict the hidden variables of the next stage according to the time step and noise; the Unet module includes an encoding layer and a decoding layer, the encoding layer and the decoding layer include the same number of Unet processing units, each of which includes a residual connected ResNet module, an attention module and a cross attention module; the attention module in each of the Unet processing units in the encoding layer adopts a self-attention mechanism; the attention module in each of the Unet processing units in the decoding layer adopts a self-attention mechanism before a predetermined number of time steps, and then adopts one or both of the attention direct fusion calculation method and the attention indirect injection method to inject the feature information of the hidden variables of the background image into the target to be detected in the picture containing the target to be detected;
[0171] The decoder module is used to input the diffused latent variable and decode it, and output a picture containing the disguised target.
[0172] In a third aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a method for generating a camouflaged target based on a diffusion model as described in any one of the first aspect of the present invention is implemented.
[0173] In one embodiment of the present invention, a computer device is provided, the computer device comprising a processor and a memory, the memory being used to store a computer program, the computer program comprising program instructions, and the processor being used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, which are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions in a computer storage medium to implement a corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used to perform the operation of a camouflage target generation method based on a diffusion model.
[0174] In one embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM (Random Access Memory) memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the method for generating a camouflaged target based on a diffusion model in the above embodiment.
[0175] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) that contain computer-usable program code.
[0176] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0177] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for generating a camouflaged target based on a diffusion model, characterized in that: The following steps are involved: Acquire a picture containing a target to be detected, and segment the picture containing the target to be detected to obtain a detection target image and a background image; Based on the picture containing the target to be detected, the detection target image and the background image, a pre-trained diffusion model is used for reasoning to obtain a picture containing the camouflaged target; Wherein, the diffusion model includes: An encoder module, used for inputting the picture containing the target to be detected, the detection target image and the background image, encoding them, and outputting an initial state latent variable; An image inversion module, used for inputting the initial state latent variables and performing image inversion processing, and outputting the latent variables after image inversion; A Unet module is used to input the hidden variables after the image is inverted and diffuse multiple times based on a scheduler, and output the diffused hidden variables; wherein the scheduler is used to predict the hidden variables of the next stage according to the time step and noise; the Unet module includes an encoding layer and a decoding layer, the encoding layer and the decoding layer include the same number of Unet processing units, each of which includes a residual connected ResNet module, an attention module and a cross attention module; the attention module in each of the Unet processing units in the encoding layer adopts a self-attention mechanism; the attention module in each of the Unet processing units in the decoding layer adopts a self-attention mechanism before a predetermined number of time steps, and then adopts one or both of the attention direct fusion calculation method and the attention indirect injection method to inject the feature information of the hidden variables of the background image into the target to be detected in the picture containing the target to be detected; The decoder module is used to input the diffused latent variable and decode it, and output a picture containing the disguised target.
2. The method for generating a camouflaged target based on a diffusion model according to claim 1, characterized in that: The step of injecting the characteristic information of the hidden variable of the background image into the target to be detected in the picture containing the target to be detected by using one or both of the direct attention fusion calculation method and the indirect attention injection method, the step of injecting the characteristic information of the hidden variable of the background image into the target to be detected in the picture containing the target to be detected by using the direct attention fusion calculation method specifically includes: Introduce an attention fusion term in the self-attention calculation of the image containing the target to be detected, use the query Query of the image containing the target to be detected, the key of the target image and the value of the background image to perform attention calculation, and add the attention calculation result to the reconstruction process of the original image containing the target to be detected; A mapping relationship between a picture containing a target to be detected and a target area in the detection target image is established, where the mapping relationship is an attention weight matrix established by a query Query of the picture containing the target to be detected and a key Key of the detection target image; using the mapping relationship, the value Value of the background image is directly injected into the picture containing the target to be detected in the form of self-attention calculation.
3. The method for generating a camouflaged target based on a diffusion model according to claim 1, characterized in that: In the step of injecting the feature information of the hidden variable of the background image into the target to be detected in the picture containing the target to be detected by using one or both of the direct attention fusion calculation method and the indirect attention injection method, The direct attention fusion calculation method is to directly use the query containing the image of the target to be detected, the key of the target image and the value of the background image to perform self-attention calculation; The indirect attention injection method is to inject the key and value of the background image into the detection target image; The complete reasoning process for the final three images is as follows: In the formula, q comp , k comp 、v comp Respectively represent the query, key, and value of the image containing the target to be detected; q for , k for 、v for Respectively represent the query, key and value of the target image to be detected; q bg , k bg 、v bg denote the query, key, and value of the background image, respectively; Attention denotes the scaled dot product attention mechanism; Softmax refers to the normalized exponential function; d k represents scaling, with a dimension of size K; w1 and w2 represent the weights of indirect injection and direct injection, respectively.
4. The method for generating a camouflaged target based on a diffusion model according to claim 1, characterized in that: In the step of performing multiple diffusions based on the scheduler, in each time step after a predetermined number of time steps, after the latent variables of the picture containing the target to be detected are updated, content-based loss guidance is performed; The steps of content-based loss guidance are: set the mode to allow gradient propagation, re-set the hidden variable z of the image containing the target to be detected t Input into the Unet module and get the result z t-1 , using the current z t-1 Predict the hidden variable z0 at time t=0; input the hidden variable z0 into the decoder module to obtain the image reconstructed at the current time step Using the reconstructed image and the image I containing the target to be detected comp Calculate the content loss L of the background area mse , using the calculated loss to optimize z t-1 ; The loss calculation expression is: In the formula, x t represents the latent variable of the diffusion model at step t; σ is a predefined hyperparameter when training the diffusion model; ε θ is the parameterized network for predicting noise; t represents the number of steps of the current diffusion model inference; z t represents I at step t comp The corresponding hidden variable; η is the learning rate of gradient descent; downsample(·) represents the downsampling operation; z refers to the hidden variable containing the target image to be detected; L mse Refers to pixel-level difference loss.
5. The method for generating a camouflaged target based on a diffusion model according to claim 4, characterized in that: The Unet module is provided with 16 layers of Unet processing units in total. As the depth of the layer in Unet increases, the resolution of each layer of Unet processing units first decreases and then increases, and the feature dimension of each pixel first increases and then decreases; wherein, the decoding layer of the Unet module is the part where the resolution increases and the number of feature dimensions of the pixels decreases.
6. The method for generating a camouflaged target based on a diffusion model according to claim 5, characterized in that: The attention module in each of the Unet processing units of the decoding layer adopts a self-attention mechanism before a predetermined time step, and then adopts one or both of an attention direct fusion calculation method and an attention indirect injection method to inject feature information of latent variables of the background image into the target to be detected in the picture containing the target to be detected, In terms of time, in the steps after time step t=30, one or both of the direct attention fusion calculation method and the indirect attention injection method are used to inject the feature information of the latent variable of the background image into the target to be detected in the picture containing the target to be detected; Spatially, in the Unet processing units after the 12th layer of the Unet module, one or both of the attention direct fusion calculation method and the attention indirect injection method are used to inject the feature information of the latent variables of the background image into the target to be detected in the picture containing the target to be detected.
7. The method for generating a camouflaged target based on a diffusion model according to claim 6, characterized in that: In the image inversion module, when performing image inversion, the hidden variables corresponding to the background image, the detection target image, and the target image to be detected are inverted to the hidden variables corresponding to the time t=30, and then in the subsequent reasoning process, the hidden variables at the time t=30 are used as the initial state to perform the reverse denoising process.
8. A camouflage target generation system based on a diffusion model, characterized in that: include: A data acquisition module is used to acquire a picture containing a target to be detected, and to segment the picture containing the target to be detected to obtain a detection target image and a background image; A disguised target generation module, used to perform reasoning using a pre-trained diffusion model based on the picture containing the target to be detected, the detection target image and the background image, to obtain a picture containing the disguised target; Wherein, the diffusion model includes: An encoder module, used for inputting the picture containing the target to be detected, the detection target image and the background image, encoding them, and outputting an initial state latent variable; An image inversion module, used for inputting the initial state latent variables and performing image inversion processing, and outputting the latent variables after image inversion; A Unet module is used to input the hidden variables after the image is inverted and diffuse multiple times based on a scheduler, and output the diffused hidden variables; wherein the scheduler is used to predict the hidden variables of the next stage according to the time step and noise; the Unet module includes an encoding layer and a decoding layer, the encoding layer and the decoding layer include the same number of Unet processing units, each of which includes a residual connected ResNet module, an attention module and a cross attention module; the attention module in each of the Unet processing units in the encoding layer adopts a self-attention mechanism; the attention module in each of the Unet processing units in the decoding layer adopts a self-attention mechanism before a predetermined number of time steps, and then adopts one or both of the attention direct fusion calculation method and the attention indirect injection method to inject the feature information of the hidden variables of the background image into the target to be detected in the picture containing the target to be detected; The decoder module is used to input the diffused latent variable and decode it, and output a picture containing the disguised target.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for generating a camouflaged target based on a diffusion model according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a camouflaged target based on a diffusion model according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Camouflage target detection method based on complementary perception cross-view fusion network
CN117593517A
Camouflage image generation method based on knowledge retrieval and reasoning enhancement
CN118052899A
Camouflage target detection method and system based on pyramid type visual Transform
CN118172540A
Regitive digital camouflage generation method, device and equipment
CN119251043A
Special latent image pattern structure, creation method of data for special latent image pattern structure
JP2019147245A
Cited By
Biologically inspired camouflage image generation method, system and equipment
CN120877022A