Camouflage image generation method and device, equipment and medium
Through the basic diffusion network, attention network and control network in the diffusion model, the problem of high cost of camouflage image generation in the prior art is solved, and high-precision camouflage image generation without training is achieved.
Patent Information
- Application Number
- CN202510414598.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Existing methods for camouflage image generation rely on a large amount of training data, resulting in excessive cost of generating camouflage images.
The diffusion model is adopted, including the basic diffusion network, attention network and control network, and the noise image is denoised by multiple rounds and the features of the background image and foreground image are aggregated into the camouflage image to achieve training-free camouflage image generation.
The cost of camouflage image generation is reduced and high-precision and high-fidelity camouflage image can be generated.
Smart Images

Figure CN119941909A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of algorithm technology, and in particular to a method, device, equipment and medium for generating a camouflage image. Background Art
[0002] Biological camouflage in nature is an important mechanism for the survival of organisms, which has inspired the widespread application of camouflage technology in the fields of military, artworks, and security, and also stimulated the development of anti-camouflage visual perception. However, the scarcity of camouflaged images greatly limits the effectiveness of anti-camouflage strategies, which has stimulated people's interest in camouflage image generation.
[0003] However, current camouflaged image generation methods rely on a large amount of training data to train the model, making the cost of generating camouflaged images too high. Summary of the invention
[0004] The present invention aims at the problem that the cost of generating a camouflage image is too high in the prior art, and proposes a method, device, equipment and medium for generating a camouflage image.
[0005] In a first aspect, an embodiment of the present invention provides a method for generating a camouflage image, the method comprising: Acquire an image data set, where the image data set includes a background image and a foreground image; Input the image dataset into the pre-built diffusion model to generate camouflaged images; Among them, the diffusion model includes the basic diffusion network, the attention network and the control network; The basic diffusion network is used to perform multiple rounds of denoising on the noisy image generated based on the background image; the attention network is used to aggregate the features of the noisy image, background image and foreground image; and the control network is used to control the denoising process of the basic diffusion network.
[0006] In one embodiment, the basic diffusion network includes: ; ; Where d is the differential symbol, represents the stochastic differential equation, represents the noisy image after the tth round of denoising, and represents the standard Brownian motion process, and denote the drift function and the fluctuation function respectively, represents the scoring function, represents the instantaneous loss, and represents the hyperparameter, represents the control network, represents the first terminal loss, represents the second terminal loss, To pretend to be restrained, Indicates that the control network satisfies the masquerade constraint Constraints, represents mathematical expectation.
[0007] In one embodiment, the first terminal loss includes: ; in, Represents the Tweedie formula The camouflaged image obtained by approximation, Represents the background image, represents the best camouflage loss indicator, To pretend to be restrained, Indicates that the control network satisfies the masquerade constraint constraints.
[0008] In one embodiment, the image dataset further includes a foreground mask, and the second terminal loss is constructed based on the foreground mask.
[0009] In one embodiment, the second terminal loss includes: ; in, Represents the Tweedie formula The camouflaged image obtained by approximation, Represents the anti-detection descriptor, Represents the foreground mask.
[0010] In one embodiment, the attention network specifically includes: Linear mapping is performed on the noise image, the background image and the foreground image respectively to obtain first feature data of the noise image, second feature data of the background image and third feature data of the foreground image; The first attention output is determined according to the first feature data and the second feature data, and the calculation formula includes: ; in, represents the first attention output, represent the query, key and value of the first feature data respectively, denote the key and value of the second feature data respectively, and D denotes the latitude of the key; The second attention output is determined according to the first feature data and the third feature data, and the calculation formula includes: ; in, represents the second attention output, denote the key and value of the third characteristic data respectively, and D denotes the latitude of the key; Aggregate based on the first attention output and the second attention output, the calculation formula includes:
[0011] in, Represents the result of feature aggregation.
[0012] In one embodiment, the image dataset also includes text data; and the attention network is further used to perform feature aggregation on the noisy image and text data.
[0013] In a second aspect, an embodiment of the present invention provides a device for generating a camouflage image, the device comprising: An acquisition module is used to acquire an image data set, where the image data set includes a background image and a foreground image; A generation module, used for inputting an image dataset into a pre-built diffusion model to generate a camouflaged image; Among them, the diffusion model includes the basic diffusion network, the attention network and the control network; The basic diffusion network is used to perform multiple rounds of denoising on the noisy image generated based on the background image; the attention network is used to aggregate the features of the noisy image, background image and foreground image; and the control network is used to control the denoising process of the basic diffusion network.
[0014] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a processor and a memory, wherein the memory stores at least one message, at least one program, a code set or a message set, and the at least one message, at least one program, a code set or a message set is loaded and executed by the processor to implement a method for generating a camouflaged image as described in any one of the first aspects.
[0015] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which at least one message or at least one program is stored, and the at least one message or at least one program is loaded and executed by a processor to implement a method for generating a camouflaged image as described in any one of the first aspects.
[0016] The present invention provides a method, device, equipment and medium for generating a camouflaged image, which has the following technical effects: The constructed noisy image is denoised through the basic diffusion network. During the denoising process, the features of the foreground image and the background image are aggregated into the camouflaged image through the attention network. At the same time, a control network is constructed to optimize the denoising process of the basic diffusion network, thereby realizing the generation of camouflaged images without training. There is no need for a large number of training samples to train the model, which reduces the generation cost of the camouflaged image and can obtain high-precision and high-fidelity camouflaged images.
[0017] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings: Figure 1 A schematic flow chart of a method for generating a camouflaged image provided by an exemplary embodiment of the present invention; Figure 2 A schematic diagram of a diffusion model provided for an exemplary embodiment of the present invention; Figure 3 A schematic diagram of a multi-head attention sub-network provided for an exemplary embodiment of the present invention; Figure 4 A schematic diagram of a mask-restricted subnetwork provided for an exemplary embodiment of the present invention; Figure 5 A schematic diagram of comparing test results provided by an exemplary embodiment of the present invention; Figure 6 A schematic diagram of a module of a camouflage image generation device provided by an exemplary embodiment of the present invention; Figure 7 The present invention provides a schematic structural diagram of an electronic device according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0019] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that the embodiments and features in the embodiments of the present invention can be combined with each other without conflict.
[0020] The following detailed description is an exemplary description, which is intended to provide further detailed description of the present invention. Unless otherwise specified, all technical terms used in the present invention have the same meaning as those generally understood by those skilled in the art to which the present invention belongs. The terms used in the present invention are only for describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present invention.
[0021] An exemplary embodiment of the present invention provides a method for generating a camouflage image. Figure 1 , methods include: Step S101: Acquire an image data set.
[0022] The image dataset includes background images and foreground images.
[0023] Step S102: input the image data set into the pre-built diffusion model to generate a camouflaged image.
[0024] Among them, the diffusion model includes basic diffusion network, attention network and control network.
[0025] Specifically, the basic diffusion network is used to perform multiple rounds of denoising on the noisy image generated based on the background image, the attention network is used to aggregate the features of the noisy image, the background image and the foreground image, and the control network is used to control the denoising process of the basic diffusion network.
[0026] In this embodiment, the constructed noisy image is denoised through the basic diffusion network. During the denoising process, the features of the foreground image and the background image are aggregated into the camouflage image through the attention network. At the same time, the control network controls the denoising process of the basic diffusion network to achieve the generation of camouflage images without training. There is no need for a large number of training samples to train the model, which reduces the generation cost of the camouflage image and can obtain high-precision and high-fidelity camouflage images.
[0027] In one embodiment, see Figure 2 , the basic diffusion network of the diffusion model is explained: Before the basic diffusion network performs denoising, the basic diffusion network is also used to generate a noise image based on the background image. The generation process of the noise image uses the forward random process of the basic diffusion network. The forward random process can be understood as the process of adding random Gaussian noise to the background image T times cumulatively, which can be expressed as the solution of the stochastic differential equation, which is expressed as follows: (1) in, and They are called drift function and fluctuation function respectively.
[0028] The denoising process of the noisy image uses the reverse random process of the basic diffusion network. The reverse random process can be understood as the solution of the reverse stochastic differential equation from T~0 time step under a slight regularity, which can be expressed by the following formula: (2) in, and represents the standard Brownian motion process, and represent the drift function and the fluctuation function respectively. Represents the scoring function, which is used in the image generation task through the U-Net neural network of the diffusion model Approximately obtain.
[0029] Based on the above two random processes, the basic diffusion network in the embodiment of the present invention specifically includes: (3) (4) Where d is the differential symbol, represents the stochastic differential equation, represents the noisy image after the tth round of denoising, is the generated image after denoising. and represents the standard Brownian motion process, and represent the drift function and the fluctuation function respectively. represents the scoring function, represents the instantaneous loss, and represents the hyperparameter, represents the control network, represents the first terminal loss, represents the second terminal loss, To pretend to be restrained, Indicates that the control network satisfies the masquerade constraint Constraints, It represents mathematical expectation. The camouflage constraint is set according to the actual diffusion scenario and is not particularly limited here.
[0030] The denoising process of the basic diffusion network is controlled by introducing a control network in the basic diffusion network. The goal of the control network is to make the area where the foreground image is located in the generated camouflage image visually merge with the background in the image, and the synthesized camouflage image can effectively evade the detection of the camouflage detector. In this embodiment, these two goals are defined as the desired terminal conditions of the diffusion model, so that the control network can modify the drift field that changes with time in the basic diffusion network during the denoising process, thereby effectively guiding the denoising process to achieve these two desired terminal conditions.
[0031] The above two targets correspond to the first terminal loss and the second terminal loss respectively, and the first terminal loss and the second terminal loss are determined by the control network.
[0032] In one embodiment, since the scoring function It is a U-Net neural network that can be used in image generation tasks through the diffusion model Approximately, represents the diffusion step, represents the network parameters, so formula (4) can be understood as mainly composed of the control network The total loss is minimized by optimizing the drift of the reverse random process. The total loss is defined as arrive Points : ; At this point, minimizing the total loss is converted to solving Minimum value of ,Right now: ; This objective can be obtained using the Hamilton-Jacobi-Bellman equation, which is used to find the optimal control strategy in a continuous-time system that minimizes some performance metric or cost function. The Hamilton-Jacobi-Bellman equation states that if the minimum There are continuous partial derivatives, which must satisfy the following partial differential equation: ; Therefore, when →∞, the optimal controller can be obtained by solving the Hamilton-Jacobi-Bellman equation, which allows us to ignore the transient cost. Therefore, the present invention only needs to focus on the optimization of the two terminal losses to implement formula (3).
[0033] Specifically, the first terminal loss includes: (5) in, Represents the Tweedie formula The camouflaged image obtained by approximation, Represents the background image, represents the best camouflage loss indicator, To pretend to be restrained, Indicates that the control network satisfies the masquerade constraint The camouflage constraint is set according to the actual diffusion scenario and is not specifically limited here.
[0034] Tweedy's formula is used to approximate the camouflaged image , and use the consistent style extractor as the best disguise loss identifier To minimize the background image and camouflage images The visual difference between the foreground and background images makes the foreground image in the generated camouflaged image visually integrated with the background in the image.
[0035] Specifically, the image data set also includes a foreground mask, and the second terminal loss is constructed based on the foreground mask. The second terminal loss includes: (6) in, Represents the Tweedie formula The camouflaged image obtained by approximation, Represents the anti-detection descriptor, Represents the foreground mask.
[0036] The camouflaged image is approximated based on Tweedy's formula Input to the anti-detection descriptor In , a pre-trained disguised object detector is used for initialization to minimize the anti-detection result and the foreground mask The consistency between them makes the synthesized camouflage image able to effectively evade the detection of the camouflage detector.
[0037] In one embodiment, see Figure 2-4 , an attention network is described, and the attention network may include at least one of a multi-head attention sub-network and a mask-restricted attention sub-network.
[0038] Specifically, the multi-head attention sub-network is used to aggregate features of the noise image, background image, and foreground image, see Figure 3 , the multi-head attention sub-network (Multi-head Attention Aggregation) specifically includes: Linear mapping is performed on the noise image, the background image and the foreground image respectively to obtain first feature data of the noise image, second feature data of the background image and third feature data of the foreground image.
[0039] The first attention output is determined according to the first feature data and the second feature data, and the calculation formula includes: (7) in, represents the first attention output, represent the query, key and value of the first feature data respectively, They represent the key and value of the second feature data respectively, and D represents the latitude of the key.
[0040] The second attention output is determined according to the first feature data and the third feature data, and the calculation formula includes: (8) in, represents the second attention output, They represent the key and value of the third feature data respectively, and D represents the latitude of the key.
[0041] Aggregate based on the first attention output and the second attention output, the calculation formula includes: (9) in, Represents the result of feature aggregation.
[0042] The multi-head attention sub-network aims to explicitly incorporate the camouflage attributes and the features of the foreground image into the camouflage image. In the multi-head attention sub-network, the hidden state of the noise image in the t-th round of denoising in the diffusion model is linearly mapped to obtain the first feature data, including: query ,key Sum , thereby capturing the global context and long-distance dependency in the hidden state of the noisy image. At the same time, in order to integrate the background image and the foreground image, in this embodiment, the CLIP encoder is first used to encode the background image and the foreground image respectively, and then the two are linearly projected to obtain the key of the second feature data Sum and the key of the third characteristic data Sum By treating these keys and values separately, we effectively separate their effects on the state variables, ensuring that the attention map obtained from the background image is different from the attention map based on the foreground image input.
[0043] Specifically, the image dataset also includes text data. The attention network also includes a mask-restricted attention subnetwork, which is used to perform feature aggregation on noisy images and text data, see Figure 4 , the Mask-restricted Attention Modulation includes: Linear mapping is performed on the text data to obtain fourth feature data of the text data.
[0044] The first cross-attention feature is determined according to the first feature data and the fourth feature data, and the calculation formula includes: (10) Among them, A represents the first cross-attention feature, It is the key of the fourth feature data.
[0045] The first cross-attention feature is adjusted by the foreground mask to obtain the second cross-attention feature. The calculation formula includes: (11) in, represents the attention weight at the object-specific token, Represents the foreground mask At the location The mask value at .
[0046] Aggregate the second attention feature and the fourth attention feature to obtain the third attention output. The calculation formula includes: (12) in, represents the third attention output, Indicates the value of the fourth feature data.
[0047] In this embodiment, the basic diffusion network does not fine-tune the model parameters. , and no low-rank parameters are introduced , but through the control network ( Figure 2 The red module in the figure is the Optimal camouflage Controller) and the attention network ( Figure 2 The blue module in the figure, Controllable Attention Manipulation, is used to modulate the drift field of the denoising process, so that the diffusion model can support various user inputs, such as Figure 2 Background image shown , foreground image , text data And the foreground mask Etc., thereby achieving camouflage image generation without modifying core parameters.
[0048] Also, see Figure 5 This embodiment also tests the diffusion model in this embodiment through a dedicated test image data set. The test image data set specifically includes 300 samples, each of which is a background image (background), a foreground image (subject), and may also include a foreground mask (Mask). Among them, the background image covers a variety of natural scenes, such as grasslands, forests, sand dunes, oceans, etc., and the foreground image includes humans, animals, plants, etc. The complete diffusion model is input and the camouflage image generated by this model is output.
[0049] The results show that ControlNet (the second example), Stable Cascade (the third example), and IP-Adapter (the fourth example) perform poorly in effectively disguising objects within the background, indicating that existing camouflaged image generation methods lack the ability to produce effective camouflage. InstantStyle (the fifth example) benefits from style consistency modulation, which helps to visually align objects and backgrounds, but is still limited in practical applications by object hiding and subject feature preservation. In addition, their synthetic images often show obvious artifacts at the embedding boundaries (see the third and fourth examples), and it is difficult to accurately match objects with smaller masks. Although traditional camouflaged image generation methods are able to blend objects into the background, they cannot preserve the attributes or identity of the objects, which contradicts the principles of nature. In nature, camouflaged objects should blend seamlessly into the background while retaining their inherent features, rather than becoming completely invisible. Overall, these methods are hampered by several limitations, including insufficient camouflage effectiveness, lack of authenticity, and poor controllability.
[0050] In contrast, the method for generating a camouflaged image provided in this embodiment effectively camouflages an object by generating a realistic appearance and texture that blends seamlessly with the background (Example 6). The method of the present invention embeds the foreground image into the specified area while retaining the main features of the foreground image by controlling the network during the denoising process of the diffusion model, thereby ensuring precise control and high fidelity throughout the camouflage process.
[0051] An exemplary embodiment of the present invention provides a device for generating a camouflage image. Figure 6 , the device comprises: An acquisition module 61 is used to acquire an image data set, where the image data set includes a background image and a foreground image; A generating module 62, for inputting the image data set into a pre-built diffusion model to generate a camouflaged image; Among them, the diffusion model includes the basic diffusion network, the attention network and the control network; The basic diffusion network is used to perform multiple rounds of denoising on the noisy image generated based on the background image; the attention network is used to aggregate the features of the noisy image, background image and foreground image; and the control network is used to control the denoising process of the basic diffusion network.
[0052] In one embodiment, the basic diffusion network includes: ; ; Where d is the differential symbol, represents the stochastic differential equation, represents the noisy image after the tth round of denoising, and represents the standard Brownian motion process, and represent the drift function and the fluctuation function respectively. represents the scoring function, represents the instantaneous loss, and represents the hyperparameter, represents the control network, represents the first terminal loss, represents the second terminal loss, To control the control constraints, Indicates that the control network satisfies the control constraints Constraints, represents mathematical expectation.
[0053] In one embodiment, the first terminal loss includes: ; in, Represents the background image, Represents the Tweedie formula Approximate camouflaged image represents the best camouflage loss indicator, To control constraints, Indicates that the control network satisfies the control constraints constraints.
[0054] In one embodiment, the image dataset further includes a foreground mask, and the second terminal loss is constructed based on the foreground mask.
[0055] In one embodiment, the second terminal loss includes: ; in, Represents the Tweedie formula The camouflaged image obtained by approximation, Represents the anti-detection descriptor, Represents the foreground mask.
[0056] In one embodiment, the attention network specifically includes: Linear mapping is performed on the noise image, the background image and the foreground image respectively to obtain first feature data of the noise image, second feature data of the background image and third feature data of the foreground image; The first attention output is determined according to the first feature data and the second feature data, and the calculation formula includes: ; in, represents the first attention output, represent the query, key and value of the first feature data respectively, denote the key and value of the second feature data respectively, and D denotes the latitude of the key; The second attention output is determined according to the first feature data and the third feature data, and the calculation formula includes: ; in, represents the second attention output, denote the key and value of the third characteristic data respectively, and D denotes the latitude of the key; Aggregate based on the first attention output and the second attention output, the calculation formula includes: ; in, Represents the result of feature aggregation.
[0057] In one embodiment, the image dataset also includes text data; and the attention network is further used to perform feature aggregation on the noisy image and text data.
[0058] The system and method embodiments in the embodiments of the present invention are based on the same inventive concept.
[0059] An exemplary embodiment of the present invention further provides an electronic device, see Figure 7 The electronic device 700 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 710 (the processor 710 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic battery FPGA), a memory 730 for storing data, and one or more storage media 720 (such as one or more mass storage devices) for storing application programs 723 or data 722. Among them, the memory 730 and the storage medium 720 can be short-term storage or permanent storage. The storage medium 720 stores at least one message, at least one program, code set or message set, and the at least one message, at least one program, code set or message set is loaded and executed by the processor to implement the above method.
[0060] Furthermore, the CPU 710 may be configured to communicate with the storage medium 720 and execute a series of message operations in the storage medium 720 on the electronic device 700. The electronic device 700 may also include one or more power supplies 760, one or more wired and wireless network interfaces 750, one or more input and output interfaces 740, and / or one or more operating systems 721, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0061] The input / output interface 740 may be used to receive or send data via a network. A specific example of the network may include a wireless network provided by a communication provider of the electronic device 700. In one example, the input / output interface 740 includes a network adapter (Network Interface Controller, NIC), which may be connected to other network devices via a base station so as to communicate with the Internet. In one example, the input / output interface 740 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0062] It can be understood by those skilled in the art that Figure 7 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 7 More or fewer components as shown, or with Figure 7 Different configurations shown.
[0063] An embodiment of the present invention also provides a computer-readable storage medium, which can be set in a server to store at least one message, at least one program, a code set or a message set related to a method in an implementation method embodiment. The at least one message, the at least one program, the code set or the message set is loaded and executed by the processor to implement the above method.
[0064] Optionally, in this embodiment, the storage medium may be located in a network server among multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0065] It is known from common technical knowledge that the present invention can be implemented by other embodiments that do not deviate from its spirit or essential features. Therefore, the embodiments of the above invention are only illustrative in all respects and are not exclusive. All changes within the scope of the present invention or within the scope equivalent to the present invention are included in the present invention.
[0066] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0067] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program messages. These computer program messages can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the messages executed by the processor of the computer or other programmable data processing device generate messages for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0068] These computer program messages can also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to work in a specific manner, so that the messages stored in the computer readable memory produce an article of manufacture including a message device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0069] These computer program messages can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the messages executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for generating a camouflaged image, characterized in that: The method comprises: Acquire an image data set, wherein the image data set includes a background image and a foreground image; Inputting the image data set into a pre-built diffusion model to generate a camouflaged image; Wherein, the diffusion model includes a basic diffusion network, an attention network and a control network; The basic diffusion network is used to perform multiple rounds of denoising on the noise image generated based on the background image; the attention network is used to perform feature aggregation on the noise image, the background image and the foreground image; and the control network is used to control the denoising process of the basic diffusion network.
2. The method for generating a disguised image according to claim 1, wherein: The basic diffusion network includes: ; ; Where d is the differential symbol, represents the stochastic differential equation, represents the noisy image after the tth round of denoising, and represents the standard Brownian motion process, and denote the drift function and the fluctuation function respectively, represents the scoring function, represents the instantaneous loss, and represents the hyperparameter, represents the control network, represents the first terminal loss, represents the second terminal loss, To control constraints, Indicates that the control network satisfies the control constraints Constraints, represents mathematical expectation.
3. The method for generating a disguised image according to claim 2, wherein: The first terminal loss includes: ; in, Represents the Tweedie formula The camouflaged image obtained by approximation, Represents the background image, represents the best camouflage loss indicator, is the control set, Indicates that the control network satisfies the control constraints constraints.
4. The method for generating a disguised image according to claim 2, wherein: The image dataset also includes a foreground mask, and the second terminal loss is constructed based on the foreground mask.
5. The method for generating a disguised image according to claim 4, wherein: The second terminal loss includes: ; in, Represents the Tweedie formula The camouflaged image obtained by approximation, Represents the anti-detection descriptor, Represents the foreground mask.
6. The method for generating a disguised image according to claim 1, wherein: The attention network specifically includes: Performing linear mapping on the noise image, the background image and the foreground image respectively to obtain first feature data of the noise image, second feature data of the background image and third feature data of the foreground image; The first attention output is determined according to the first feature data and the second feature data, and the calculation formula includes: ; in, represents the first attention output, represent the query, key and value of the first feature data respectively, represent the key and value of the second feature data respectively, and D represents the latitude of the key; The second attention output is determined according to the first feature data and the third feature data, and the calculation formula includes: ; in, represents the second attention output, represent the key and value of the third feature data respectively, and D represents the latitude of the key; Aggregation is performed based on the first attention output and the second attention output, and the calculation formula includes: ; in, Represents the result of feature aggregation.
7. The method for generating a disguised image according to claim 1 or 6, wherein: The image data set also includes text data; the attention network is also used to perform feature aggregation on the noisy image and the text data.
8. A device for generating a disguised image, characterized in that: The device comprises: An acquisition module, used to acquire an image data set, wherein the image data set includes a background image and a foreground image; A generating module, used for inputting the image data set into a pre-built diffusion model to generate a camouflaged image; Wherein, the diffusion model includes a basic diffusion network, an attention network and a control network; The basic diffusion network is used to perform multiple rounds of denoising on the noise image generated based on the background image; the attention network is used to perform feature aggregation on the noise image, the background image and the foreground image; and the control network is used to control the denoising process of the basic diffusion network.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, a code set or an instruction set is loaded and executed by the processor to implement the method for generating a camouflaged image according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the method for generating a camouflaged image according to any one of claims 1 to 7.
Citation Information
Patent Citations
Limited angle CT (Computed Tomography) reconstruction method, system and equipment based on wavelet multi-channel scoring
CN117671054A
Speech synthesis method and device, electronic equipment and readable storage medium
CN117854470A
Unmanned aerial vehicle image target detection method and system based on remote sensing basic model auxiliary framework
CN118505980A
Mask-guided camouflage target image generation method
CN118570628A
Ultra-short-term wind power probability prediction method based on nonparametric stochastic differential equation
CN119109037A
Cited By
Method and device for denoising normal diagram of building pitched roof
CN121147059A
A method and device for denoising a normal map of a building hip roof
CN121147059B