A method, apparatus, device, and medium for generating a camouflage image

Through the combination of the basic diffusion network and the attention network, the control of the network is controlled to generate high-precision and high-fidelity camouflage images, solving the problem of high generation cost in the existing technology.

CN119941909BActive Publication Date: 2025-07-11HUNAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510414598.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

Existing methods of camouflage image generation rely on a large amount of training data, resulting in excessive generation cost.

Method used

The basic diffusion network is used to perform multiple rounds of denoising on the noise image, and the attention network is combined with the feature aggregation of the noise, background and foreground images, and control the denoising process through the control network to generate a camouflage image.

Benefits of technology

It realizes high-precision, high-fidelity camouflage image generation without a large number of training samples, reducing the generation cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941909B_ABST
    Figure CN119941909B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, device and medium for generating a camouflage image. The method includes: acquiring an image data set, which includes a background image and a foreground image; inputting the image data set into a pre-constructed diffusion model to generate a camouflage image; wherein the diffusion model includes a basic diffusion network, an attention network and a control network; the basic diffusion network is used for performing multiple rounds of denoising on a noise image generated based on the background image; the attention network is used for feature aggregation of the noise image, the background image and the foreground image; the control network is used for controlling the denoising process of the basic diffusion network. A large number of training samples are required to train the model, reducing the generation cost of the camouflage image, and at the same time, a high-precision and high-fidelity camouflage image can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of algorithms, and in particular, to a method, apparatus, device, and medium for generating camouflage images. Background Art

[0002] Biological camouflage in nature is an important mechanism for the survival of organisms, which has inspired the wide application of camouflage technology in the fields of military, art, and security, and has also stimulated the development of anti-camouflage visual perception. However, the scarcity of camouflage images greatly limits the effectiveness of anti-camouflage strategies, which has inspired people's interest in generating camouflage images.

[0003] However, the current methods for generating camouflage images rely on a large amount of training data to train the model, resulting in too high a cost for generating camouflage images. Summary of the Invention

[0004] In view of the problem of too high a cost for generating camouflage images in the prior art, the present invention provides a method, apparatus, device, and medium for generating camouflage images.

[0005] In a first aspect, an embodiment of the present invention provides a method for generating a camouflage image, the method including:

[0006] Obtaining an image data set, the image data set including a background image and a foreground image;

[0007] Inputting the image data set into a pre-constructed diffusion model to generate a camouflage image;

[0008] wherein the diffusion model includes a basic diffusion network, an attention network, and a control network;

[0009] The basic diffusion network is used for performing multiple rounds of denoising on a noise image generated based on the background image; the attention network is used for aggregating features of the noise image, the background image, and the foreground image; and the control network is used for controlling the denoising process of the basic diffusion network.

[0010] In one embodiment, the basic diffusion network includes:

[0011] ;

[0012] ;

[0013] where d is the differential symbol, represents a stochastic differential equation, represents the noise image after the t-th round of denoising, and represent the standard Brownian motion process, and represent the drift function and the volatility function respectively, represents a scoring function, represents the instantaneous loss, and represents a hyperparameter, represents a control network, represents the first terminal loss, represents the second terminal loss, is a camouflage constraint, represents that the control network satisfies the camouflage constraint constraint, represents the mathematical expectation.

[0014] In one embodiment, the first terminal loss includes:

[0015] ;

[0016] wherein, represents the camouflage image approximately estimated based on the Tweedie formula for , represents the background image, represents the best camouflage loss identifier, is a camouflage constraint, represents that the control network satisfies the camouflage constraint constraint.

[0017] In one embodiment, the image dataset further includes a foreground mask, and the second terminal loss is constructed based on the foreground mask.

[0018] In one embodiment, the second terminal loss includes:

[0019] ;

[0020] wherein, represents the camouflage image approximately estimated based on the Tweedie formula for , represents the anti-detection descriptor, represents the foreground mask.

[0021] In one embodiment, the attention network specifically includes:

[0022] Perform linear mapping on the noise image, background image, and foreground image respectively to obtain the first feature data of the noise image, the second feature data of the background image, and the third feature data of the foreground image;

[0023] Determine the first attention output according to the first feature data and the second feature data, and the calculation formula includes:

[0024] ;

[0025] wherein, represents the first attention output, respectively represent the query, key, and value of the first feature data, respectively represent the key and value of the second feature data, and D represents the dimension of the key;

[0026] Determine the second attention output according to the first feature data and the third feature data, and the calculation formula includes:

[0027] ;

[0028] wherein, represents the second attention output, respectively represent the key and value of the third feature data, and D represents the dimension of the key;

[0029] Aggregate according to the first attention output and the first attention output, and the calculation formula includes:

[0030]

[0031] wherein, represents the result of feature aggregation.

[0032] In one embodiment, the image data set further includes text data; the attention network is further used to perform feature aggregation on the noisy image and the text data.

[0033] In a second aspect, an embodiment of the present invention provides a device for generating a camouflage image, and the device includes:

[0034] An acquisition module, configured to acquire an image data set, where the image data set includes a background image and a foreground image;

[0035] A generation module, configured to input the image data set into a pre-constructed diffusion model to generate a camouflage image;

[0036] wherein, the diffusion model includes a basic diffusion network, an attention network, and a control network;

[0037] The basic diffusion network is used to perform multi-round denoising on the noisy image generated based on the background image; the attention network is used to perform feature aggregation on the noisy image, the background image, and the foreground image; the control network is used to control the denoising process of the basic diffusion network.

[0038] In a third aspect, an embodiment of the present invention provides an electronic device, which includes a processor and a memory. At least one message, at least one program, a code set, or a message set is stored in the memory, and at least one message, at least one program, a code set, or a message set is loaded and executed by the processor to implement a method for generating a camouflage image according to any one of the first aspects.

[0039] Fourthly, an embodiment of the present invention provides a computer-readable storage medium, in which at least one message or at least one segment of program is stored, and the at least one message or at least one segment of program is loaded and executed by a processor to implement a method for generating a camouflage image according to any one of the first aspect.

[0040] The method, device, equipment and medium for generating a camouflage image provided by the present invention have the following technical effects:

[0041] The noise image constructed is denoised by the basic diffusion network, and during the denoising process, the features of the foreground image and the background image are aggregated into the camouflage image through the attention network. At the same time, a control network is constructed to optimize the denoising process of the basic diffusion network, realizing the generation of a camouflage image without training. There is no need to use a large number of training samples to train the model, reducing the generation cost of the camouflage image, and at the same time, a high-precision and high-fidelity camouflage image can be obtained.

[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The attached drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0044] Figure 1 is a schematic flowchart of a method for generating a camouflage image provided by an exemplary embodiment of the present invention;

[0045] Figure 2 is a schematic diagram of a diffusion model provided by an exemplary embodiment of the present invention;

[0046] Figure 3 is a schematic diagram of a multi-head attention sub-network provided by an exemplary embodiment of the present invention;

[0047] Figure 4 is a schematic diagram of a mask restriction sub-network provided by an exemplary embodiment of the present invention;

[0048] Figure 5 is a schematic diagram for comparing test results provided by an exemplary embodiment of the present invention;

[0049] Figure 6 is a schematic diagram of modules of a device for generating a camouflage image provided by an exemplary embodiment of the present invention;

[0050] Figure 7 is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0052] The following detailed descriptions are all exemplary descriptions, aiming to provide further details of the present invention. Unless otherwise specified, all technical terms adopted in the present invention have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention pertains. The terms used in the present invention are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0053] An exemplary embodiment of the present invention provides a method for generating a camouflage image. Refer to Figure 1 , the method includes:

[0054] Step S101, obtaining an image data set.

[0055] Among them, the image data set includes a background image and a foreground image.

[0056] Step S102, inputting the image data set into a pre-constructed diffusion model to generate a camouflage image.

[0057] Among them, the diffusion model includes a basic diffusion network, an attention network, and a control network.

[0058] Specifically, the basic diffusion network is used to perform multiple rounds of denoising on a noise image generated based on the background image, the attention network is used to perform feature aggregation on the noise image, the background image, and the foreground image, and the control network is used to control the denoising process of the basic diffusion network.

[0059] In this embodiment, the basic diffusion network is used to denoise the constructed noise image. During the denoising process, the attention network aggregates the features of the foreground image and the background image into the camouflage image, and at the same time, the control network controls the denoising process of the basic diffusion network, realizing the generation of a camouflage image without training, without the need for a large number of training samples to train the model, reducing the generation cost of the camouflage image, and at the same time being able to obtain a high-precision and high-fidelity camouflage image.

[0060] In one embodiment, refer to Figure 2 , an explanation of the basic diffusion network of the diffusion model is given:

[0061] Before the basic diffusion network performs denoising, the basic diffusion network is also used to generate a noise image based on the background image. The generation process of the noise image uses the forward stochastic process of the basic diffusion network. The forward stochastic process can be understood as a process of cumulatively adding random Gaussian noise to the background image T times, which can be expressed as the solution of a stochastic differential equation and is represented by the following formula:

[0062] (1)

[0063] Among them, and are respectively called the drift function and the fluctuation function.

[0064] The denoising process of the noisy image utilizes the reverse stochastic process of the basic diffusion network. The reverse stochastic process can be understood as the solution of the reverse stochastic differential equation from the time step of T~0 under slight regularity, and is expressed by the formula as follows:

[0065] (2)

[0066] Among them, and represent the standard Brownian motion process, and are respectively called the drift function and the fluctuation function. represents the scoring function. The scoring function is approximately obtained through the U-Net neural network of the diffusion model in the image generation task for approximation.

[0067] Based on the above two stochastic processes, the basic diffusion network in the embodiments of the present invention specifically includes:

[0068] (3)

[0069] (4)

[0070] Among them, d is the differential symbol, represents the stochastic differential equation, represents the noisy image after denoising in the t-th round, is the generated image after denoising is completed, and represent the standard Brownian motion process, and are respectively called the drift function and the fluctuation function. represents the scoring function, represents the instantaneous loss, and represent hyperparameters, represents the control network, represents the first terminal loss, represents the second terminal loss, is the camouflage constraint, represents that the control network satisfies the constraint of the camouflage constraint of, denotes the mathematical expectation, and the camouflage constraint is set according to the actual diffusion scenario, which is not specifically limited here.

[0071] In the basic diffusion network, the denoising process of the basic diffusion network is controlled by introducing a control network. The goal of the control network is to make the area of the foreground image in the generated camouflage image visually blend with the background in the image, and the synthesized camouflage image can effectively evade the detection of the camouflage detector. In this embodiment, these two goals are defined as the expected terminal conditions of the diffusion model, enabling the control network to modify the time-varying drift field in the basic diffusion network during the denoising process, thereby effectively guiding the denoising process to achieve these two expected terminal conditions.

[0072] The above two goals respectively correspond to representing the first terminal loss and the second terminal loss, which are determined by the control network.

[0073] In a specific embodiment, since the scoring function can be approximately obtained through the U-Net neural network of the diffusion model in the image generation task where represents the diffusion step, represents the network parameters, so formula (4) can be understood as mainly minimizing the total loss by the control network by optimizing the drift of the reverse stochastic process. Further, the total loss regarding the control network can be defined as from to the integral

[0074] :

[0075] At this time, minimizing the total loss is transformed into solving the minimum value of i.e.:

[0076] ;

[0077] This goal can be obtained through the Hamilton-Jacobi-Bellman equation, which is used to find the optimal control strategy in a continuous-time system to minimize a certain performance index or cost function. The Hamilton-Jacobi-Bellman equation states that if the minimum value has continuous partial derivatives, it must satisfy the following partial differential equation:

[0078] ;

[0079] Therefore, when When →∞, the optimal controller can be obtained by solving the Hamilton-Jacobi-Bellman equation, which allows us to disregard the transient cost. Therefore, the present invention only needs to focus on optimizing the two terminal losses to achieve formula (3).

[0080] Specifically, the first terminal loss includes:

[0081] (5)

[0082] Among them, represents the camouflage image approximately estimated based on the Tweedie formula for , represents the background image, represents the best camouflage loss identifier, is the camouflage constraint, represents the constraint that the control network satisfies the camouflage constraint . The camouflage constraint is set according to the actual diffusion scenario and is not specifically limited here.

[0083] Approximately estimate the camouflage image based on the Tweedie formula, and use a consistent style extractor as the best camouflage loss identifier to minimize the visual difference between the background image and the camouflage image so that the area where the foreground image is located in the generated camouflage image visually blends with the background in the image.

[0084] Specifically, the image dataset also includes a foreground mask. The second terminal loss is constructed based on the foreground mask and includes:

[0085] (6)

[0086] Among them, represents the camouflage image approximately estimated based on the Tweedie formula for , represents the anti-detection descriptor, represents the foreground mask.

[0087] Input the camouflage image approximately estimated based on the Tweedie formula into the anti-detection descriptor , and initialize it with a pre-trained camouflage target detector to minimize the consistency between the anti-detection result and the foreground mask so that the synthesized camouflage image should be able to effectively evade the detection of the camouflage detector.

[0088] In one embodiment, refer to Figure 2-4, the attention network is described. The attention network may include at least one of a multi-head attention sub-network and a masked restricted attention sub-network.

[0089] Specifically, the multi-head attention sub-network is used to perform feature aggregation on the noise image, background image, and foreground image. Refer to Figure 3 , specifically, the multi-head attention sub-network (Multi-head Attention Aggregation) includes:

[0090] Perform linear mapping on the noise image, background image, and foreground image respectively to obtain the first feature data of the noise image, the second feature data of the background image, and the third feature data of the foreground image.

[0091] Determine the first attention output based on the first feature data and the second feature data. The calculation formula includes:

[0092] (7)

[0093] Where, represents the first attention output, respectively represent the query, key, and value of the first feature data, respectively represent the key and value of the second feature data, and D represents the dimension of the key.

[0094] Determine the second attention output based on the first feature data and the third feature data. The calculation formula includes:

[0095] (8)

[0096] Where, represents the second attention output, respectively represent the key and value of the third feature data, and D represents the dimension of the key.

[0097] Aggregate based on the first attention output and the first attention output. The calculation formula includes:

[0098] (9)

[0099] Where, represents the result of feature aggregation.

[0100] The multi-head attention sub-network aims to explicitly incorporate the features of the camouflage attribute and the foreground image into the camouflage image. In the multi-head attention sub-network, the hidden layer state of the noise image in the t-th denoising process of the diffusion model is linearly mapped to obtain the first feature data, including: query , key and value , so as to capture the global context and long-distance dependencies in the hidden layer state of the noisy image. At the same time, in order to integrate the background image and the foreground image, in this embodiment, the CLIP encoder is first used to encode the background image and the foreground image respectively, and then the two are respectively linearly projected to obtain the keys and values of the second feature data and the keys and values of the third feature data. By processing these keys and values separately, we effectively separate their influence on the state variables, ensuring that the attention map obtained from the background image is different from the attention map based on the foreground image input.

[0101] Specifically, the image dataset also includes text data. The attention network further includes a mask-restricted attention sub-network, which is used to perform feature aggregation on the noisy image and the text data. See Figure 4 , and the mask-restricted attention sub-network (Mask-restricted Attention Modulation) includes:

[0102] Perform a linear mapping on the text data to obtain the fourth feature data of the text data.

[0103] Determine the first cross-attention feature according to the first feature data and the fourth feature data, and the calculation formula includes:

[0104] (10)

[0105] Where, A represents the first cross-attention feature, is the key of the fourth feature data.

[0106] Adjust the first cross-attention feature through the foreground mask to obtain the second cross-attention feature, and the calculation formula includes:

[0107] (11)

[0108] Where, represents the attention weight at the object-specific token, represents the foreground mask at the position of is the mask value.

[0109] Aggregate the second attention feature and the fourth attention feature to obtain the third attention output, and the calculation formula includes:

[0110] (12)

[0111] Where, represents the third attention output, Represents the value of the fourth feature data.

[0112] In this embodiment, the basic diffusion network does not fine-tune the model parameters , nor does it introduce low-rank parameters , but instead modulates the drift field of the denoising process through the control network ( Figure 2 the red module in, Optimal camouflage Controller) and the attention network ( Figure 2 the blue module in, Controllable Attention Manipulation) to enable the diffusion model to support various user inputs, such as Figure 2 the background image shown , the foreground image , the text data and the foreground mask etc., so as to achieve camouflage image generation without modifying the core parameters.

[0113] In addition, referring to Figure 5 , this embodiment also tests the diffusion model in this embodiment through a dedicated test image dataset. The test image dataset specifically includes 300 samples, each sample being a background image (background), a foreground image (subject), and may also include a foreground mask (Mask). Among them, the background images cover a variety of natural scenes, such as grasslands, forests, sand dunes, oceans, etc., while the foreground images include humans, animals, plants, etc. Input the complete diffusion model and output the camouflage images generated by this model.

[0114] The results show that the controlnet network (the second example), Stable Cascade (the third example), and IP - Adapter (the fourth example) perform poorly in effectively camouflaging objects within the background, indicating that existing camouflage image generation methods lack the ability to produce effective camouflage. InstantStyle (the fifth example) benefits from style consistency modulation, which helps to visually align the object and the background, but is still limited in object hiding and subject feature retention in practical applications. In addition, their synthesized images usually show obvious artifacts at the embedding boundaries (see the third and fourth examples), and it is difficult to precisely match the object with a small mask. Traditional camouflage image generation methods can integrate the object into the background, but cannot retain the attributes or identity of the object, which contradicts the natural principle. In nature, the camouflaged object should seamlessly blend into the background while retaining its inherent features, rather than becoming completely invisible. Overall, these methods are hindered by some limitations, including insufficient camouflage effectiveness, lack of authenticity, and poor controllability.

[0115] In contrast, the method for generating a camouflage image provided in this embodiment effectively camouflages an object by generating a realistic appearance and texture that seamlessly blends with the background (the sixth example). The method of the present invention embeds the foreground image into a specified area while retaining the main features of the foreground image by adding a control network during the denoising process of the diffusion model, ensuring precise control and high fidelity throughout the camouflage process.

[0116] An exemplary embodiment of the present invention provides a device for generating a camouflage image. Refer to Figure 6 , the device includes:

[0117] An acquisition module 61, configured to acquire an image data set, where the image data set includes a background image and a foreground image;

[0118] A generation module 62, configured to input the image data set into a pre-constructed diffusion model to generate a camouflage image;

[0119] Among them, the diffusion model includes a basic diffusion network, an attention network, and a control network;

[0120] The basic diffusion network is used to perform multiple rounds of denoising on a noise image generated based on the background image; the attention network is used to perform feature aggregation on the noise image, the background image, and the foreground image; the control network is used to control the denoising process of the basic diffusion network.

[0121] In one embodiment, the basic diffusion network includes:

[0122] ;

[0123] ;

[0124] Among them, d is the differential symbol, represents the stochastic differential equation, represents the noise image after the t-th round of denoising, and represent the standard Brownian motion process, and represent the drift function and the volatility function respectively. represents the score function, represents the instantaneous loss, and represent hyperparameters, represents the control network, represents the first terminal loss, represents the second terminal loss, is the control constraint, represents that the control network satisfies the control constraint of the constraint, represents the mathematical expectation.

[0125] In one embodiment, the first terminal loss includes:

[0126] ;

[0127] Wherein, represents the background image, represents the pseudo-image approximately estimated based on the Tweedie formula for the approximation, represents the best pseudo-loss identifier, is the control constraint, represents that the control network satisfies the control constraint constraint.

[0128] In one embodiment, the image dataset further includes a foreground mask, and the second terminal loss is constructed based on the foreground mask.

[0129] In one embodiment, the second terminal loss includes:

[0130] ;

[0131] Wherein, represents the pseudo-image approximately estimated based on the Tweedie formula for the approximation, represents the anti-detection descriptor, represents the foreground mask.

[0132] In one embodiment, the attention network specifically includes:

[0133] Perform linear mapping on the noise image, background image, and foreground image respectively to obtain the first feature data of the noise image, the second feature data of the background image, and the third feature data of the foreground image;

[0134] Determine the first attention output according to the first feature data and the second feature data, and the calculation formula includes:

[0135] ;

[0136] Wherein, represents the first attention output, respectively represent the query, key, and value of the first feature data, respectively represent the key and value of the second feature data, and D represents the latitude of the key;

[0137] Determine the second attention output according to the first feature data and the third feature data, and the calculation formula includes:

[0138] ;

[0139] Wherein, represents the second attention output, respectively represent the key and value of the third feature data, and D represents the dimension of the key;

[0140] Aggregate according to the first attention output and the first attention output, and the calculation formula includes:

[0141] ;

[0142] Among them, represents the result of feature aggregation.

[0143] In one embodiment, the image data set further includes text data; the attention network is also used to perform feature aggregation on the noise image and the text data.

[0144] The system and method embodiments in the embodiments of the present invention are based on the same inventive concept.

[0145] An exemplary embodiment of the present invention further provides an electronic device. Refer to Figure 7 , the electronic device 700 may vary greatly due to configuration or performance differences, and may include one or more central processing units (Central Processing Units, CPUs) 710 (the processor 710 may include, but is not limited to, a microprocessor MCU or a programmable logic battery FPGA, etc.), a memory 730 for storing data, and one or more storage media 720 for storing application programs 723 or data 722 (for example, one or more mass storage devices). Among them, the memory 730 and the storage media 720 may be transient storage or persistent storage. At least one message, at least one program, a code set or a message set is stored in the storage media 720, and at least one message, at least one program, a code set or a message set is loaded and executed by the processor to implement the above method.

[0146] Furthermore, the central processing unit 710 may be configured to communicate with the storage media 720 and perform a series of message operations in the storage media 720 on the electronic device 700. The electronic device 700 may further include one or more power supplies 760, one or more wired and wireless network interfaces 750, one or more input / output interfaces 740, and / or, one or more operating systems 721, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0147] The input / output interface 740 can be used to receive or transmit data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the electronic device 700. In one example, the input / output interface 740 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the input / output interface 740 can be a RadioFrequency (RF) module, which is used to communicate with the Internet wirelessly.

[0148] Those of ordinary skill in the art can understand that Figure 7 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the electronic device 700 may further include more or fewer components than Figure 7 shown, or have a different configuration from Figure 7 that shown.

[0149] An embodiment of the present invention also provides a computer-readable storage medium. The storage medium can be disposed in the server to store at least one message, at least one program, a code set or a message set related to a method in the method embodiment. The at least one message, the at least one program, the code set or the message set are loaded and executed by the processor to implement the above method.

[0150] Optionally, in this embodiment, the above storage medium may be a network server among multiple network servers of a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs.

[0151] As is known by common technical knowledge, the present invention can be implemented by other embodiments that do not depart from its spiritual essence or essential features. Therefore, the above embodiments of the present invention are illustrative in all aspects and are not the only ones. All changes within the scope of the present invention or within the scope equivalent to the present invention are encompassed by the present invention.

[0152] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0153] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that it is still possible to modify the specific implementation manners of the present invention or make equivalent substitutions. Any modification or equivalent substitution that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A method for generating a camouflage image, characterized in that, The method includes: Obtaining an image data set, where the image data set includes a background image and a foreground image; Inputting the image data set into a pre-constructed diffusion model to generate a camouflage image; Wherein, the diffusion model includes a basic diffusion network, an attention network, and a control network; The basic diffusion network is used for performing multiple rounds of denoising on a noise image generated based on the background image; the attention network is used for feature aggregation of the noise image, the background image, and the foreground image; the control network is used for controlling the denoising process of the basic diffusion network; The basic diffusion network includes: ; ; Among them, is the differential symbol, represents a stochastic differential equation, represents the noise image after the -th round of denoising, and represents a standard Brownian motion process, and represent the drift function and the volatility function respectively, represents a scoring function, represents the instantaneous loss, and represent hyperparameters, represents a control network, represents the first terminal loss, represents the second terminal loss, is the control constraint, represents that the control network satisfies the control constraint constraint, represents the mathematical expectation.

2. The method for generating a camouflage image according to claim 1, characterized in that, The first terminal loss includes: ; Among them, represents the camouflage image approximately estimated based on the Tweedie formula for the approximation, represents the background image, represents the best camouflage loss identifier, is the control set, represents that the control network satisfies the control constraint of the constraint.

3. The method for generating a camouflage image according to claim 1, wherein, The image data set further includes a foreground mask, and the second terminal loss is constructed based on the foreground mask.

4. The method for generating a camouflage image according to claim 3, characterized in that, The second terminal loss includes: ; Among them, represents the camouflage image approximately estimated based on the Tweedie formula for the approximate estimation, represents the anti-detection descriptor, and represents the foreground mask.

5. The method for generating a camouflage image according to claim 1, wherein The attention network specifically includes: Performing linear mapping on the noise image, the background image, and the foreground image respectively to obtain first feature data of the noise image, second feature data of the background image, and third feature data of the foreground image; Determining a first attention output according to the first feature data and the second feature data, and the calculation formula includes: ; Among them, represents the first attention output, respectively represent the query, key, and value of the first feature data, respectively represent the key and value of the second feature data, represents the dimension of the key; Determining a second attention output according to the first feature data and the third feature data, and the calculation formula includes: ; Among them, represents the second attention output, respectively represent the key and value of the third feature data, represents the dimension of the key; Performing aggregation according to the first attention output and the second attention output, and the calculation formula includes: ; Among them, represents the result of feature aggregation.

6. The method for generating a camouflage image according to claim 1 or 5, characterized in that, The image data set further includes text data; the attention network is further used for feature aggregation of the noise image and the text data.

7. An apparatus for generating a camouflage image, characterized in that, The device includes: An acquisition module, configured to acquire an image data set, where the image data set includes a background image and a foreground image; A generation module, configured to input the image data set into a pre-constructed diffusion model to generate a camouflage image; Wherein, the diffusion model includes a basic diffusion network, an attention network, and a control network; The basic diffusion network is used for performing multiple rounds of denoising on a noise image generated based on the background image; the attention network is used for feature aggregation of the noise image, the background image, and the foreground image; the control network is used for controlling the denoising process of the basic diffusion network; The basic diffusion network includes: ; ; Among them, is the differential symbol, represents a stochastic differential equation, represents the noise image after the th round of denoising, and represent the standard Brownian motion process, and represent the drift function and the volatility function respectively, represents the scoring function, and represent hyperparameters, represents the control network, represents the first terminal loss, represents the second terminal loss, is the control constraint, represents that the control network satisfies the control constraint constraint, represents the mathematical expectation.

8. An electronic device, characterized in that, The electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and at least one instruction, at least one program, a code set or an instruction set is loaded and executed by the processor to implement the method for generating a camouflage image according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, At least one instruction or at least one program is stored in the computer-readable storage medium, and at least one instruction or at least one program is loaded and executed by the processor to implement the method for generating a camouflage image according to any one of claims 1-6.

Citation Information

Patent Citations

  • Regitive digital camouflage generation method, device and equipment

    CN119251043A