An infrared ship image controllable generation method for a marine scene

By constructing a controllable ship image generation framework with foreground-background separation, and adopting a denoising diffusion model architecture and a progressive generation strategy, the problem of generating small target ship images in infrared scenes was solved, and high-quality, controllable infrared ship image generation was achieved.

CN119809946BActive Publication Date: 2025-10-17AEROSPACE SCI & IND GRP INTELLIGENT TECH RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411748955.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-10-17
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

Existing technologies struggle to generate high-quality, controllable images of small target ships in infrared scenarios, especially at sea where the targets are small, their attributes are vague, and sea conditions are complex, resulting in poor generation outcomes.

Method used

A controllable ship image generation framework with foreground and background separation is constructed. A denoising diffusion model architecture is adopted. The sub-diffusion model is trained in three stages: foreground target generation, background target generation, and foreground and background target fusion. A progressive generation strategy and an autoencoder are introduced to achieve accurate binding and high-realism fusion of small-scale targets.

Benefits of technology

It generates clear images of target ships that showcase their unique characteristics, achieving highly realistic foreground and background fusion and position-aware fusion, thus solving the challenge of controllable generation of small-scale targets in infrared scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119809946B_ABST
    Figure CN119809946B_ABST
Patent Text Reader

Abstract

The application provides an infrared ship image controllable generation method for a marine scene, and the method comprises the following steps: preparing an infrared marine ship image dataset; constructing a front-background separation controllable ship image generation framework, the framework adopts a denoising diffusion model architecture, and comprises three stages of foreground target generation, background target generation and foreground-background target fusion; setting a model training constraint; training sub-diffusion models in the generation framework respectively; inputting a ship target to be generated, a position, a camera view angle and a sea condition into the generation framework, and outputting a model as a target ship generation result in the infrared marine scene meeting a control information requirement, thereby completing ship image generation of the marine infrared scene. By using the technical scheme of the application, the technical problem that the existing diffusion model control method cannot realize controllable generation under the conditions of small ship target scale, fuzzy ship attribute information, different sea conditions and other interference can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer vision, and particularly relates to an infrared ship image controllable generation method for a sea scene. BACKGROUND

[0002] In recent years, significant progress has been made in the field of visual synthesis. Especially since the rise of deep learning technology, generative adversarial networks (GAN) and diffusion models have been proposed, which have greatly promoted related research. These technologies make the generated results more diverse and more natural and realistic, realizing the transition from random generation to controllable generation of synthetic data. At present, significant progress has been made in the field in terms of synthesis quality, attribute controllability, cross-modal synthesis and driving (such as text-to-image translation), data accumulation, and the extension of pre-trained generative models to specific downstream tasks.

[0003] However, due to the small scale of sea target data in the infrared scene and the low contrast and details of infrared images, the image quality is low, which greatly hinders the development of visual synthesis in this field, and related research is relatively scarce. In addition, the unclear texture of small targets and complex sea conditions further increase the difficulty of generating specific types of ships. At present, the mainstream diffusion model, such as the stable diffusion model (SD), is mainly pre-trained in the visible light range, and direct fine-tuning based on these models may produce color artifacts in some cases and is difficult to meet the requirements in terms of target control.

[0004] Recently, some researches have begun to focus on how to control the diffusion model to achieve a balance between controllability and diversity. The DreamBooth method proposed by Nataniel Ruiz inputs a small number of sample images of specific objects or styles together with predefined class labels into the model, and uses the reconstruction loss to make the model learn how to generate images with these characteristics. During the fine-tuning process, DreamBooth also adjusts the weights of the model to retain the ability to generate diverse images while generating images with specific personalized features. Another approach is to use weight separation to learn specific attribute features while maintaining the diversity and prior knowledge of the generated images. ControlNet proposed by Lvmin Zhang is a representative of this approach. ControlNet first locks the original model (locked copy) so that it does not participate in training, and then copies a trainable network model (trainable copy) for training. By inputting control conditions (such as prompt embeddings or other image latent space feature information) into the network, passing through zero convolution and trainable copy modules in turn, and finally adding the results of the locked copy module to the final generated results. This method allows a wide range of control of the generated results based on image edge contours, segmentation maps, depth maps, pose maps, and other types of images.

[0005] Most of the research on background control focuses on target transfer, that is, transferring target objects to a specific background to achieve control of the background. Methods such as Paint-by-Example and ObjectStitch use the guidance of image-text encoders to embed reference objects into images, but only retain semantic similarity with the inserted objects. AnyDoor uses a self-supervised representation of the reference object and its high-frequency mapping as a condition to enhance object identity preservation, and supplements the identity features of the self-supervised representation with the detailed features of the high-frequency mapping. This design allows a variety of local changes (such as lighting, direction, pose, etc.) to be implemented, supporting effective fusion of objects with different environments, thereby improving the fidelity of AnyDoor generated images. However, in infrared scenes, the color information prior contained in the pre-trained weights of these models may interfere with the generation of infrared images. At the same time, some physical priors from visible light (such as shadows) may be misleading in infrared scene generation. In addition, there is still a lack of mature solutions for controllable generation of small targets and special backgrounds, and this challenge needs in-depth research. SUMMARY

[0006] The present invention aims to at least solve one of the technical problems existing in the prior art.

[0007] The present invention provides an infrared ship image controllable generation method for marine scenes, which comprises:

[0008] Prepare an infrared maritime vessel image dataset;

[0009] A front-background separation controllable vessel image generation framework is constructed, which adopts a denoising diffusion model architecture, including three stages of foreground target generation, background target generation and foreground-background target fusion;

[0010] Set model training constraints;

[0011] Respectively train the sub-diffusion models in the generation framework;

[0012] The vessel target to be generated, the position, the camera perspective and the sea conditions are input into the generation framework, and the model output is the target vessel generation result in the infrared maritime scene that meets the control information requirements, completing the vessel image generation of the maritime infrared scene.

[0013] The technical scheme of the present application provides an infrared vessel image controllable generation method for maritime scenes, which constructs a perfect controllable diffusion model generation framework with separately controllable foreground and background. After training, the framework can generate target vessels that clearly exhibit unique category features, and realize high-realism foreground-background fusion and position perception fusion. The method of the present application particularly emphasizes the accurate binding of small-scale targets in categories and attributes (such as direction information), and focuses on how to accurately simulate the modal characteristics in the infrared scene. Compared with the prior art, the technical scheme of the present application can solve the technical problems that the existing diffusion model control method cannot realize controllable generation of small-scale maritime vessel targets, blurred vessel attribute information and different sea conditions and other disturbances. BRIEF DESCRIPTION OF DRAWINGS

[0014] The included drawings are used to provide a further understanding of the embodiments of the present application, constitute a part of the specification, serve to illustrate the embodiments of the present application, and together with the text description, explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0015] Figure 1 A flowchart of the infrared vessel image controllable generation method for maritime scenes provided by the specific embodiments of the present application is shown;

[0016] Figure 2 A schematic diagram of the front-background separation controllable vessel image generation framework provided by the specific embodiments of the present application is shown;

[0017] Figure 3 A synthesis effect diagram of the vessel infrared image of the maritime scene provided by the specific embodiments of the present application is shown. DETAILED DESCRIPTION

[0018] It should be noted that the embodiments and features of the embodiments in the present application can be combined with each other in the case of no conflict. The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The description of the at least one exemplary embodiment is actually only illustrative, but not as any limitation on the present application and its application or use. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0019] It should be noted that the terms used herein are only intended to describe specific embodiments, and are not intended to limit the exemplary embodiments according to the present application. As used herein, the singular form is intended to include the plural form, unless the context clearly indicates otherwise, and it should also be understood that when the terms "comprise" and / or "include" are used in the specification, there is a presence of the features, steps, operations, devices, components and / or combinations thereof.

[0020] Unless specifically stated otherwise, the relative arrangements of the components and steps illustrated in these embodiments and the numerical expressions and values set forth herein are not limiting of the scope of the application. It should be understood that the various parts of the drawings are not necessarily drawn to scale, and that, for the purpose of convenience and clarity, not all components can be shown in a given figure. Techniques, methods, and apparatus known to those of ordinary skill in the relevant art can not be discussed in detail, but are to be considered as part of the description of the present application. In all examples shown and discussed herein, any specific values are to be interpreted as illustrative only and are not to be construed as limiting. Other examples of the exemplary embodiments can have different values.

[0021] As Figure 1 According to a specific embodiment of the present application, a controllable generation method of an infrared ship image for a marine scene is provided, which comprises:

[0022] Preparation of an infrared marine ship image dataset;

[0023] Construction of a controllable ship image generation framework for foreground-background separation, which adopts a denoising diffusion model architecture, including three stages of foreground target generation, background target generation and foreground-background target fusion;

[0024] Setting model training constraints;

[0025] Training the sub-diffusion models in the generation framework respectively;

[0026] The ship target to be generated, the position, the camera perspective and the sea condition are input into the generation framework, and the model output is a generated result of the target ship in the infrared sea scene meeting the control information requirements, so as to complete the ship image generation of the infrared sea scene.

[0027] By applying the configuration mode, an infrared ship image controllable generation method for a sea scene is provided. The method constructs a perfect controllable diffusion model generation framework with controllable foreground and background. After training, the framework can generate target ships clearly showing unique category features, and realize high realistic foreground and background fusion and position perception fusion. The method particularly emphasizes the accurate binding of small-scale targets in categories and attributes (such as direction information), and focuses on how to accurately simulate the modal features in the infrared scene.

[0028] Further, in the present application, first, an infrared sea ship image dataset is prepared.

[0029] In the present application, a suitable infrared ship image dataset is selected, image preprocessing is completed, and the data is divided into a training set, a validation set and a test set to ensure the effectiveness of model training.

[0030] Further, in the present application, after obtaining the infrared sea ship image dataset, a foreground and background separation controllable ship image generation framework is constructed.

[0031] In the present application, the framework adopts a denoising diffusion model architecture, which includes four sub-type diffusion models. Among them, the first and second sub-type diffusion models focus on learning the accurate category and attribute binding of small-scale targets, and adopt a pipeline architecture. Specifically, the first sub-type diffusion model is responsible for learning the contour information of the target, and the generated image mask is transmitted as control information to the second sub-type diffusion model; the second sub-type diffusion model learns the internal texture details of the target on this basis, and finally generates an infrared foreground target with obvious category attributes. In addition, the third sub-type diffusion model generates an infrared sea surface background with a specified perspective and sea condition through the joint control of text information and sketch information. The fourth sub-type diffusion model realizes the high realistic fusion and position controllable fusion of the foreground and background in the infrared scene.

[0032] As shown in Figure 2 , the foreground and background separation controllable ship image generation framework includes three stages of foreground target generation, background target generation and foreground and background target fusion.

[0033] (1) Foreground target generation stage

[0034] The application adopts a novel progressive generation strategy, sequentially learns the contour and texture of the target, and transmits the learned contour in the first stage to the second stage as control information to form constraints and guidance. This method effectively improves the problem of poor generation quality caused by low resolution of small-scale targets and blurred attribute information. The specific structure process of the foreground target generation stage is as follows:

[0035] 1) View and direction information text embedding: embed the required view and angle information through the frozen text encoder to ensure the accuracy of the generation effect.

[0036] 2) Control sketch image embedding: based on the text, control sketches can be added and embedded through the frozen image encoder to enhance the control information and reduce the learning difficulty.

[0037] 3) ControlNet contour generation structure: ControlNet locks all parameters in the UNet of the stable diffusion model and clones them to a trainable copy as a control network for training. It is mainly responsible for extracting control sketch information in the first stage and guiding the model to generate the corresponding results.

[0038] 4) Mask image embedding: the target mask generated by the first stage training is used as the control information of the second stage, and the frozen image encoder is used for embedding to reduce the learning difficulty.

[0039] 5) ControlNet texture generation structure: similar to the third step ControlNet contour generation structure, it is responsible for extracting the information of the mask in the second stage and guiding the model to perform the final foreground generation.

[0040] Since the UNet model of the stable diffusion model receives latent features (dimension 64x64) instead of original images, it is necessary to convert the image-based input conditions to the 64x64 feature space to match the convolution size in the model. Unlike processing images into latent space, a lightweight convolution network ε is used here to extract condition information features:

[0041] c f =ε(c i ),

[0042] The purpose of this step is to encode the input condition c i into control feature map c f .

[0043] ControlNet replicates the encoder and intermediate mapping module in the original stable diffusion model UNet as a controllable encoding network, and then connects it to the UNet decoder through a skip connection. Its forward process can be represented as:

[0044]

[0045] where m and y c represent the image deep features and transformed features encoded by neural network respectively, and c represents the provided additional conditions. denotes the mapping operation of the intermediate module in the neural network. z1 , Θ z2 and Θ c represent the parameters of the first zero convolution layer, the second zero convolution layer and the trainable copy respectively.

[0046] Furthermore, the main training object of ControlNet is the weight of the copied copy network, while the UNet in the stable diffusion model remains in a frozen state. The loss function used in training is the loss function L simple used in the original stable diffusion model to fit the noise.

[0047]

[0048] where c t and c f are the text input and control feature map respectively, which means that the stable diffusion model in ControlNet essentially receives double conditional input. denotes the calculation of expectation, and ∈ denotes the latent space noise. θ denotes the latent space noise calculated by the stable diffusion model, and z t denotes the latent space feature at diffusion time step t, and t denotes the diffusion time step.

[0049] (2) Background target generation stage

[0050] The present application converts the image channel in the post-processing link on the basis of the existing stable diffusion model to avoid the generation of visible light ghost. This method effectively improves the visible light prior interference that may be encountered when migrating from visible light to infrared scene. The specific structure process of the background target generation stage is as follows:

[0051] 1) Background description information text embedding: embed the generated view and sea information through the frozen text encoder to ensure that the generated image meets the expected conditions.

[0052] 2) Stable diffusion model structure: Stable diffusion model is a text-to-image model implemented based on latent diffusion model (LDM). Compared with denoising diffusion probabilistic model (DDPM), LDM introduces an autoencoder, so that the diffusion process is carried out in the latent space, improving the efficiency of image generation. At the same time, the conditional mechanism is added, so that the image generation can be controlled by other modal data, and the conditional generation control is realized through attention mechanism.

[0053] Further, since the diffusion model of LDM acts on the latent space, an encoder ε b Embed the original image x into the latent space to obtain z:

[0054] z=ε b (x),

[0055] Therefore, the objective function of the stable diffusion model in the latent space can be simplified as:

[0056]

[0057] This is still a loss calculation on the noise prediction result, where E ε(x),∈~N(0,1),t represents the calculation of expectation.

[0058] In addition, LDM converts the diffusion model into a more flexible conditional image generator by using cross-attention in the UNet model. The means is to introduce a new encoder τ θ to map the condition c to Then map to the middle layer of UNet through cross-attention layer. Therefore, the objective function of LDM under the condition control becomes:

[0059]

[0060] Encoder τ θ and UNet∈ θ are trained jointly through the above formula.

[0061] 3) Image conversion module: The weighted synthesis method is adopted to simulate the infrared sensing effect by combining different weights. Since the infrared camera responds differently to light of different wavelengths, and the red, green and blue channels of the RGB image correspond to different spectral ranges respectively, a reasonable weight distribution can be used to approximate the performance of the object in the infrared band.

[0062] (3) Foreground and background target fusion stage

[0063] The AnyDoor model is used to represent the foreground target as "ID-related" and "detail-related" features, which are then recombined with the generated background features. The specific structure of the foreground-background target fusion stage is as follows:

[0064] 1) High-frequency map extraction: The required perspective and angle information is embedded through a frozen text-side encoder to obtain high-frequency features of the background.

[0065] Specifically, to obtain the high-frequency map, a high-pass filter is used to extract the high-frequency region, and then a Hadamard product is used to extract the RGB color. In addition, an erosion mask is added to filter out information near the outer contour of the target object. This can be represented as:

[0066]

[0067] where K h and K v represent the horizontal and vertical Sobel kernels, respectively, used as high-pass filters; M erode represents the erosion mask; I represents the background map, and I h represents the high-frequency map.

[0068] 2) Custom position embedding: By pasting the high-frequency map of the foreground target into the specified position in the background, a position-controllable target fusion is achieved. This process ensures natural fusion of the foreground and background, enhancing the spatial consistency of the generated image.

[0069] 3) ID extractor: A pre-trained autoregressive model is used to extract the identity features of the foreground target, providing important semantic information for the generation process.

[0070] 4) Detail extractor: A ControlNet-style UNet encoder is used to generate a series of detail maps with hierarchical resolution. These detail maps enhance the visual representation of the foreground target, making it more realistic and vivid.

[0071] 5) Image generation module: The information obtained from the aforementioned feature extraction and detail extraction is injected into a stable diffusion model for image generation. In this process, the encoder module remains frozen, while the decoder module and other unfrozen parts are jointly trained to achieve more detailed image generation.

[0072] Furthermore, the AnyDoor model optimizes the objective function using a variant of the LDM optimization objective function, which predicts the original data. The formula is written as:

[0073]

[0074] where x is the real image, t is the diffusion time step, and the pre-trained UNet is denoted as x θ , which starts from the initial latent space noise and generates new latent space features using the text embedding c as a condition where z t represents the latent space feature at diffusion time step t, and t and t are denoising hyperparameters.

[0075] Further, in the present application, after the construction of the controllable ship image generation framework is completed, the model training constraints are set.

[0076] The present application adopts the simplified optimization objective based on the prediction noise ∈ proposed by the denoising diffusion probabilistic model (DDPM) and its variants as the loss function to constrain the model training process.

[0077] Further, in the present application, after setting the model training constraints, the sub-diffusion models in the generation framework are trained respectively.

[0078] In the present application, the network parameters are iteratively updated and optimized using the backpropagation algorithm until the model loss converges, thereby ensuring that each model can effectively learn its target features.

[0079] Further, in the present application, after the training of the sub-diffusion models in the generation framework is completed, the ship target, position, camera perspective, and sea conditions to be generated are input into the generation framework, and the model output is the target ship generation result in the infrared sea scene that meets the control information requirements, thereby completing the ship image generation in the infrared sea scene.

[0080] In summary, in view of the shortcomings of existing diffusion model control methods, the present application proposes a method for controllable generation of infrared ship image foreground and background in a sea scene. The method framework is based on the series and parallel generation strategy of the diffusion model, successfully solving the problem of existing algorithms in the comprehensive control of foreground and background, and achieving good results in the infrared scene. For foreground target generation, the present application introduces a novel progressive generation strategy, effectively improving the generation quality problem caused by low resolution and fuzzy attribute information of small-scale targets. This strategy significantly improves the clarity and accuracy of the generation results. On the basis of ensuring high-quality foreground and background generation, the fusion method of the present application realizes high-fidelity foreground and background fusion and has the function of position perception, meeting the requirement of position controllability.

[0081] The framework of the application can generate target ships clearly showing the characteristics of the specific species after training, and realize high realistic foreground and background fusion and position perception fusion. The method particularly emphasizes the accurate binding of small-scale targets in class and attributes (such as direction information), and focuses on how to accurately simulate the modal characteristics in the infrared scene.

[0082] In order to have a further understanding of the application, the offshore scene-oriented infrared ship image controllable generation method of the application will be described in detail below in combination with specific embodiments.

[0083] The embodiment provides an offshore scene-oriented infrared ship image controllable generation method, which specifically comprises the following steps.

[0084] 1. Infrared offshore ship data set preparation. Including data set selection, data preprocessing and data set division.

[0085] 1.1 Data set selection A high-quality infrared offshore ship data set. The data set includes more than 8400 labeled infrared data collected in different scenes. The image resolution is:

[0086] 384*288, 640*512, 1280*1024, and seven types of ship targets in the image are labeled.

[0087] 1.2 Data set preprocessing includes using liner, bulk carrier, sailboat, canoe, containership,

[0088] fishing boat as the label of cruise ship, bulk carrier, kayak, canoe, container ship and fishing boat, respectively, using a rectangular box to label the target in the picture, taking the upper left corner of the picture as the coordinate origin [0, 0], using the form [x1, y1, x2, y2] to record the position of the rectangular box, x1 represents the horizontal coordinate of the upper left corner of the rectangular box, y1 represents the vertical coordinate of the upper left corner of the rectangular box, x2 represents the horizontal coordinate of the right lower corner of the rectangular box, and y2 represents the vertical coordinate of the right lower corner of the rectangular box. All label information is saved in the form of xml file.

[0089] 1.3 Data set is divided according to the standard given by the publisher. A total of 8402 pictures are trained,

[0090] 1000 pictures without labels are tested.

[0091] 2. Design foreground and background generation fusion framework. For example Figure 2As shown, the foreground and background generation fusion framework is based on four diffusion models, including a novel progressive generation strategy to achieve fine control of small-scale foreground targets and high-realistic position-aware fusion of foreground and background. The specific model design has been discussed in the foregoing, and will not be repeated here. In this embodiment, the background target generation, foreground target generation, and foreground and background target fusion respectively use Stable Diffusion v1.4, Stable Diffusion v1.5, and Stable Diffusion v2.1 as the basic generator.

[0092] 3. Design model training constraints. For a conditional diffusion model such as Stable Diffusion, classifier-free guidance (CFG) is used in the inference stage:

[0093] ∈ pred =∈ uc +β cfg (∈ c -∈ uc ),

[0094] where β cfg is a fixed scaling factor, which is set to 0.5 in this embodiment, and ∈ uc and ∈ c are the noises predicted by the unconditional diffusion model (text is empty) and the conditional diffusion model, respectively. Since the foreground target generation and foreground and background target fusion stages are both double-condition diffusion models, the Balanced mode is adopted, and the conditional information is added to ∈ uc and ∈ c , i.e. and ∈ u =∈ θ (z t ,t,c t ,c f ), represents empty, i.e. no text input, so that the model constrains the text and image to be equally important.

[0095] 4. Train the model. Update the network parameter weights using the backpropagation algorithm until the model loss converges. In this embodiment, the training and evaluation of the four sub-diffusion models are all completed on the PyTorch platform. The model is trained on a single NVIDIA L20 GPU (48 GB) with a batch size of 1. The diffusion model is optimized using the Adam optimizer with an initial learning rate of 1×10 -5 for training, and then the learning rate is adjusted to 1×10 -6 for fine-tuning the model.

[0096] 4.1 For the foreground target generation stage: 8 epochs are needed for model training to fit, of which 2 epochs are needed for fine-tuning.

[0097] 4.2 For the background target generation stage: 5 epochs are needed for model training to fit, of which 2 epochs are needed for fine-tuning.

[0098] 4.3 For the foreground-background target fusion stage: only 2 epochs are needed for fine-tuning since it is from visible light to infrared modal.

[0099] 5. After the model training is completed, the sea ship infrared image generation is performed. The control condition is input into the generated model obtained by training, and the output of the model is the generation result. The sea scene ship infrared image generation task can be effectively completed by using the present application, and the attribute binding performance is good on different ship types, as shown in FIG. 6. Figure 3

[0100] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.​

Claims

1. A controllable generation method of infrared ship images for marine scenes, characterized in that: The controllable generation method of infrared ship images for marine scenes includes: Prepare infrared marine ship image dataset; A controllable ship image generation framework with foreground and background separation is constructed. The framework adopts a denoising diffusion model architecture and includes three stages: foreground target generation, background target generation, and foreground and background target fusion. The foreground and background separation controllable ship image generation framework includes four sub-diffusion models. The first and second sub-diffusion models focus on learning and generating precise categories and attribute bindings for small-scale targets. Using a pipeline architecture, the first sub-diffusion model is responsible for learning the contour information of the target and passing the generated image mask as control information to the second sub-diffusion model. The second sub-diffusion model learns the internal texture details of the target on this basis, and ultimately generates infrared foreground targets with obvious category attributes. The third sub-diffusion model generates an infrared sea surface background of a specified perspective and sea condition through the joint control of text information and sketch information. The fourth sub-diffusion model realizes high-realism fusion and position-controllable fusion of foreground and background in infrared scenes. Set model training constraints; Train the sub-diffusion models in the generation framework separately; The ship target, position, camera perspective and sea conditions to be generated are input into the generation framework, and the model output is the target ship generation result in the infrared marine scene that meets the control information requirements, completing the ship image generation of the marine infrared scene.

2. The controllable generation method of infrared ship images for marine scenes according to claim 1 is characterized in that: The specific structural process of the foreground target generation stage is as follows: 1) Text embedding of view and direction information: The required view and angle information is embedded through a frozen text encoder; 2) Control sketch image embedding: Based on the text, a control sketch is added and embedded through a frozen image encoder; 3) ControlNet outline generation structure: ControlNet locks all parameters in the UNet of the stable diffusion model and clones them into a trainable copy. It is trained as a control network and is responsible for extracting control sketch information in the first stage to guide the model to generate corresponding results. 4) Mask image embedding: The target mask generated by the first stage of training is used as the control information for the second stage and embedded using a frozen image encoder; 5) ControlNet texture generation structure: Similar to the ControlNet contour generation structure in the third step, it is responsible for extracting the information of the mask map in the second stage and guiding the model to perform the final foreground generation.

3. The controllable generation method of infrared ship images for marine scenes according to claim 2 is characterized in that: ControlNet replicates the encoder and intermediate mapping modules in the original stable diffusion model UNet as a controllable encoding network, which is then connected to the decoder of UNet through a jump connection. Its forward process is expressed as: Among them, m and y c They represent the deep features of the image encoded by the neural network and the transformed features, respectively, and c represents the additional conditions provided; represents zero convolution operation, Represents the mapping operation of the intermediate module of the neural network; parameter Θ z1 、Θ z2 and Θ c Represent the parameters of the first zero-convolutional layer, the second zero-convolutional layer, and the trainable copy, respectively.

4. The controllable generation method of infrared ship images for marine scenes according to claim 3 is characterized in that: The loss function L used in ControlNet training simple for: Among them, c t and c f They are text input and control feature maps; represents the computational expectation, ∈ represents the latent space noise, ∈ θ represents the latent space noise calculated by the stable diffusion model, z t represents the latent space feature when the diffusion time step is t, and t represents the diffusion time step.

5. The controllable generation method of infrared ship images for marine scenes according to claim 4 is characterized in that: The specific structural process of the background target generation stage is as follows: 1) Background description information text embedding: The required viewing angle and sea condition information are embedded through a frozen text encoder to ensure that the generated image meets the expected conditions; 2) Stable Diffusion Model Structure: The stable diffusion model is a text-based graph model implemented based on the latent space diffusion model. The latent space diffusion model introduces an autoencoder, allowing the diffusion process to proceed in the latent space. Furthermore, a conditional mechanism is added to control image generation by data from other modalities. Conditional generation control is achieved through an attention mechanism. 3) Image conversion module: It uses weighted synthesis method to simulate infrared sensing effects through different weight combinations.

6. The controllable generation method of infrared ship images for marine scenes according to claim 5 is characterized in that: In the stable diffusion model structure, the objective function L of the latent space diffusion model is LDM for: in, represents the expected value, τ θ (c) is the condition c through the encoder τ θ After the mapping.

7. The controllable generation method of infrared ship images for marine scenes according to claim 6 is characterized in that: The AnyDoor target fusion diffusion model is used in the foreground and background target fusion stage. The specific structure process is as follows: 1) High-frequency image extraction: Generate the required perspective and angle information and embed it through the frozen text-side encoder to obtain high-frequency features of the background; 2) Customized position embedding: By pasting the high-frequency image of the foreground object to the specified position in the background, position-controlled object fusion is achieved; 3) ID Extractor: Extracts identity features of foreground objects using a pre-trained autoregressive model; 4) Detail Extractor: Uses a ControlNet-style UNet encoder to generate a series of detail maps with hierarchical resolution; 5) Image generation module: Image generation is performed by injecting the information obtained from feature extraction and detail extraction into a stable diffusion model.

8. The controllable generation method of infrared ship images for marine scenes according to claim 7 is characterized in that: The objective function L optimized by the AnyDoor model is: in, represents the computational expectation, x is the real image, and the pre-trained UNet is denoted as x θ , α t and σ t are all denoising hyperparameters.

9. The controllable generation method of infrared ship images for marine scenes according to claim 1 is characterized in that: The simplified optimization objective based on prediction noise proposed by the denoising probability diffusion model and its variants are used as the loss function to constrain the model training process.

Citation Information

Patent Citations

  • Detection image generation method and system

    CN113936135A

  • Infrared small target detection method and device based on data enhancement

    CN117409192A