Training method of infrared coupling editing diffusion generation model

Through the proposed training method of infrared coupled editing diffusion generation model, the problems of poor interpretability, high cost and low efficiency in the existing infrared image editing generation methods are solved, and low cost, high efficiency generation and diversified data editing of infrared images are realized.

CN120182744APending Publication Date: 2025-06-20XIDIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510205478.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing infrared image coupled editing and generation methods have problems such as poor image interpretability, lack of artificial controllability, poor generation effect, and high training cost and low efficiency. It is especially difficult to achieve low-cost and high-efficiency diversified data editing and generation under infrared mode.

Method used

A training method for infrared coupled editing diffusion generation model is proposed. By obtaining the sample pure infrared target image and background image, the pre-trained feature extractor is used to extract the overall features and detailed features, and the model is constructed based on U-Net for fine-tuning training, gradually adding noise and denoising, achieving efficient generation of infrared images.

Benefits of technology

It realizes low-cost and efficient generation of infrared images, has stronger generalization and editability, and can accurately and efficiently realize the goal and background coupled editing training tasks of infrared modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182744A_ABST
    Figure CN120182744A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, in particular to a training method of an infrared coupling editing diffusion generation model, which comprises the following steps of: inputting a pure sample infrared target image into an overall feature extractor, respectively encoding the pure sample infrared target image into a global token and a patch token in a token mode, splicing the two kinds of tokens and then aligning the tokens to a submerged space required by training to obtain an overall feature; processing the sample pure infrared target image through high-pass filtering, wavelet transform or discrete cosine transform to obtain a high-frequency mapping graph, fitting the high-frequency mapping graph to a specified position of a sample background image, and inputting the high-frequency mapping graph into a detail feature extractor to obtain detail features; and constructing an infrared coupling editing diffusion generation model, and injecting the overall features and the detail features into the infrared coupling editing diffusion generation model together for training, thereby rapidly training to obtain the infrared coupling editing diffusion generation model capable of regulating and controlling the shape and posture of the target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer vision technology, and particularly to a training method for an infrared-coupled editing diffusion generation model. Background Art

[0002] At present, artificial intelligence technology has made great progress in the generation and editing of images and videos. Generative intelligent large models have also received extensive attention and applications, mainly due to their excellent content generation capabilities and broad application potential. Generative intelligent large models aim to learn the latent distribution of data and generate new samples similar to the training data. Through deep learning technology, they learn and extract key information from massive data, and then generate high-quality and diverse new content. Whether it is text, images, or audio, generative intelligent large models can present in a highly realistic and creative manner. In the fields of graphic design, image synthesis, medical diagnosis, etc., generative intelligent large models are gradually assisting or replacing humans to complete tedious creative work, greatly improving production efficiency and creative quality, and significantly promoting the in-depth application of artificial intelligence technology in various industries.

[0003] With the popularity of generative intelligent large models such as ChatGPT4V and Sora for generating images from text and videos from text, the topic of intelligent data generation has attracted much attention. Since traditional data acquisition methods are basically based on actual shooting or simulation using a simulation platform, the data acquisition efficiency is low and the speed is slow. For some scarce special scenarios or target data, they cannot even be obtained. In contrast, with the development of deep learning technology, the advantages of intelligent generation algorithms in terms of high efficiency and low cost have become more prominent. They can effectively reduce the acquisition cost of specific data, and can also create new data content, and even some content is difficult to directly obtain in the real world. Generally speaking, generative intelligent large models are key tools for discovering the essence of data and finding the best representation form. With the help of these models, people can quickly generate samples that are not in the training set but follow the same distribution, thus solving the problems of difficult data acquisition and lack of diversity.

[0004] In recent years, in the field of infrared image coupled editing and generation, research teams at home and abroad have proposed many novel methods. For example, the relatively classic Dreambooth model, which fine-tunes a pre-trained diffusion model by binding scarce characters representing the target and the actual prior category. It can generate edited images of the target with 3 to 5 images and text descriptions of the specified target. Another example is the Instruct-pix2pix method, which is completely driven by training data. It obtains editing instruction text and input-output image pairs before and after editing through intelligent methods, and retrains the pre-trained model to achieve editing tasks such as style editing of the target.

[0005] However, there are still some problems with the image coupling editing and generation methods proposed so far. For example, the method of generating images solely through text editing instructions results in poor image interpretability and lack of human controllability. There are also phenomena such as poor generation effects of the trained generative intelligent large model and inability to accurately understand the user's intention. The method of fine-tuning and training the generative intelligent large model completely based on training data can achieve good results and has strong editability, but it requires a large amount of paired data before and after editing, and often other controllable parameters involved in editing, which leads to a doubling of training costs and low generation and editing efficiency. Directly applying the visible light-based image coupling editing and generation method to the infrared modality will result in inaccurate target characteristics and poor controllability of the target, thus leading to low quality of the infrared data generated by coupling editing and being unable to achieve the editing and generation of diverse data at low cost and high efficiency. Summary of the Invention

[0006] In view of this, an embodiment of the present application proposes a training method for an infrared coupled editing diffusion generation model, aiming to quickly train an infrared coupled editing diffusion generation model for the infrared modality that can regulate the shape and pose of the target, so as to generate the required infrared images at low cost and high efficiency.

[0007] To achieve the above object, an embodiment of the present application proposes a training method for an infrared coupled editing diffusion generation model, the method comprising the following steps: obtaining a sample pure infrared target image and a sample background image; inputting the sample pure infrared target image into a pre-trained overall feature extractor, and encoding the sample pure infrared target image into a global token and a patch token respectively in the form of tokens, splicing the global token and the patch token and then aligning them into the latent space required for training the infrared coupled editing diffusion generation model to obtain the overall feature of the sample pure infrared target image; processing the sample pure infrared target image through high-pass filtering, wavelet transform, or discrete cosine transform to obtain a high-frequency mapping of the sample pure infrared target image, fitting the high-frequency mapping to a specified position of the sample background image and then inputting it into a pre-trained detail feature extractor to obtain the detail feature of the sample pure infrared target image; constructing an infrared coupled editing diffusion generation model based on U-Net, injecting the overall feature and the detail feature into the infrared coupled editing diffusion generation model together to guide fine-tuning training, and obtaining the trained infrared coupled editing diffusion generation model through the way of gradually adding noise and gradually removing noise to the model.

[0008] To achieve the above object, an embodiment of the present application further provides a training system for an infrared coupled editing diffusion generation model, the system comprising: a general feature extractor training module for constructing and training a general feature extractor; a detailed feature extractor training module for constructing and training a detailed feature extractor; a sample acquisition module for acquiring a sample pure infrared target image and a sample background image; a general feature extraction module for inputting the sample pure infrared target image into the general feature extractor, encoding the sample pure infrared target image into global tokens and patch tokens respectively in a token manner, splicing the global tokens and patch tokens and then aligning them into the latent space required for the training of the infrared coupled editing diffusion generation model to obtain the general features of the sample pure infrared target image; a detailed feature extraction module for processing the sample pure infrared target image through high-pass filtering, wavelet transform, or discrete cosine transform to obtain a high-frequency mapping of the sample pure infrared target image, fitting the high-frequency mapping to a specified position of the sample background image and then inputting it into a pre-trained detailed feature extractor to obtain the detailed features of the sample pure infrared target image; a model construction module for constructing an infrared coupled editing diffusion generation model based on U-Net; a model training module for injecting the general features and the detailed features into the infrared coupled editing diffusion generation model together to guide fine-tuning training, and obtaining a trained infrared coupled editing diffusion generation model by gradually adding noise and gradually removing noise to the model.

[0009] To achieve the above object, an embodiment of the present application further provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute a training method for an infrared coupled editing diffusion generation model as described above.

[0010] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it can implement a training method for an infrared coupled editing diffusion generation model as described above.

[0011] A training method for an infrared coupled editing diffusion generation model proposed in an embodiment of the present application encodes an image into global tokens and patch tokens during the process of extracting target overall feature information, and then connects these two types of tokens to retain more feature information. When extracting target detailed feature information, methods including but not limited to high-pass filtering, wavelet transform, discrete cosine transform, etc. are used to supplement the detailed information ignored during overall feature extraction. The pure infrared target after removing the background is extracted as a high-frequency mapping graph and tiled to a given position of the background image, and an information bottleneck is set to prevent too many appearance constraints during the tiling process, thereby enhancing the diversity of the model generation results. During the training process, in order to simulate that the user is not fine enough when drawing the shape of the infrared target, the real infrared target mask input during training is downsampled at different ratios, and random dilation or erosion is applied to remove the details of the infrared target mask in the training data. Finally, the input infrared target and the infrared target mask are connected. The infrared coupled editing diffusion generation model trained in this way can regulate the target shape and posture, and achieve low-cost and high-efficiency coupled generation. Compared with the proposed technologies, it has stronger generalization and editability, and can more accurately and efficiently implement the infrared modality target and background coupled editing training task.

[0012] Optionally, the infrared coupled editing diffusion generation model constructed based on U-Net is a symmetric U-shaped fully convolutional neural network, which consists of a large number of convolutional layers, pooling layers, fully connected layers, and skip connections for supplementing lost semantic information. The infrared coupled editing diffusion generation model constructed based on U-Net mainly includes two parts: an encoder and a decoder; the encoder is responsible for further feature extraction and learning of the input overall features and detailed features, increasing the robustness to perturbation noise, reducing the risk of overfitting, reducing the computational amount, and increasing the receptive field; each convolutional layer of the encoder consists of two 3×3 convolutions, and each convolutional layer is followed by a rectified linear unit and a 2×2 max pooling operation with a stride of 2. The max pooling operation is used to achieve downsampling, and the number of channels of the feature is doubled during each downsampling process; the decoder is responsible for restoring the feature map extracted by the encoder to the original resolution, and fuses the shallow position information and the deep semantic information; each step of the decoder includes upsampling the feature, then halving the number of channels of the feature through a 2×2 convolution, connecting it with the corresponding cropped feature map in the contraction path, and then passing through two 3×3 convolutions. Each 3×3 convolution is followed by a ReLU activation function layer, and a 1×1 convolution is used in the last layer to map each feature vector to the required number of classes.

[0013] Optionally, the operation of the infrared coupled editing diffusion generation model mainly consists of two stages: a forward process and a backward process;

[0014] For the initial sample \(x_0\in q(x_0)\), by define the forward process as a Markov chain with a stationary distribution being a Gaussian distribution. Let \(T\) be the end time, i.e., the total number of time steps. is a mathematical symbol used to describe the asymptotic behavior of a function, and := means to make a definition.

[0015] In the forward process, random Gaussian noise is slowly added to the data sample. Therefore, the forward process is also called the diffusion process. The transition probability at each step is constructed by the following formula to add Gaussian noise:

[0016]

[0017] where \(\beta\) t is the variance of the noise, \(\beta\) t \(\in(0,1)\), \(\beta\) t \(>\beta\) t-1 and \(\beta\) t with the values at time step \(t\) being called the variance table, which is default specified. As the time step \(t\) increases, the noisy sample \(x\) t gets closer and closer to \(x\) T , and more and more noise is added.

[0018] Let \(\beta\) t \(=1 - \alpha\) t . Through the following formula, \(q(x\) t |x\) t-1 ) is extended to the form with respect to the initial sample \(x_0\):

[0019]

[0020] where \(\alpha\) t is a function decreasing with respect to the time step \(t\). As the time step \(t\) increases, the data weight gradually decreases and the noise weight gradually increases. When \(t\rightarrow\infty\), that is to say, when \(t\) is large enough, \(q(x\) t |x_0)\) will converge to a standard Gaussian distribution independent of \(x_0\), i.e., \(\lim\) t→∞ \(q(x\) t ) = 0.

[0021] Based on this, in the continuous-time model, the diffusion process is described as a continuously changing process in time. Assume the termination time of the continuous-time diffusion process is 1, then \(t\in(0,1)\), and \(t = 0\) and \(t = 1\) correspond to the two states of no noise and full noise respectively.

[0022] Starting from the data sample \(x\), the noisy sample corresponding to the diffusion process at any \(t\in(0,1)\) is \(z\) t , and can be written as:

[0023]

[0024] where α t and are both strictly positive scalar functions of the time step t;

[0025] In addition, define the signal-to-noise ratio function as SNR(t), SNR(t) is a strictly monotonically decreasing function at the time step t, visualizing the concept that the noisy sample z t becomes noisier and noisier as the time step t progresses;

[0026] where,

[0027] Let Then we can get That is Substituting it into SNR(t), we can get SNR(t) = exp[-γ(t)];

[0028] When γ(t) = log[expml(e -4 + 10t 2 ), the diffusion model in discrete time can be related to the continuous time model, regarding the discrete time model as the discrete form of the continuous time model, and the two can be transformed into each other.

[0029] Optionally, the reverse process is the inverse process of the forward process, also known as the denoising process, mainly denoising a random noise matrix step by step until an image is generated. The noise distribution predicted and removed at each step needs to be learned by the infrared-coupled editing diffusion generation model during training;

[0030] In the reverse process, according to Bayes' formula q(z t-1 |z t ) = [q(z t-1 |z t )′q(z t-1 )] / q(z t ), we can obtain the q(x t-1 |x t ) of the reverse process. Since the marginal distributions q(z t ) and q(z t-1 ) are unknown, q(z t-1 |z t ) is rewritten as:

[0031]

[0032] Due to the Markov chain property of the forward process, q(z t-1 |z t ,x) can be simplified to:

[0033]

[0034] Substituting the simplified formula again, we get:

[0035]

[0036] Among them, the mean value is a function that depends on x and z t .

[0037] Since the value of x is unknown, a neural network is designed to estimate the result of x. However, it is very difficult to directly estimate the generated samples from the noise. Therefore, using the reparameterization trick, we get:

[0038]

[0039] Among them,

[0040] Through simple transformation, we get:

[0041]

[0042] Based on this, given the noisy sample z t , only the noise t needs to be estimated from z to calculate x, and then the next sample z t-1 can be calculated. In this process, x is used as an intermediate variable, so that the reverse process can be freed from the time step limit in the discrete diffusion model and the time step can be set arbitrarily.

[0043] For the continuous time model, define the samples corresponding to any two time steps s and t satisfying 0 ≤ s ≤ t ≤ 1 in the reverse process as z s and z t , and we have:

[0044]

[0045] Further rewrite p(z s |z t ) as:

[0046]

[0047] γ(t) = log[expml(e -4 + 10t 2 )]

[0048] Among them, ε ∈ N(0, I).

[0049] Optionally, an infrared coupled editing diffusion generation model is constructed based on U-Net. The overall features and detailed features are injected into the infrared coupled editing diffusion generation model to guide the fine-tuning training. Through the way of gradually adding noise and gradually denoising by the model, a trained infrared coupled editing diffusion generation model is obtained, including:

[0050] First, sample training samples \(x\in q(x)\) from the data distribution \(q(x)\), then randomly generate a time \(t\) and calculate the variance and \(\alpha\) t , and then sample a from the standard Gaussian distribution to add noise to the training sample \(x\) and generate \(z\) t and input it into the infrared coupled editing diffusion generation model to obtain the predicted noise Finally, calculate and the added noise to calculate the loss between them and perform gradient update;

[0051] In the sampling process of generating samples, set the sampling step to 1000 steps, then the step size of each sampling is 1 / 1000. First, sample a standard Gaussian noise i.e., \(z_1\) from the noise space, and use \(z_1\) and to calculate the further denoised sample \(z\) 0.999 , and so on until \(z\) 0.001 is denoised to obtain the final desired generated sample \(z_0\);

[0052] Represent the infrared coupled editing diffusion generation model as Start denoising from the initial latent noise and generate a new image latent function conditioned on the overall feature \(c\) The mean squared error loss for training supervision is:

[0053]

[0054] where \(X\) is the true image latent function, \(t\) is the diffusion time step, \(\alpha\) t and \(\sigma\) t are denoising hyperparameters, and \(c\) is the overall feature;

[0055] The overall feature is injected into each U-Net layer through cross-attention, and the detailed feature is connected to the U-Net decoder feature at each resolution. During training, the pre-trained parameters of the U-Net encoder are frozen to maintain the prior, and the U-Net decoder is adjusted to adapt to the current infrared target and background coupling editing task.

[0056] Optionally, the backbone of the overall feature extractor is the DINOv2 network, and the overall feature extractor is trained using a vision self-supervised learning method, namely a knowledge distillation method for training without labels;

[0057] In terms of self-supervised learning, DINOv2 uses a discriminative self-supervised method to learn features, which can be regarded as a combination of the DINO loss centered on SwAV and the iBOT loss, and also adds a regularizer to propagate features and a short high-resolution training stage;

[0058] In the image-level information, the training considers the cross-entropy loss between the features extracted from the student network and the teacher network. These two features both come from the class tokens of ViT and are obtained from different targets of the same image. The student class token is passed through the student DINO head, whose head is an MLP model that outputs a score vector called the student prototype score, and then Softmax is applied to obtain the probability distribution p of the student network s , the teacher DINO head is applied to the teacher class token to obtain the teacher prototype score, and then Softmax is applied to obtain the probability distribution p of the teacher network t ;

[0059] The DINO loss is expressed by the formula:

[0060] L DINO = -∑p t log(p s )

[0061] where L DiNO represents the DINO loss;

[0062] In the pixel-level information, the training randomly masks some input segments given to the student network, and then the student iBOT head is applied to the student masked tokens. Similarly, the teacher iBOT head is applied to the visible teacher patch tokens, which correspond to the tokens masked in the student. Then Softmax is applied to obtain the probability distribution p of the student network si and the probability distribution p of the teacher network ti ;

[0063] The iBOT loss is expressed by the formula:

[0064] L DINO = -∑ i p ti log(p si )

[0065] where i is the patch index of the masked token;

[0066] Both the DINO loss and the iBOT loss use a learnable MLP projection head, which is applied to the output token information and the loss is calculated, finally obtaining the overall pre-trained feature extractor.

[0067] Optionally, the sample pure infrared target image is processed by high-pass filtering to obtain the high-frequency mapping of the sample pure infrared target image, which is implemented by the following formula:

[0068]

[0069] where K h and K v respectively represent the horizontal high-pass Sobel operator and the vertical high-pass Sobel operator, acting as the high-pass filter convolution kernel, I represents the sample pure infrared target image, and I gray represents the grayscale image of the sample pure infrared target image, M erode is the erosion mask, used to filter out the feature information near the outer contour of the infrared target, and ⊙ represents the Hadamard product operation, represents the convolution operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0071] Figure 1 is a flowchart of a training method for an infrared coupled editing diffusion generation model provided in an embodiment of the present application;

[0072] Figure 2 is a schematic diagram of deep convolution provided in an embodiment of the present application;

[0073] Figure 3 is a structural diagram of model training provided in an embodiment of the present application;

[0074] Figure 4 is a structural diagram of the U-Net provided in an embodiment of the present application;

[0075] Figure 5 is a schematic diagram of the principle of model training provided in another embodiment of the present application;

[0076] Figure 6 is a structural schematic diagram of a training system for an infrared coupled editing diffusion generation model provided in another embodiment of the present application;

[0077] Figure 7 It is a schematic structural diagram of an electronic device provided in another embodiment of the present application. Specific implementation manners

[0078] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. Those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are proposed for readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation to the specific implementation manners of the present application. The various embodiments can be combined and cross-referenced with each other on the premise of not conflicting with each other.

[0079] An embodiment of the present application proposes a training method for an infrared-coupled editing diffusion generation model, which is applied to an electronic device. Among them, the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the server is taken as an example for illustration. The implementation details of a training method for an infrared-coupled editing diffusion generation model proposed in this embodiment will be specifically described below. The following content is only implementation details provided for easy understanding and is not necessary for implementing this solution.

[0080] The specific process of a training method for an infrared-coupled editing diffusion generation model proposed in this embodiment can be as Figure 1 shown and includes:

[0081] Step 101, obtain a sample pure infrared target image and a sample background image.

[0082] In specific implementation, the server first needs to obtain a sample pure infrared target image and a sample background image. The sample background image is relatively easy to obtain, but the sample pure infrared target image needs to be obtained based on a sample infrared image containing an infrared target, that is, a pre-trained segmentation model needs to be used to remove the background of the infrared target in the sample infrared image and align the infrared target to the center of the image, and then a sample pure infrared target image can be obtained.

[0083] The segmentation model selects a convolutional neural network. Taking Mobilenetv2 in SegmenTron as an example, Mobilenet is a lightweight convolutional neural network that mainly proposes depthwise separable convolution. Depthwise separable convolution splits a conventional convolution into a depth convolution and a point convolution composed of 1×1 convolution kernels. As Figure 1 shown, the use of depth convolution greatly reduces the number of parameters under the condition of ensuring accuracy.

[0084] Depth convolution performs convolution operations separately for each channel of the input feature map. This operation is used to extract the spatial information of the input feature map while greatly reducing the number of parameters. Point convolution can be regarded as a conventional convolution operation. The difference is that the spatial size of its convolution kernel is 1×1. A 1×1 convolution kernel is used to perform point-by-point convolution on the feature map obtained in the previous step, thereby restoring the number of channels of the input feature map. This step is mainly used to fuse the useful information between channels.

[0085] Before performing depth convolution operations, Mobilenetv2 first expands the number of channels using point convolution and does not use a non-linear activation function at the end. At the same time, MobileNetv2 also borrows from the ResNet structure, designs a residual structure, and replaces the convolutional layers in it with depth convolution.

[0086] Step 102: Input the sample pure infrared target image into the pre-trained overall feature extractor. By means of tokens, the sample pure infrared target image is encoded into global tokens and patch tokens respectively. After concatenating the global tokens and patch tokens, align them into the latent space required for training the infrared coupled editing diffusion generation model to obtain the overall features of the sample pure infrared target image.

[0087] In a specific implementation, after the server obtains the sample pure infrared target image and the sample background image, it needs to input the sample pure infrared target image into the pre-trained overall feature extractor. By means of tokens, the sample pure infrared target image is encoded into global tokens and patch tokens respectively. Subsequently, the global tokens and patch tokens are concatenated, and then the concatenated tokens are aligned into the latent space required for training the infrared coupled editing diffusion generation model to obtain the overall features of the sample pure infrared target image.

[0088] In an example, a self-supervised model is selected as the overall feature extractor. Currently, most methods mainly use the CLIP network. The CLIP network includes a text encoder and an image encoder. This network has strong generalization performance and excellent performance in zero-shot tasks, so it is widely used in extracting semantic information. However, the CLIP network is slightly inferior to the DINO network in terms of performance. Therefore, the server selects the DINOv2 network as the backbone of the overall feature extractor.

[0089] The training of the overall feature extractor adopts a visual self-supervised learning method, which is a knowledge distillation method for training without labels. Knowledge distillation is a learning paradigm where a student network matches a given teacher network, and both networks output probability distributions on the K-dimension for a given input image. DINO simplifies the self-supervised training process by directly predicting the output of the teacher network using the standard cross-entropy loss. Two perspectives of an image are respectively input into the student network and the teacher network. The two networks have the same architecture but different parameters. The output of the teacher network is centralized through the average value of one round. Each network outputs a multi-dimensional feature, which is normalized using Softmax, and then the cross-entropy loss is used as the objective function to calculate the similarity between the student network and the teacher network. The stop-gradient operator is used on the teacher network to block the propagation of gradients, and only the gradients are passed to the student network to update its parameters, and then the teacher network is updated using the parameters of the student network.

[0090] In contrast, DINOv2 retains the best parts of its predecessor and adds multiple enhancements to achieve more efficient training and better performance. In terms of pre-training the model on a large amount of data, by generating general visual features, the model can greatly simplify the use of images in any system. DINOv2 summarizes existing methods and combines different techniques to expand the model pre-training in terms of data and model size.

[0091] In terms of self-supervised learning, DINOv2 uses a discriminative self-supervised method to learn features, which can be regarded as a combination of the DINO loss centered on SwAV and the iBOT loss, and also adds a regularizer to propagate features and a short high-resolution training stage.

[0092] In the image-level information, the training considers the cross-entropy loss between the features extracted from the student network and the teacher network. These two features both come from the class tokens of ViT and are obtained from different objects of the same image. The student class token is passed through the student DINO head, whose head is an MLP model that outputs a score vector called the student prototype score, and then Softmax is applied to obtain the probability distribution p of the student network s , the teacher DINO head is applied to the teacher class token to obtain the teacher prototype score, and then Softmax is applied to obtain the probability distribution p of the teacher network t .

[0093] The DINO loss is expressed by the formula:

[0094] L DINO = -∑p t log(p s )

[0095] where LDINO Represents the DINO loss.

[0096] In pixel-level information, during training, some input segments to the student network are randomly masked, and then the student iBOT head is applied to the student masked tokens. Similarly, the teacher iBOT head is applied to the visible teacher patch tokens, which correspond to the tokens masked in the student. Then, Softmax is applied to obtain the probability distribution p of the student network si and the probability distribution p of the teacher network ti .

[0097] The iBOT loss is expressed by the formula:

[0098] L DINO = -∑ i p ti log(p si )

[0099] where i is the patch index of the masked token.

[0100] Both the DINO loss and the iBOT loss use a learnable MLP projection head, which is applied to the output token information and calculates the loss, finally obtaining a pre-trained overall feature extractor.

[0101] Step 103, process the sample pure infrared target image through high-pass filtering, wavelet transform, or discrete cosine transform to obtain the high-frequency mapping of the sample pure infrared target image. After fitting the high-frequency mapping to the specified position of the sample background image, input it into the pre-trained detail feature extractor to obtain the detail features of the sample pure infrared target image.

[0102] In specific implementation, in addition to the overall features, detail features also need to be extracted for supplementation. The server can process the sample pure infrared target image through high-pass filtering, wavelet transform, or discrete cosine transform to obtain the high-frequency mapping of the sample pure infrared target image. After fitting the high-frequency mapping to the specified position of the sample background image, input it into the pre-trained detail feature extractor to obtain the detail features of the sample pure infrared target image.

[0103] Taking high-pass filtering as an example, processing the sample pure infrared target image through high-pass filtering to obtain the high-frequency mapping of the sample pure infrared target image can be achieved through the following formula:

[0104]

[0105] where K h and K vrespectively represent the horizontal high-pass Sobel operator and the vertical high-pass Sobel operator, acting as the convolution kernels of the high-pass filter. I represents the sample pure infrared target image, and I gray represents the grayscale image of the sample pure infrared target image, and M erode is the erosion mask, used to filter out the feature information near the outer contour of the infrared target. ⊙ represents the Hadamard product operation, represents the convolution operation.

[0106] Then, given the sample pure infrared target image I, first use these methods for obtaining high-pass information to extract the high-frequency mapping diagram of the infrared target, and then use the Hadamard product to extract the RGB color. In addition, an erosion mask M erode is added, which is used to filter out the feature information near the outer contour of the infrared target and better retain the required detailed features.

[0107] It should be noted that fitting the high-frequency mapping diagram to the specified position of the sample background image can provide powerful prior knowledge during model training. Through this tiling, the generated fidelity of the coupling is relatively good, but the output result is too similar to the given infrared target, and the generated result is too single. To address this issue, an information bottleneck needs to be set to prevent too many appearance constraints during the tiling process, thereby enhancing the diversity of the model generation results.

[0108] Correspondingly, in order to simulate the user's rough shape drawing of the infrared target when using the trained model, the mask of the infrared target is required to indicate the pose of the infrared target. During the network training process, in order to make the model understand that the user's drawing of the infrared target shape is not fine enough, the real infrared target mask input during training is downsampled at different ratios, and random dilation or erosion is applied to remove the details of the infrared target mask in the training data. Finally, the input infrared target and the infrared target mask are connected and input into a U-Net encoder as a detail feature extractor, and a series of detail features with hierarchical resolutions can be output.

[0109] Step 104: Based on U-Net, construct an infrared coupled editing diffusion generation model, inject the overall features and detail features into the infrared coupled editing diffusion generation model to guide fine-tuning training, and obtain the trained infrared coupled editing diffusion generation model through the way of gradually adding noise and gradually removing noise in the model.

[0110] In the specific implementation, after the server obtains the overall features and detail features, it is necessary to construct an infrared coupled editing diffusion generation model based on U-Net, inject the overall features and detail features into the infrared coupled editing diffusion generation model to guide fine-tuning training, and obtain the trained infrared coupled editing diffusion generation model through the way of gradually adding noise and gradually removing noise in the model.

[0111] In one example, the training of the infrared-coupled editing diffusion generation model is as follows Figure 3 As shown, the overall features and detailed features are injected into the infrared-coupled editing diffusion generation model together.

[0112] In one example, as Figure 4 shown, the infrared-coupled editing diffusion generation model constructed based on U-Net is a symmetric U-shaped fully convolutional neural network, which consists of a large number of convolutional layers, pooling layers, fully connected layers, and skip connections for supplementing the lost semantic information. The infrared-coupled editing diffusion generation model constructed based on U-Net mainly includes two parts: an encoder and a decoder.

[0113] The encoder is responsible for further feature extraction and learning of the input overall features and detailed features, increasing the robustness to perturbation noise, reducing the risk of overfitting, reducing the computational amount, and increasing the receptive field.

[0114] Each convolutional layer of the encoder is composed of two 3×3 convolutions. Behind each convolutional layer, there follows a rectified linear unit and a 2×2 max pooling operation with a stride of 2. The max pooling operation is used to achieve downsampling, and during each downsampling process, the number of channels of the features is doubled.

[0115] The decoder is responsible for restoring the feature map extracted by the encoder to the original resolution and fusing the shallow position information and the deep semantic information.

[0116] Each step of the decoder includes upsampling of the features, then reducing the number of channels of the features by a 2×2 convolution, connecting with the corresponding cropped feature map in the contracting path, and then passing through two 3×3 convolutions. Behind each 3×3 convolution, there is connected a ReLU activation function layer. In the last layer, a 1×1 convolution is used to map each 64-component feature vector to the required number of classes. The decoder has a total of 23 convolutional layers.

[0117] The operation of the infrared-coupled editing diffusion generation model mainly consists of two stages: a forward process and a backward process.

[0118] For the initial sample \(x_0\in q(x_0)\), through the forward process is defined as a Markov chain with a stationary distribution being a Gaussian distribution. \(T\) is the end time, that is, the total number of time steps, is a mathematical symbol used to describe the asymptotic behavior of a function, and := means to make a definition.

[0119] In the forward process, random Gaussian noise is slowly added to the data sample. Therefore, the forward process is also called the diffusion process. The transition probability of each step is constructed by the following formula to add Gaussian noise:

[0120]

[0121] Among them, β t is the variance of the noise, β t ∈(0,1), β t >β t-1 , β t The value of β with respect to the time step t is called the variance table, which is default specified. As the time step t increases, the noisy sample x t gets closer and closer to x T , and more and more noise is added.

[0122] It should be noted that the variance of the noise increases linearly from β1 = 10 -4 to β T = 0.02.

[0123] Let β t = 1 - α t , and through the following formula, q(x t |x t-1 ) is extended to the form with respect to the initial sample x0:

[0124]

[0125] Among them, α t is a function that decreases with respect to the time step t. As the time step t increases, the data weight gradually decreases and the noise weight gradually increases. When t → ∞, That is to say, when t is large enough, q(x t |x0) will converge to a standard Gaussian distribution independent of x0, that is, lim t→∞ q(x t ) = 0.

[0126] Based on this, in the continuous-time model, the diffusion process is described as a continuously changing process in time. Assuming that the termination time of the continuous-time diffusion process is 1, then t ∈ (0,1), and t = 0 and t = 1 correspond to the two states of no noise and full noise respectively.

[0127] Starting from the data sample x, the noisy sample corresponding to the diffusion process at any t ∈ (0,1) is z t , which can be written as:

[0128]

[0129] Among them, α t and are both strictly positive scalar functions with respect to the time step t.

[0130] In addition, the signal-to-noise ratio function is defined as SNR(t), SNR(t) is a strictly monotonically decreasing function at time step t, visualizing the noisy sample z t The concept of becoming noisier as time step t progresses.

[0131] Among them,

[0132] Let Then we can obtain That is Substituting into SNR(t), we can get SNR(t) = exp[-γ(t)].

[0133] When γ(t) = log[expml(e -4 +10t 2 ), the diffusion model in discrete time can be related to the continuous time model. The discrete time model can be regarded as the discrete form of the continuous time model, and the two can be transformed into each other.

[0134] The reverse process is the inverse process of the forward process, also known as the denoising process. It mainly denoises a random noise matrix step by step until an image is generated. The noise distribution predicted and removed at each step needs to be learned by the infrared-coupled editing diffusion generation model during training.

[0135] In the reverse process, according to Bayes' formula q(z t-1 |z t ) = [q(z t-1 |z t )′q(z t-1 )] / q(z t ), we can obtain the q(x t-1 |x t ) of the reverse process. Since the marginal distributions q(z t ) and q(z t-1 ) are unknown, q(z t-1 |z t ) is rewritten as:

[0136]

[0137] Due to the Markov chain property of the forward process, q(z t-1 |z t ,x) can be simplified to:

[0138]

[0139] Substituting the simplified formula again, we can obtain:

[0140]

[0141] Among them, the mean is a function that depends on x and z t .

[0142] Since the value of x is unknown, a neural network is designed to estimate the result of x. However, it is difficult to directly estimate the generated samples from the noise. Therefore, using the reparameterization trick, we can get:

[0143]

[0144] where

[0145] Through simple transformation, we can get:

[0146]

[0147] Based on this, given the noisy sample z t , we only need to estimate the noise t from z to calculate x, and then we can calculate the next sample z t-1 . In this process, x is used as an intermediate variable, which can make the reverse process get rid of the limitation of the time step in the discrete diffusion model and set the time step arbitrarily;

[0148] For a continuous time model, define the samples corresponding to any two time steps s and t satisfying 0 ≤ s ≤ t ≤ 1 in the reverse process as z s and z t , and we have:

[0149]

[0150] Further rewrite p(z s |z t ) as:

[0151]

[0152] γ(t) = log[expml(e -4 + 10t 2 )]

[0153] where ε ∈ N(0, I).

[0154] In an example, the training principle of the infrared coupling editing diffusion generation model is as Figure 5 shown. The model is trained using the derivation result of the continuous form of the diffusion process. During training, first sample the training sample x ∈ q(x) from the data distribution q(x), then randomly generate a time t and calculate the variance and α t , and then sample a Add noise to the training sample x to generate z t And input it into the infrared-coupled editing diffusion generation model to obtain the predicted noise Finally, calculate The loss between and the added noise ò, and perform gradient update.

[0155] During the sampling process of generating samples, set the sampling step to 1000 steps, then the step size of each sampling is 1 / 1000. First, sample a standard Gaussian noise ò from the noise space, that is, z1, and use z1 and Calculate the further denoised sample z 0.999 , and so on, until z 0.001 Is denoised to obtain the finally desired generated sample z0.

[0156] Represent the infrared-coupled editing diffusion generation model as Start denoising from the initial latent noise And generate a new image latent function conditioned on the overall feature c The training supervised mean squared error loss is:[[]]

[0157]

[0158] Where X is the true value image latent function, t is the diffusion time step, α t And σ t Are denoising hyperparameters, and c is the overall feature.

[0159] The overall feature is injected into each U-Net layer through cross-attention, and the detailed features are connected to the U-Net decoder features at each resolution. During training, the pre-trained parameters of the U-Net encoder are frozen to maintain the prior, and the U-Net decoder is adjusted to adapt to the current infrared target and background coupling editing task.

[0160] A training method for an infrared coupled editing diffusion generation model proposed in this embodiment encodes an image into global tokens and patch tokens during the process of extracting target overall feature information, and then connects these two types of tokens to retain more feature information. When extracting target detailed feature information, methods including but not limited to high-pass filtering, wavelet transform, discrete cosine transform, etc. are used to supplement the detailed information ignored during overall feature extraction. The pure infrared target after removing the background is extracted as a high-frequency mapping and tiled to a given position of the background image. At the same time, an information bottleneck is set to prevent too many appearance constraints during the tiling process, thereby enhancing the diversity of the model generation results. During the training process, in order to simulate the lack of fineness when users draw the shape of the infrared target, different ratios are selected to downsample the real infrared target mask input during training, and random dilation or erosion is applied to remove the details of the infrared target mask in the training data. Finally, the input infrared target and the infrared target mask are connected. The infrared coupled editing diffusion generation model trained in this way can regulate the target shape and posture, and achieve low-cost and high-efficiency coupled generation. Compared with the proposed technologies, it has stronger generalization and editability, and can more accurately and efficiently implement the target-background coupled editing training task of the infrared modality.

[0161] A training method for an infrared coupled editing diffusion generation model proposed in this embodiment, the infrared coupled editing diffusion generation model trained by this method generates coupled images at a speed of about 10 seconds per image. While achieving efficient generation, it solves the current problems of difficult infrared data acquisition, scarce data volume, unclear images, etc., realizes data augmentation of high-value targets under different infrared backgrounds, reduces the cost of artificial shooting or simulation of infrared images, and thus achieves the purpose of obtaining a large number of diverse infrared image coupled data with high efficiency and low cost.

[0162] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step, or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process are within the protection scope of this application.

[0163] Another embodiment of this application proposes a training system for an infrared coupled editing diffusion generation model. The details of the training system for an infrared coupled editing diffusion generation model proposed in this embodiment will be specifically described below. The following content is only implementation details provided for convenient understanding and is not necessary for implementing this example.

[0164] Figure 6It is a schematic structural diagram of a training system for an infrared coupled editing diffusion generation model proposed in this embodiment, including: a global feature extractor training module 201, a detailed feature extractor training module 202, a sample acquisition module 203, a global feature extraction module 204, a detailed feature extraction module 205, a model construction module 206, and a model training module 207.

[0165] The global feature extractor training module 201 is used to construct and train a global feature extractor.

[0166] The detailed feature extractor training module 202 is used to construct and train a detailed feature extractor.

[0167] The sample acquisition module 203 is used to acquire a sample pure infrared target image and a sample background image.

[0168] The global feature extraction module 204 is used to input the sample pure infrared target image into the global feature extractor, and encode the sample pure infrared target image into global tokens and patch tokens respectively in the form of tokens. After splicing the global tokens and patch tokens, align them into the latent space required for training the infrared coupled editing diffusion generation model to obtain the global features of the sample pure infrared target image.

[0169] The detailed feature extraction module 205 is used to process the sample pure infrared target image through high-pass filtering, wavelet transform, or discrete cosine transform to obtain a high-frequency mapping of the sample pure infrared target image. After fitting the high-frequency mapping to a specified position of the sample background image, input it into the pre-trained detailed feature extractor to obtain the detailed features of the sample pure infrared target image.

[0170] The model construction module 206 is used to construct an infrared coupled editing diffusion generation model based on U-Net.

[0171] The model training module 207 is used to inject the global features and detailed features into the infrared coupled editing diffusion generation model together to guide fine-tuning training, and obtain a trained infrared coupled editing diffusion generation model by gradually adding noise and gradually removing noise to the model.

[0172] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details and technical effects mentioned in the above method embodiment are still valid in this embodiment, and for the sake of reducing repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiment.

[0173] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0174] Another embodiment of this application proposes an electronic device, and its specific structure can be as Figure 7 shown, including: at least one processor 301; and a memory 302 communicatively connected to the at least one processor 301; wherein, the memory 302 stores instructions executable by the at least one processor 301, and the instructions are executed by the at least one processor 301 to enable the at least one processor 301 to execute a training method of an infrared coupling editing diffusion generation model as described in each of the above method embodiments.

[0175] Among them, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0176] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store data used by the processor when executing operations.

[0177] Another embodiment of this application proposes a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a training method of an infrared coupling editing diffusion generation model as described in the above method embodiments.

[0178] That is, those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disks, or optical discs.

[0179] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made to them in form and details without departing from the spirit and scope of the present application.

Claims

1. A training method for an infrared coupled editing diffusion generation model, characterized in that: include: Acquire a sample pure infrared target image and a sample background image; The sample pure infrared target image is input into the pre-trained overall feature extractor, and the sample pure infrared target image is encoded into global tokens and patch tokens respectively by tokenization. The global tokens and patch tokens are spliced ​​and aligned to the latent space required for the training of the infrared coupled edited diffusion generation model to obtain the overall features of the sample pure infrared target image. The sample pure infrared target image is processed by high-pass filtering, wavelet transform, or discrete cosine transform to obtain a high-frequency mapping image of the sample pure infrared target image, and the high-frequency mapping image is fitted to a specified position of the sample background image and then input into a pre-trained detail feature extractor to obtain detail features of the sample pure infrared target image; An infrared coupled editing diffusion generation model is constructed based on U-Net. The overall features and detail features are injected into the infrared coupled editing diffusion generation model to guide fine-tuning training. The trained infrared coupled editing diffusion generation model is obtained by gradually adding noise and denoising the model.

2. A training method for an infrared coupled editing diffusion generation model as claimed in claim 1, characterized in that: The infrared coupled edit diffusion generation model built based on U-Net is a symmetrical U-shaped fully convolutional neural network, which consists of a large number of convolutional layers, pooling layers, fully connected layers, and skip connections for supplementing lost semantic information. The infrared coupled edit diffusion generation model built based on U-Net mainly includes two parts: encoder and decoder. The encoder is responsible for further feature extraction and learning of the overall and detailed features of the input, increasing the robustness to disturbance noise, reducing the risk of overfitting, reducing the amount of computation, and increasing the receptive field; Each convolutional layer of the encoder consists of two 3×3 convolutions, each of which is followed by a linear rectifier unit and a 2×2 maximum pooling operation with a stride of 2. The maximum pooling operation is used to achieve downsampling, doubling the number of feature channels in each downsampling process; The decoder is responsible for restoring the feature map extracted by the encoder to the original resolution and fusing the shallow position information with the deep semantic information; Each step of the decoder consists of upsampling the features, then halving the number of channels of the features through a 2×2 convolution, concatenating them with the corresponding cropped feature maps in the contraction path, followed by two 3×3 convolutions, each of which is followed by a ReLU activation function layer, and finally using a 1×1 convolution to map each feature vector to the desired number of classes.

3. A method for training an infrared coupled editing diffusion generation model as claimed in claim 2, characterized in that: The work of the infrared coupled editing diffusion generation model mainly consists of two stages: the forward process and the backward process; For the initial sample x0∈q(x0), by The forward process is defined as a Markov chain whose stationary distribution is a Gaussian distribution, T is the end time, that is, the total number of time steps, It is a mathematical symbol used to describe the asymptotic behavior of a function, := indicates definition; In the forward process, random Gaussian noise is slowly added to the data sample, so the forward process is also called a diffusion process. The transition probability of each step is constructed by the following formula to add Gaussian noise: Among them, β t is the variance of the noise, β t ∈(0,1),β t >β t-1 , β t The values ​​taken over time step t are called variance tables, which are specified by default. As time step t increases, the noisy sample x t Getting closer to x T , more and more noise is added; Let β t =1-α t , through the following formula, q(x t |x t-1 ) is expanded to the form of the initial sample x0: Among them, α t is a function that decreases with time step t. As time step t increases, the data weight gradually decreases and the noise weight gradually increases. When t→∞, That is, when t is large enough, q(x t |x0) will converge to a standard Gaussian distribution that is independent of x0, that is, lim t→∞ q(x t )=0; Based on this, in the continuous-time model, the diffusion process is described as a transformation process that is continuous in time. Assuming that the termination time of the continuous-time diffusion process is 1, then t∈(0,1), t=0 and t=1 correspond to the two states of no noise and full noise respectively; Starting from the data sample x, the corresponding noisy sample of the diffusion process at any t∈(0,1) is z t , can be written as: Among them, α t and are all strictly positive scalar functions with respect to the time step t; In addition, the signal-to-noise ratio function is defined as SNR(t), SNR(t) is a strictly monotonically decreasing function at time step t, which visualizes the noisy sample z t The concept of becoming increasingly noisy as time step t passes; in, make Then we can get Right now Substituting into SNR(t) we get SNR(t) = exp[-γ(t)]; When γ(t)=log[expml(e -4 +10t 2 )], the discrete time diffusion model and the continuous time model can be linked together, and the discrete time model can be regarded as the discrete form of the continuous time model. The two can be transformed into each other.

4. A method for training an infrared coupled editing diffusion generation model as claimed in claim 3, characterized in that: The reverse process is the inverse process of the forward process, also known as the denoising process. It mainly denoises a random noise matrix step by step until an image is generated. The noise distribution predicted and removed at each step needs to be learned by the infrared coupled edited diffusion generation model during training. In the reverse process, according to the Bayesian formula q(z t-1 |z t )=[q(z t-1 |z t )′q(z t-1 )] / q(z t ), we can get the reverse process q(x t-1 |x t ), due to the marginal distribution q(z t ) and q(z t-1 ) is unknown, so q(z t-1 |z t ) is rewritten as: Due to the Markov chain property of the forward process, q(z t-1 |z t ,x) is simplified to: Substituting the simplified formula into the equation, we get: Among them, the mean is a function that depends on x and z t The function of; Since the x value is unknown, a neural network is designed to estimate the result of x. However, it is difficult to directly estimate the generated samples from the noise, so the reparameterization technique is used to obtain: Among them, ò∈N(0,I); Through simple transformation, we can get: Based on this, given a noisy sample z t , just need to start from z t By estimating the noise ò, we can calculate x, and then calculate the next step sample z t-1 In this process, x is used as an intermediate variable, so that the reverse process can get rid of the time step restriction in the discrete diffusion model and set the time step arbitrarily; For the continuous time model, the sample corresponding to any two time steps s and t satisfying 0≤s≤t≤1 in the reverse process is defined as z s and z t ,have: Further, p(z s |z t ) is rewritten as: c = -exp[γ(s) - γ(t)] γ ( t ) log [ exppml ( e -4 +10h 2 )] Among them, ε∈N(0,I).

5. The training method of the infrared coupled editing diffusion generation model as claimed in claim 4, characterized in that: Based on U-Net, an infrared coupled editing diffusion generation model is constructed. The overall features and detail features are injected into the infrared coupled editing diffusion generation model to guide fine-tuning training. By gradually adding noise and denoising the model, the trained infrared coupled editing diffusion generation model is obtained, including: First, we sample a training sample x∈q(x) from the data distribution q(x), then randomly generate a time t and calculate the variance and α t , then sample an ò∈N(0,I) from the standard Gaussian distribution, add noise to the training sample x, and generate z t And input into the infrared coupling editing diffusion generation model to get the predicted noise Finally calculated The loss between adding noise ò is used to update the gradient; In the sampling process of generating samples, the sampling step number is set to 1000 steps, and the step length of each sampling is 1 / 1000. First, a standard Gaussian noise ò, i.e. z1, is sampled from the noise space. Calculate the sample z after further denoising 0.999 , and so on, until z 0.001 Denoising to obtain the final desired generated sample z0; The infrared coupled edit diffusion generation model is expressed as Denoising starts from the initial latent noise ò and generates a new image latent function conditioned on the overall feature c The mean squared error loss for training supervision is: Among them, X is the true image latent function, t is the diffusion time step, α t and σ t is the denoising hyperparameter, c is the overall feature; The overall features are injected into each U-Net layer through cross-attention, and the detailed features are connected with the U-Net decoder features of each resolution. During training, the pre-trained parameters of the U-Net encoder are frozen to maintain the prior, and the U-Net decoder is adjusted to adapt it to the current infrared target and background coupled editing task.

6. A method for training an infrared coupled editing diffusion generation model according to any one of claims 1 to 5, characterized in that: The main body of the overall feature extractor is divided into the DINOv2 network. The training of the overall feature extractor adopts the visual self-supervised learning method, which is a knowledge distillation method for unlabeled training; In terms of self-supervised learning, DINOv2 uses a discriminative self-supervised method to learn features, which can be seen as a combination of SwAV-centered DINO loss and iBOT loss. A regularizer is added to propagate features and a short high-resolution training phase. In the image-level information, the training considers the cross entropy loss between the features extracted from the student network and the teacher network, both of which come from the class labels of ViT, obtained from different objects in the same image. The student class token is passed through the student DINO head, whose head is an MLP model that outputs a score vector called the student prototype score, and then Softmax is applied to obtain the probability distribution p of the student network. s , apply the teacher DINO head to the teacher class tokens to obtain the teacher prototype scores, and then apply Softmax to obtain the probability distribution p of the teacher network t ; The DINO loss is expressed by the formula: L DINO =-∑p t log(p s ) Among them, L DINO represents DINO loss; In the pixel-level information, the training randomly masks some input fragments to the student network, and then applies the student iBOT head to the student masked tokens. Similarly, the teacher iBOT head is applied to the visible teacher patch tokens, which correspond to the masked tokens in the student. Softmax is then applied to obtain the probability distribution p of the student network. si and the probability distribution p of the teacher network ti ; iBOT loss is expressed by the formula: L DINO =-∑ i p ti log(p si ) Where i is the patch index of the mask token; Both DINO loss and iBOT loss use a learnable MLP projection head, which is applied to the output token information and the loss is calculated, resulting in a pre-trained overall feature extractor.

7. A method for training an infrared coupled editing diffusion generation model as claimed in any one of claims 1 to 5, characterized in that: The sample pure infrared target image is processed by high-pass filtering to obtain the high-frequency mapping of the sample pure infrared target image, which is achieved by the following formula: Among them, L h and L v Respectively represent the horizontal high-pass Soble operator and the vertical high-pass Soble operator, acting as the high-pass filter convolution kernel, I represents the sample pure infrared target image, I gray Represents the grayscale image of the sample pure infrared target image, M erode is an erosion mask, which is used to filter out the feature information near the outer contour of the infrared target. ⊙ represents the Hadamard product operation. Represents a convolution operation.

8. A training system for infrared coupled editing diffusion generation model, characterized in that: include:. The overall feature extractor training module is used to build and train the overall feature extractor; Detail feature extractor training module, used to build and train detail feature extractors A sample acquisition module is used to acquire a sample pure infrared target image and a sample background image; The overall feature extraction module is used to input the sample pure infrared target image into the overall feature extractor, encode the sample pure infrared target image into global tokens and patch tokens respectively by tokens, splice the global tokens and patch tokens and align them to the latent space required for the training of the infrared coupled edit diffusion generation model, and obtain the overall features of the sample pure infrared target image; A detail feature extraction module is used to process the sample pure infrared target image through high-pass filtering, wavelet transform, or discrete cosine transform to obtain a high-frequency mapping image of the sample pure infrared target image, fit the high-frequency mapping image to a specified position of the sample background image, and then input it into a pre-trained detail feature extractor to obtain detail features of the sample pure infrared target image; Model building module, used to build an infrared coupled editing diffusion generation model based on U-Net; The model training module is used to inject overall features and detail features into the infrared coupled editing diffusion generation model to guide fine-tuning training. The trained infrared coupled editing diffusion generation model is obtained by gradually adding noise and denoising the model.

9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a training method for an infrared coupled editing diffusion generation model as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement a training method for an infrared coupling editing diffusion generation model as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Medical image segmentation method and system based on diffusion difference learning model

    CN121811411A