Text condition guided image external expansion method based on diffusion model and terminal
Through the combination of the multimodal large language model and dual UNet network, an external expansion image is generated, which solves the problem of out-expansion of any pixel point in the image in the prior art and semantic inconsistency, and achieves a reasonable and beautiful external expansion effect.
Patent Information
- Application Number
- CN202510741426.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
In the prior art, the method based on the generation of adversarial networks cannot perform out-expanding of any pixel point in the image, and the denoising process of the diffusion model is less guided by the original sub-graph information, resulting in the semantics of the out-expanding content and the generation results are not beautiful enough.
The multimodal large language model is used to generate outscaling text conditions, and combined with frozen and trainable dual UNet networks, image features and text features are processed through data augmentation and zero convolutional layers to generate outscaling images.
The external expansion of any pixel point in the image is realized. The external expansion content conforms to the semantic logic, improves the rationality and aesthetics of the generated image, enhances the semantic coherence between the external expansion content and the original image, and breaks through the limitations of the external expansion range of the generative adversarial network.
Smart Images

Figure CN120259113A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model-based image extrapolation, and particularly relates to an image extrapolation method and terminal guided by text conditions based on a diffusion model. Background Art
[0002] Currently, two main technical routes are adopted in the field of image extrapolation: one is the method based on generative adversarial networks, which constructs an adversarial training framework of a generator and a discriminator, enabling the generator to learn the distribution characteristics of the boundary region of the input image to generate a visually coherent extended region; the other is the method based on diffusion models, which maps the sub-image region of the input image to the latent space through an encoder, gradually denoises using the diffusion process, and reconstructs the latent space features into an extrapolated image through a decoder.
[0003] However: (1) The method based on generative adversarial networks cannot perform extrapolation of arbitrary pixel points of an image; (2) The denoising process of the diffusion model is less guided by the original sub-image information, and the semantic connection between the extrapolated content and the original content is not coherent; (3) The generated extrapolated content lacks rationality, and the final generated result is not beautiful enough. Summary of the Invention
[0004] The technical problem to be solved by the present invention is: to provide an image extrapolation method and terminal guided by text conditions based on a diffusion model, which support extrapolation of arbitrary pixel points, and are reasonable, beautiful, and semantically coherent.
[0005] To solve the above technical problem, the technical solution adopted by the present invention is: An image extrapolation method guided by text conditions based on a diffusion model, comprising the steps of: S1. Receive the original image input by the user, and for the original image, generate extrapolation text conditions using a pre-trained multi-modal large language model; S2. Perform feature encoding on the original image to generate image features, and perform feature encoding on the extrapolation text conditions to generate text features; S3. Input the image features and the text features into a pre-trained latent diffusion model based on a dual UNet network, and generate an extrapolated image based on the latent diffusion model.
[0006] To solve the above technical problem, the technical solution adopted by the present invention is: An image extrapolation terminal guided by text conditions based on a diffusion model, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the above-mentioned image extrapolation method guided by text conditions based on a diffusion model are implemented.
[0007] The beneficial effects of the present invention are as follows: An image out-expansion method and terminal based on a diffusion model with text-conditioned guidance according to the present invention introduce a multi-modal large language model to generate text conditions, replacing the discarded text guidance in the traditional diffusion model, making the out-expanded content conform to semantic logic and improving rationality and aesthetics; the dual UNet structure processes text semantics and original image features in separate modules, avoiding the overburden of cross-attention in a single UNet and enhancing the semantic coherence between the out-expanded content and the original image; through data augmentation and the dual UNet architecture, it supports out-expansion of any pixel in the image, breaking through the out-expansion range limitation of the generative adversarial network. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 It is a schematic diagram of a brief process example of an image out-expansion method based on a diffusion model with text-conditioned guidance according to an embodiment of the present invention; Figure 2 It is a schematic diagram of the overall method process architecture example of an image out-expansion method based on a diffusion model with text-conditioned guidance according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the structure of a mixture-of-experts module of an image out-expansion method based on a diffusion model with text-conditioned guidance according to an embodiment of the present invention; Figure 4 It is an interaction structure diagram of a Q-former module of an image out-expansion method based on a diffusion model with text-conditioned guidance according to an embodiment of the present invention; Figure 5 It is a training connection diagram of a Q-former and an LLM decoder of an image out-expansion method based on a diffusion model with text-conditioned guidance according to an embodiment of the present invention; Figure 6 It is a schematic diagram of the structure of an image out-expansion terminal based on a diffusion model with text-conditioned guidance according to an embodiment of the present invention; LABEL DESCRIPTION: 1. An image out-expansion terminal based on a diffusion model with text-conditioned guidance; 2. A processor; 3. A memory. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0009] To describe in detail the technical content, achieved objectives and effects of the present invention, the following is described in conjunction with the embodiments and with reference to the accompanying drawings.
[0010] Please refer to Figures 1 to 5 , an image out-expansion method based on a diffusion model with text-conditioned guidance, comprising the steps of: S1. Receive the original image input by the user, and for the original image, use a pre-trained multi-modal large language model to generate out-expansion text conditions; S2. Perform feature encoding on the original image to generate image features, and perform feature encoding on the out-expansion text conditions to generate text features; S3. Input the image features and the text features into a pre-trained latent diffusion model based on a dual UNet network, and generate an expanded image based on the latent diffusion model.
[0011] As can be seen from the above description, the beneficial effects of the present invention are as follows: A text-conditioned image expansion method and terminal based on a diffusion model according to the present invention introduce a multi-modal large language model to generate text conditions, replacing the discarded text guidance in the traditional diffusion model, making the expanded content conform to semantic logic, and improving rationality and aesthetics; The dual UNet structure processes text semantics and original image features in separate modules, avoiding the excessive burden of cross-attention in a single UNet, and enhancing the semantic coherence between the expanded content and the original image; Through data augmentation and the dual UNet architecture, it supports the expansion of any pixel of the image, breaking through the expansion range limitation of the generative adversarial network.
[0012] Further, for the original image, generating the expanded text condition using a pre-trained multi-modal large language model includes the steps of: Input the original image into the multi-modal large language model, extract image features through the CLIP visual encoder therein, and generate the expanded text condition through the mixture of experts module and the Q-former module.
[0013] As can be seen from the above description, CLIP extracts local patch features of the image, ConvNeXt extracts global semantics, and the mixture of experts module dynamically weights and fuses local and global information, enabling the model to understand the overall structure and local details of the image; The Q-former strengthens the cross-modal alignment between image features and text through image-text contrast, generation, and pairing tasks, ensuring that the generated expanded text condition accurately matches the semantics of the original image and improving the guidance accuracy of the subsequent diffusion model.
[0014] Further, the latent diffusion model adopts a dual UNet network, including a frozen UNet network with frozen original parameters and a pre-trained trainable UNet network; Step S3 includes the steps of: Inject the text features into the cross-attention module of the frozen UNet network, input the image features into the cross-attention module of the trainable UNet network, and combine the features of the two UNets through a zero convolution layer, and gradually denoise to generate an expanded latent variable; Decode the expanded latent variable into an expanded image through a decoder.
[0015] As described above, the frozen UNet utilizes the pre-trained text-to-image generation ability and focuses on the direction of text semantic-guided external expansion; the trainable UNet combines the original image features (such as edges and colors) to constrain the visual coherence of the externally expanded content; the zero convolutional layer dynamically adjusts the interaction intensity between the two UNets to balance semantic guidance and content constraint, avoiding the externally expanded area deviating from the original image style or semantic breakage.
[0016] Furthermore, the construction of the training data for the multimodal large language model includes the steps of: Obtain a certain number of image data, generate corresponding text descriptions for each image through an open-source model, and input them into a text embedding model to generate text embedding features, obtaining image-text pairs; For each image-text pair, use data augmentation processing to generate enhanced image-text pairs, obtaining a training dataset; The choices of the data augmentation processing include instance cropping of the image data, random cropping at a preset ratio, and random horizontal or vertical mixing of the image-text pairs.
[0017] As described above, instance cropping forces the model to learn the external expansion logic of "local instance → global scene" (such as cropping a sheep's head to generate a complete sheep); random cropping at a preset ratio forces the model to infer large areas of missing content, improving the flexibility of the external expansion range; image-text hybrid augmentation simulates the requirements of multi-scene fusion and enhances the model's ability to process complex external expansion semantics (such as splicing images of "sheep" and "grassland" and generating mixed text descriptions).
[0018] Furthermore, the training of the multimodal large language model includes the steps of: For each image-text pair after data augmentation processing, input the image data therein into a CLIP visual encoder to perform image encoding, obtaining image embedding features, and extract global features from the original image data through a convolutional neural network; Input the image embedding features and the global features into a mixture of experts module to generate new image embedding features; Input the image embedding features processed by the mixture of experts module and the text embedding features into a Q-former module, and perform the first-stage training through three tasks: cross-modal contrastive learning, text generation, and image-text pairing; Link a layer of MLP to the vector output by the Q-former module and input it into the decoder of the large language model, calculate the cross-entropy loss of the context window, and fine-tune the parameters of the Q-former model.
[0019] As described above, the mixture-of-experts module dynamically assigns feature weights through K MLP expert networks to adapt to the feature fusion requirements under different cropping ratios (e.g., enhancing the global feature weight when cropping at a large ratio); the Q-former optimizes image-text alignment in the first stage and calibrates text logic through the large language model decoder in the second stage, making the generated extended text conform to the image semantics and possess natural language fluency.
[0020] Furthermore, the loss parameters of the Q-former model are: ; where , ,..., is the given text token sequence, and θ is the parameter of the model.
[0021] As described above, the cross-entropy loss based on the context window forces the text token sequence generated by the Q-former to conform to the language prior knowledge (such as grammar and semantic coherence), avoiding logical breaks or semantic contradictions in the extended text and improving the quality of the text conditions.
[0022] Furthermore, the mixture-of-experts module uses K multi-layer perceptrons as expert networks : ; where represents the i-th new image embedding feature output by the mixture-of-experts module, represents the i-th image embedding feature generated by the CLIP visual encoder, and Cat represents the concatenation operation of tensors, which is used for calculating the average weight of the expert networks: ; where is a multi-layer perceptron that converts the tensor from the feature dimension to the K dimension.
[0023] As described above, the gating network dynamically calculates the expert weights, enabling the model to adaptively select experts according to the input features (such as the relevance between image patches and global semantics), improving the flexibility and efficiency of feature fusion; sparsely activating some experts (non-full connection), while reducing the computational complexity, enhancing the model's ability to cluster multi-modal features.
[0024] Furthermore, the construction of the latent diffusion model includes: Obtain an open-source UNet network, generate a copy of the open-source UNet network, freeze the parameters of the open-source UNet network to serve as a frozen UNet network, and use the copy of the open-source UNet network for training as a trainable UNet network; The frozen UNet network and the trainable UNet network are connected by zero convolution and use different inputs. The frozen UNet network inputs the latent noise vector z t , and the trainable UNet network inputs the latent noise vector z t , the latent representation z of the masked image m , and the downsampled binary mask m; The calculation formula for the interaction between each layer of the frozen UNet network and the trainable UNet network is: ; ; where represents the output feature of the th layer of the latent diffusion model, and represent the output features of the nth layer of the frozen UNet network and the trainable UNet network respectively, represents the number of steps of adding noise in the diffusion process, represents the zero convolution layer, is a parameter that adjusts the interaction strength between the two networks, represents the encoder in the variational autoencoder VAE, represents the masked image obtained by padding zeros around the original image to get the size of the target extended image, E I represents the image feature generated by encoding the features of the original image, E T represents the text feature generated by encoding the features of the extended text condition.
[0025] As can be seen from the above description, the frozen UNet retains the pre-trained generation ability. The trainable UNet clearly distinguishes the original image from the extended area through the masked latent representation and the binary mask, guiding the diffusion process to preferentially repair the extended area; the zero convolution layer linearly combines the features of the two UNets to achieve a dual denoising mechanism of "semantic guidance (frozen UNet) + structural constraint (trainable UNet)", improving the rationality of the generated content.
[0026] Furthermore, the latent diffusion model uses the noise loss function of the diffusion model as the loss function: ; where represents the output feature of the th layer of the latent diffusion model, t represents the number of steps of adding noise in the diffusion process, is the random noise sampled from the standard normal distribution, z t is the latent noise vector obtained by adding noise at the tth step, Denotes taking the mathematical expectation.
[0027] As can be seen from the above description, by minimizing the L2 distance between the predicted noise and the true noise, the dual UNet is forced to accurately model the inverse distribution of the diffusion process, ensuring the high fidelity of the extrapolated latent variables; the optimization objective of the noise loss is consistent with the physical meaning of the diffusion model, ensuring that the generated extrapolated images are visually coherent, natural, and without obvious noise artifacts.
[0028] A method and terminal for text-conditioned guided image extrapolation based on a diffusion model according to the present invention are applicable to semantic-guided extrapolation based on a diffusion model in the field of image editing.
[0029] Please refer to Figures 1 to 5 , Example 1 of the present invention is: A method for text-conditioned guided image extrapolation based on a diffusion model, comprising the steps of: S1. Receive the original image input by the user, and for the original image, use a pre-trained multimodal large language model to generate an extrapolation text condition; For the original image, generating an extrapolation text condition using a pre-trained multimodal large language model includes the steps of: Input the original image into the multimodal large language model, extract image features through the CLIP visual encoder therein, and generate an extrapolation text condition through a mixture of experts module and a Q-former module.
[0030] The construction of the training data of the multimodal large language model includes the steps of: Obtain a certain number of image data, generate corresponding text descriptions for each image through an open-source model, and input them into a text embedding model to generate text embedding features, obtaining image-text pairs; For each image-text pair, use data augmentation processing to generate enhanced image-text pairs, obtaining a training data set; The selection of the data augmentation processing includes instance cropping of the image data, random cropping at a preset ratio, and random horizontal or vertical mixing of the image-text pairs.
[0031] In this embodiment, first, a large number of color images I G are collected, and corresponding text descriptions T G are generated for each image through an existing open-source model (other open-source multimodal large language models for describing image content), and input into a text embedding model to obtain text embeddings E G . Instance cropping means using a detection model to anchor all instances in the image, randomly selecting the bounding box of an instance to crop a sub-image to obtain I crop , and random cropping at a preset ratio means significantly reducing the cropped sub-image I crop and the original image IG The ratio is such that only 20%-50% of the original image is retained. The text embedding E of the corresponding images under the two cropping schemes remains unchanged. The random horizontal or vertical mixing of image-text pairs means that two images are spliced together horizontally or vertically, and their corresponding text embeddings are mixed proportionally. During the training process, one of the three enhancement methods is randomly selected for the original image to obtain the corresponding image after hybrid data enhancement. G The random horizontal or vertical mixing of image-text pairs means that two images are spliced together horizontally or vertically, and their corresponding text embeddings are mixed proportionally. During the training process, one of the three enhancement methods is randomly selected for the original image to obtain the corresponding image after hybrid data enhancement. and the text embedding E G are paired and input into the multi-modal large language model for training.
[0032] The training of the multi-modal large language model includes the steps of: For each image-text pair after data enhancement processing, the image data therein is input into the CLIP visual encoder to perform image encoding to obtain image embedding features, and the original image data is extracted through a convolutional neural network to obtain global features.
[0033] The image embedding features and the global features are input into the mixture of experts module to generate new image embedding features.
[0034] In this embodiment, the image obtained through data enhancement is input into the CLIP visual encoder to perform image encoding, and then the obtained image embedding feature V patch , and the global feature V obtained by the convolutional neural network ConvNeXt extracting the image I G are input into the mixture of experts module together to obtain a new image embedding global , as shown. The specific calculation formula is as follows: Figure 3 shown. The specific calculation formula is as follows: ; ; The mixture of experts module uses K multi-layer perceptrons (MLPs) as expert networks : ; where, represents the i-th new image embedding feature output by the mixture of experts module, represents the i-th image embedding feature generated by the CLIP visual encoder, Cat represents the concatenation operation of tensors, which is used for the weight calculation of the average expert network: ; where, is a multi-layer perceptron that converts the tensor from the feature dimension to the K dimension.
[0035] The role of Q-former is to pass the learnable vector VQ Link the image features and text features, with the structure as Figure 4 shown. Input the image embedding features obtained after being processed by the mixture-of-experts module and the text embedding feature E G into the Q-former module, and conduct the first-stage training through three tasks: image-text contrast learning, text generation, and image-text pairing; As Figure 5 shown, link the vector V Q output by the Q-former module to a layer of MLP and then input it into the decoder of the large language model, calculate the cross-entropy loss of the context window, and fine-tune the parameters of the Q-former model.
[0036] The loss parameter of the Q-former model is: ; where , ,..., is the given text token sequence, and θ is the parameter of the model.
[0037] S2. Perform feature encoding on the original image to generate image features, and perform feature encoding on the extended text condition to generate text features; S3. Input the image features and the text features into a pre-trained latent diffusion model based on a dual UNet network, and generate an extended image based on the latent diffusion model; The latent diffusion model adopts a dual UNet network, including a frozen UNet network with frozen original parameters and a pre-trained trainable UNet network; Step S3 includes the steps of: Inject the text features into the cross-attention module of the frozen UNet network, input the image features into the cross-attention module of the trainable UNet network, and combine the features of the two UNets through a zero convolution layer to gradually denoise and generate an extended latent variable; Decode the extended latent variable into an extended image through a decoder.
[0038] In this embodiment, the text features and image features are respectively injected into different UNets. The aim is to fully utilize the original subgraph while receiving the text condition guidance of the multi-modal large language model, so as to generate more reasonable extended content and be more aligned with the semantics of the original image. Specifically, the text condition C T is passed through the CLIP text encoder to obtain the text feature E T , and is injected into the cross-attention module of each layer of the frozen UNet. The original image I G is passed through the image encoder Obtain the image feature E I , after aligning the tensor dimension with the latent space feature dimension d through a multi-layer perceptron, inject it into the cross-attention module of each layer of the trainable UNet, as Figure 2 shown. It can be expressed as the following calculation: ; ; ; ; ; ; Here, W Q , W K , W V are learnable linear layers, , are the outputs of the cross-attention layer of the frozen UNet network with frozen weights and the output of the cross-attention layer of the trainable UNet network respectively.
[0039] The construction of the said latent diffusion model includes: Obtain an open-source UNet network, generate a copy for the said open-source UNet network, freeze the parameters of the said open-source UNet network as the frozen UNet network, and use the copy of the said open-source UNet network for training as the trainable UNet network.
[0040] In this embodiment, as Figure 2 shown, use the UNet network in the pre-trained latent diffusion model , this architecture includes residual blocks, self-attention layers and cross-attention layers, and generate a copy of it. The two UNets are connected through zero convolutions, and only the parameters in the copy UNet network are adjusted during the training process, and the parameters of the original UNet are frozen.
[0041] The frozen UNet network and the trainable UNet network are connected through zero convolutions and use different inputs. The frozen UNet network inputs the latent noise vector z t , and the trainable UNet network inputs the latent noise vector z t , the latent representation z m of the masked image, and the downsampled binary mask m; The calculation formula for the interaction between each layer of the frozen UNet network and the trainable UNet network is: ; ; Among them, represents the output feature of the -th layer of the said potential diffusion model, and respectively represent the output features of the n-th layer of the frozen UNet network and the trainable UNet network, represents the number of steps of adding noise in the diffusion process, Z represents the zero convolution layer, is a parameter for adjusting the interaction intensity between the two networks, represents the encoder in the variational autoencoder VAE, represents a mask image obtained by padding zeros around the original image to get the size of the target extended image, E I represents the image feature generated by encoding the features of the said original image, E T represents the text feature generated by encoding the features of the said extended text condition.
[0042] The said potential diffusion model uses the noise loss function of the diffusion model as the loss function: ; Among them, represents the output feature of the -th layer of the said potential diffusion model, t represents the number of steps of adding noise in the diffusion process, is the random noise sampled from the standard normal distribution, z t is the latent noise vector obtained by adding noise at the t-th step, represents taking the mathematical expectation.
[0043] In the forward process of the diffusion model, Gaussian noise is sampled and added to the data sample z0, so as to obtain a noisy sample z at the time step t t . In the reverse process, the model needs to predict the noise intensity at the current time t.
[0044] In this embodiment, the training strategy: Use BLIP2 to generate text conditions for training images. Use the UNet from StableDiffusion 1.5 as the pre-trained model. The entire potential diffusion model was trained for 200,000 iterations on 8 NVIDIA GeForce RTX4090 GPUs with a learning rate of 1e-5. Apply the gradient accumulation strategy in the Accelerate package to increase the equivalent batch size, and use the 8-bit Adam optimizer to reduce GPU memory usage. In order to enable the model to handle text-free condition generation, randomly discard the text conditions with a probability of 20% during training.
[0045] Please refer to Figure 6 , Embodiment 2 of the present invention is: An image expansion terminal 1 guided by text conditions based on a diffusion model, comprising a processor 2, a memory 3, and a computer program stored in the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, the steps in the above-mentioned image expansion method guided by text conditions based on a diffusion model are implemented.
[0046] In summary, the image expansion method and terminal guided by text conditions based on a diffusion model provided by the present invention introduce a multi-modal large language model to generate text conditions, replacing the discarded text guidance in the traditional diffusion model, making the expanded content conform to semantic logic, and improving rationality and aesthetics; the dual UNet structure processes text semantics and original image features in separate modules, avoiding the overburden of cross-attention in a single UNet, and enhancing the semantic coherence between the expanded content and the original image; through data augmentation and the dual UNet architecture, it supports the expansion of any pixel of the image, breaking through the expansion range limitation of the generative adversarial network.
[0047] In view of the text guidance characteristics of the pre-trained text-to-image diffusion model, the present invention introduces a multi-modal large language model to generate text conditions for the image expansion task to generate more reasonable and beautiful expanded images. A dual UNet diffusion model network is constructed, and text conditions and image features are respectively injected into different UNets, so that the denoising process of the diffusion model incorporates the guidance of the original sub-image information to generate more semantically continuous expanded content.
[0048] Addition of text conditions. Currently, most of the external painting methods based on diffusion models utilize pre-trained text-to-image models. However, these methods discard text conditions during the generation process, resulting in a decline in the performance of the text-to-image model. In contrast, the present invention uses a dedicated multi-modal large language model to generate reasonable image expansion text conditions for the text-to-image model, enabling it to achieve optimal generation performance and thus produce more reasonable and beautiful image expansion results.
[0049] Guidance of original image information. Understanding the internal pattern of the original image is a key factor in the image expansion task. By constructing a diffusion model with a dual UNet network and injecting the original image features into the cross-attention layer of an additional UNet, the denoising process of the diffusion model receives the original image information, reducing the linguistic incoherence between the original image and the generated expanded content. At the same time, by processing text conditions and original image features through two networks respectively, the burden on the cross-attention layer of a single UNet network for processing multi-modal data is reduced.
[0050] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent transformation made using the specification and drawings of the present invention, or directly or indirectly applied in related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for text-conditioned guided image extrapolation based on a diffusion model, characterized in that, Including the steps: S1. Receive the original image input by the user, and for the original image, use a pre-trained multi-modal large language model to generate an extended text condition; S2. Perform feature encoding on the original image to generate image features, and perform feature encoding on the extended text condition to generate text features; S3. Input the image features and the text features into a pre-trained latent diffusion model based on a dual UNet network, and generate an extended image based on the latent diffusion model.
2. The method for image extrapolation guided by text conditions based on a diffusion model according to claim 1, wherein For the original image, using a pre-trained multi-modal large language model to generate an extended text condition includes the steps: Input the original image into the multi-modal large language model, extract image features through the CLIP visual encoder therein, and generate an extended text condition through the mixture of experts module and the Q-former module.
3. The method for expanding an image guided by text conditions based on a diffusion model according to claim 1, wherein The latent diffusion model adopts a dual UNet network, including a frozen UNet network with frozen original parameters and a pre-trained trainable UNet network; Step S3 includes the steps: Inject the text features into the cross-attention module of the frozen UNet network, input the image features into the cross-attention module of the trainable UNet network, and combine the features of the two UNets through a zero convolution layer to gradually denoise and generate an extended latent variable; Decode the extended latent variable into an extended image through a decoder.
4. A method for text-conditioned image extrapolation based on a diffusion model according to claim 1, characterized in that, The construction of the training data of the multi-modal large language model includes the steps: Obtain a certain number of image data, generate corresponding text descriptions for each image through an open-source model, and input them into a text embedding model to generate text embedding features, obtaining image-text pairs; For each image-text pair, use data augmentation processing to generate enhanced image-text pairs, obtaining a training data set; The selection of the data augmentation processing includes instance cropping of the image data, random cropping at a preset ratio, and random horizontal or vertical mixing of the image-text pairs.
5. A method for text-conditioned image extrapolation based on a diffusion model according to claim 4, characterized in that, The training of the multi-modal large language model includes the steps: For each image-text pair after data augmentation processing, input the image data therein into the CLIP visual encoder to perform image encoding, obtaining image embedding features, and extract global features from the original image data through a convolutional neural network; Input the image embedding features and the global features into the mixture of experts module to generate new image embedding features; Input the image embedding features obtained after being processed by the mixture of experts module and the text embedding features into the Q-former module, and perform the first-stage training through three tasks: image-text contrast learning, text generation, and image-text pairing; Link a layer of MLP to the vector output by the Q-former module and input it into the decoder of the large language model, calculate the cross-entropy loss of the context window, and fine-tune the parameters of the Q-former model.
6. The method for image out-expansion guided by text conditions based on a diffusion model according to claim 5, wherein The loss parameter of the Q-former model is: ; where , ,..., is a given sequence of text tokens, and θ is the parameter of the model.
7. A method for text-conditioned image extrapolation based on a diffusion model according to claim 5, characterized in that, The mixture-of-experts module uses K multi-layer perceptrons as the expert networks : ; Among them, represents the i-th new image embedding feature output by the mixture of experts module, represents the i-th image embedding feature generated by the CLIP visual encoder, and Cat represents the concatenation operation of tensors, which is used for the weight calculation of the average expert network: ; Among them, is a multi-layer perceptron that converts a tensor from the feature dimension to the K dimension.
8. A method for text-conditioned guided image extrapolation based on a diffusion model according to claim 1, characterized in that The construction of the latent diffusion model includes: Obtain an open-source UNet network, generate a copy for the open-source UNet network, freeze the parameters of the open-source UNet network as the frozen UNet network, and use the copy of the open-source UNet network to receive training as the trainable UNet network; The frozen UNet network and the trainable UNet network are connected by zero convolution and use different inputs. The frozen UNet network inputs the latent noise vector z t , and the trainable UNet network inputs the latent noise vector z t , the latent representation z of the masked image m and the downsampled binary mask m; The calculation formula for the interaction between each layer of the frozen UNet network and the trainable UNet network is as follows: ; ; Among them, represents the output feature of the th layer of the potential diffusion model, and respectively represent the output features of the nth layer of the frozen UNet network and the trainable UNet network, represents the number of steps for adding noise in the diffusion process, represents the zero convolution layer, is a parameter for adjusting the interaction intensity between the two networks, represents the encoder in the variational autoencoder VAE, represents a mask image obtained by padding zeros around the original image to get the size of the target extended image, E I represents the image features generated by encoding the features of the original image, E T represents the text features generated by encoding the features of the extended text condition.
9. A method for text-conditioned guided image outpainting based on a diffusion model according to claim 8, wherein The latent diffusion model uses the noise loss function of the diffusion model as the loss function: ; Among them, represents the output feature of the th layer of the said potential diffusion model, t represents the number of steps of adding noise in the diffusion process, is the random noise sampled from the standard normal distribution, z t is the latent noise vector obtained by adding noise at the t-th step, represents taking the mathematical expectation.
10. An image extrapolation terminal with text-conditioned guidance based on a diffusion model, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps in any one of the above claims 1-9 for an image out-expansion method guided by text conditions based on the diffusion model.
Citation Information
Patent Citations
Image anomaly detection method based on character-to-image diffusion model
CN117218413A
Method for generating continuous pictures by long text based on diffusion model
CN117521672A
Method for realizing phrase-level positioning by using text-to-image diffusion model
CN118247799A
Real world image super-resolution method based on stable diffusion
CN118918009A
Dream decoding method based on functional magnetic resonance imaging signal
CN119006617A
Cited By
Building image generation method and device, and medium
CN120807715A
Intelligent clothing pattern generation method and system based on designer acceptance
CN120823280A
Image boundary filling method and medical image analysis method
CN122265330A
Image boundary filling methods and medical image analysis methods
CN122265330B