Cross-modal text image generation method based on distribution regularization

Through the combination of BiLSTM and VAE, the slow convergence problem caused by complex distribution of generated images is solved, faster model convergence and image generation that is more in line with text description are achieved, and the diversity and quality of generated images are improved.

CN120355803APending Publication Date: 2025-07-22XIAN TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510408998.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art ignores the image diversity problem when generating images, resulting in the generated image distribution being too complex and the model convergence speed is slow.

Method used

A cross-modal text generation method based on distribution regularization is adopted, text encoding is performed through BiLSTM, a variational autoencoder VAE is introduced for regularization of potential embedded vectors, and combined training is carried out for combined training in combination with the generative adversarial network. The image reconstruction loss is measured using the L1 paradigm and KL divergence, and the cross entropy is used as the discriminant loss, and the hyperparameter λ2 is adjusted to avoid overfitting.

Benefits of technology

The convergence speed of the model and the diversity and quality of the generated images are improved, and the generated images are more in line with text description and have enhanced discrimination ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355803A_ABST
    Figure CN120355803A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal text image generation method based on distribution regularization. Firstly, a text encoder encodes texts of data sets such as COCO to obtain word feature vectors and global sentence feature vectors; secondly, generating images with different resolutions from the feature vectors through a three-stage generator; and thirdly, a variational auto-encoder is introduced into a discriminator module, distribution regularization is carried out on the generated image, and a discriminator carries out authenticity judgment based on the encoded image. Then, the real image and the generated image are used as input to calculate the loss of the discriminator, and the model is optimized through multiple iterations; and finally, evaluating the trained optimal image model by using IS and FID indexes, and measuring the quality of the generated image and the performance of the model. Through experimental verification, the method can effectively generate the corresponding image based on the semantics of the text, and effectively solves the problem that the discrimination model is difficult to distinguish the authenticity of the input image. The performance of the obtained index data is superior to that of the original model AttnGAN.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and natural language processing, and is a cross-modal text-to-image generation method based on distribution regularization. Background Art

[0002] In the field of image processing, cross-modal technology significantly improves the ability of image processing and broadens its application scope by fusing information from images and other modalities (such as text, audio, video, sensor data, etc.). This technology is mainly used in the fields of image annotation and description generation, cross-modal retrieval, image generation and editing. AttnGAN is a generative adversarial network based on the attention mechanism, which is specially used to generate high-quality images, especially to generate images based on text descriptions. By introducing the attention mechanism, it can capture the relationship between text and image more finely, thereby generating images that are more consistent with the text description.

[0003] The patent with application number "CN202111183091.3" discloses "a text-generated image learning model based on deep learning". This model refers to the architecture in AttnGAN, uses the joint attention stacking generation module to extract detailed feature information from word-level information, and generates images through the global sentence attention mechanism; in addition, the reverse correction and correction of the text generation module improves the quality of the initial image by reversely generating and matching word-level feature vectors through text descriptions. The patent with application number "CN202111271415.9" discloses "a cross-modal text-to-image generation method based on adversarial generative network", which improves the effect of optimizing defective images by introducing adversarial learning in the regeneration module; in addition, semantic distance metric optimization is used to ensure the semantic consistency between image pairs, and the generated images have better semantic consistency performance. However, the essence of the generative adversarial model is to learn and fit the distribution of the target image, because the diversity of visual information in the image makes the image distribution complex. In the above two solutions, the diversity of generated images is ignored, and the distribution of directly generated images is too complex. Therefore, the discriminator’s ability to discriminate image distribution is slightly inferior, resulting in a slower model convergence speed. Summary of the invention

[0004] The present invention provides a cross-modal text image generation method based on distribution regularization, which overcomes the problem of ignoring the diversity of generated images in the prior art, resulting in the problem that the distribution of directly generated images is too complex and the model convergence speed is slow.

[0005] In order to achieve the purpose of the present invention, the technical solution provided by the present invention is: a method for generating images from cross-modal text based on distribution regularization, and the specific steps are as follows:

[0006] Step 1: The text encoder preprocesses the text corresponding to the public dataset through a bidirectional long short-term memory recurrent network (BiLSTM). The preprocessing refers to encoding the text to obtain word feature vectors and global sentence feature vectors.

[0007] Step 2: The feature vectors and randomly sampled noise z are sent as inputs to the generator module. After being processed and transformed through upsampling convolutional layers and residual layers, a series of hidden layer features h i (i = 0, 1, 2) are generated. Then, different resolution images are generated by three-stage generators respectively, and finally an image x corresponding to the text is generated i (i = 0, 1, 2);

[0008] Step 3: The image x given by the generator module i is first fed to the encoder to obtain a latent embedding vector v. A variational autoencoder (VAE) is introduced in the discriminator module. The discriminator is divided into two sub-modules: the discriminator module and the variational autoencoder module. The discriminator module consists of a discriminator encoder and a logical classifier. The logical classifier combines text information to determine whether the image is generated by the generator or belongs to the original dataset. An image generated by the generator that conforms to the semantics is determined to be a real image. The variational autoencoder module consists of a variational encoder, a variational sampling module, and a decoder. The obtained latent embedding vector v is regularized through the variational sampling module to obtain z * , and then combined with the sentence vector and sent to the decoder to obtain the generated image;

[0009] Step 4: The real image and the generated image are used as inputs to the distribution regularization loss function, and jointly trained with the adversarial loss of the discriminator generative adversarial network. The parameters are updated backward to obtain an optimal parameter model.

[0010] Furthermore, in the above Step 3, the process of regularizing z through the variational sampling module * is as follows:

[0011] First, a Gaussian distribution is constructed

[0012] Then, through operations, the normalization of the latent distribution is achieved,

[0013] Finally, feature resampling is adopted to obtain

[0014] Furthermore, in the above Step 4, the loss function in the generative adversarial network is divided into two parts: the conditional loss function and the distribution regularization loss function. Among them, the conditional loss function has two types: unconditional loss and conditional loss.

[0015] Further, in the above step 4, based on the design of the variational lower bound of the variational autoencoder network, the distribution regularization module DRM i in the variational autoencoder network has the following loss function

[0016]

[0017] The loss of the variational autoencoder (VAE) uses the L1 norm to measure the distance between the input image and the image after encoding, transformation, and decoding; the image reconstruction loss uses the KL divergence to measure the difference between two probability distributions; the loss function D of the discriminative model i , uses the cross-entropy loss as the adversarial loss, denoted as

[0018] Further, in the above step 4, the loss of the discriminator consists of the adversarial loss and the distribution regularization loss, and the defining formula:

[0019]

[0020] λ2 is a hyperparameter used to adjust the loss to avoid overfitting of the model.

[0021] Compared with the existing technologies, the beneficial effects of the present invention are:

[0022] (1) In step 1 of the present invention, BiLSTM is adopted. The generality and powerful modeling ability of BiLSTM enable it to adapt to text data in different fields and styles, better process noisy data, and improve the reliability of preprocessing. Moreover, its hidden state can be used to analyze the context information of the text. In the text preprocessing of the present invention, the decision-making process of the model can be understood by visualizing the hidden state.

[0023] (2) In step 3 of the present invention, a variational autoencoder VAE is introduced into the discriminator module. The variational autoencoder VAE is used to normalize and denoise the latent distribution of the latent embedding vector v, reducing the complexity of the image latent distribution. Using VAE also promotes the encoded image feature vector v to record important image semantics.

[0024] (3) In step 4 of the present invention, the loss of the variational autoencoder (VAE) uses the L1 norm to measure the distance between the input image and the image after encoding, transformation, and decoding. The image reconstruction loss uses the KL divergence to measure the difference between two probability distributions. The loss function D of the discriminative model i , uses the cross-entropy loss as the adversarial loss, denoted as In the training stage of the discriminative model, L DiIt can help the discriminant model better identify whether the current image sampling is from the generated image distribution or the real distribution, and better learn the distribution decision boundary between the two potential distributions.

[0025] (4) The real images and the decoded and generated images output in Step 3 of the present invention have extremely strong normalization due to the step-by-step regularization processing of the images. Therefore, after entering Step 4, the convergence speed of the model jointly trained with the adversarial loss becomes faster. In this way, corresponding images can be effectively generated based on the semantics of the text. After testing, the index data obtained by the present invention are all better than those of the original model AttnGAN. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of the network model of the present invention;

[0027] Figure 2 is a structural diagram of the DRM-GAN network model of the present invention;

[0028] Figure 3 is a structural diagram of the discriminator distribution regularization. DETAILED DESCRIPTION OF THE INVENTION

[0029] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the appended drawings and embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, which are only used to illustrate the present invention but not to limit the scope of the present invention.

[0030] As Figure 1 shown, the design principle of the present invention is as follows: First, the text encoder encodes and describes the text of the public dataset to obtain word feature vectors and global sentence feature vectors.

[0031] Then, the text feature vectors and the randomly sampled noise z pass through a three-stage generator to generate images with different resolutions respectively.

[0032] Subsequently, a variational autoencoder is introduced into the discriminator module to perform distribution regularization on the generated images, normalize their potential distributions and achieve denoising, and reduce the distribution complexity of the images. The discriminator then makes a true / false judgment based on the encoded images.

[0033] Finally, the real images and the generated images are used as inputs to calculate the loss of the discriminator, and the parameters are updated backward, and jointly trained with the adversarial loss. After multiple iterations, an optimized model is obtained.

[0034] Based on the above basic idea, the present invention provides a cross-modal text-to-image generation method based on step-by-step regularization, and the specific implementation steps are as follows:

[0035] Step 1. The text encoder preprocesses the input text through a bidirectional long short-term memory recurrent network (BiLSTM), and encodes the input text to obtain a word feature vector e and a global sentence feature vector s, (e, s) = F LSTM (t).

[0036] The global sentence feature contains the overall semantic information of the text, while the word feature represents the feature of each word in the text. To obtain richer text features, the conditional enhancement module converts the sentence feature into a conditional sentence feature, and randomly samples a latent space vector feature representation from a continuous and independent Gaussian distribution.

[0037] Step 2. The feature vector and the randomly sampled noise z are sent as inputs to the generator module. After processing through an upsampling convolutional layer and a residual layer, etc., h0 = F0(z, F s ), a series of hidden layer features h i (i = 0, 1, 2) are generated. Then, through three stages of generators, images with different resolutions are generated, and finally an image x i (i = 0, 1, 2) corresponding to the text is generated. Among them, the alignment between words and subgraphs uses an attention mechanism F attn , h i = F i (h i-1 , F i attn (e, h i-1 )) such that the generation process can pay attention to the details of each word in the text, thereby generating more accurate local image features.

[0038] Step 3. The image x i generated by the generator module is first fed to the encoder to obtain a latent embedding vector v. A variational autoencoder (VAE) is introduced in the discriminator module to learn the latent representation of the data by maximizing the variational lower bound of the data. The discriminator is divided into two sub-modules: the discriminator module and the variational autoencoder module. The discriminator module consists of a discriminative encoder and a logical classifier. By combining text information through the logical classifier, it discriminates whether the image is generated by the generator or belongs to the original dataset. If the image generated by the generator conforms to the semantics, it is determined as a real image; the variational autoencoder module consists of a variational encoder, a variational sampling module, and a decoder. The obtained latent embedding vector v is regularized through the variational sampling module to obtain z * , which is then combined with the sentence vector and fed into the decoder to obtain the decoded generated image.

[0039] The structure diagram of the distribution regularization method (DRM) is as shown in Figure 3 . Steps 3.1 and 3.2 are the detailed steps of the discriminator module, and the remaining steps are the detailed steps of the variational autoencoder module. The specific steps are as follows:

[0040] Step 3.1: Input an image x i , the image x i can be a real image I i and the decoded generated image x i is first input into the discriminative encoder E D (·) to extract the features of the sample, and E D (·) outputs the feature latent embedding vector

[0041] Step 3.2: Combine the embedding vector v with the text embedding s and feed it to the logic classifier ψ(·) to identify whether x i is a real image or a generated image.

[0042] Step 3.3: Use a neural network to fit and calculate the mean μ and variance σ of the feature embedding vector v, and construct a Gaussian distribution in the variational encoder

[0043] Step 3.4: Through operations, further constrain this Gaussian distribution to approximate the standard normal distribution to achieve the normalization of the latent distribution.

[0044] Step 3.5: During the training process, the operation of sampling the latent feature z from each Gaussian distribution is non-differentiable. To make the above process differentiable, the technique of feature resampling is adopted, and z is scaled and translated through the mean μ and variance σ of the Gaussian distribution where v is located to obtain

[0045] Step 3.6: After z * is connected to the text embedding vector s, it is sent to the decoder to reconstruct the image x i , and the image generation is conditioned on the text description to better fuse the information of the text and the image.

[0046] VAE can effectively denoise the latent distribution and reduce its complexity. Based on the advantage of reconstructing images by VAE, the normalized embedding vector can retain the key semantic visual information. Therefore, a VAE module is constructed in the discriminator to normalize the latent distribution of the image to help the discriminator better distinguish real images.

[0047] Step Four: Use the real image and the generated image as the input of the distribution regularization module DRM i of the loss function, jointly train by combining the adversarial loss of multiple discriminators, and then update the parameters through backpropagation to obtain the optimized model.

[0048] Among them, multiple discriminators and generators form a cascaded network for adversarial learning, promoting the generator to generate more realistic images that conform to the text description. The loss function in the generative adversarial network is divided into two parts: the unconditional loss function is used to judge the authenticity of the picture during the training process, and the conditional loss function is used to judge the relevance between the picture and the text; the distribution regularization loss function learns the distribution decision boundary between the latent distributions of the generated images and the real images. After multiple iterative trainings to update the parameters, the loss function is minimized.

[0049] In this step, following the alternating optimization principle of the generative adversarial network, the variational autoencoder network and the discriminative model are trained together.

[0050] The distribution regularization loss function is designed based on the variational lower bound of the variational autoencoder network. The variational autoencoder network loss function in the distribution regularization module DRM i in the i-th stage can be defined as the loss function that combines the loss of the variational autoencoder (VAE) and the image reconstruction loss.

[0051]

[0052] The loss of the variational autoencoder (VAE) uses the L1 norm to measure the distance between the input image and the image after encoding, transformation, and decoding, promoting the reconstructed image to be as similar as possible to the original image.

[0053]

[0054] The image reconstruction loss uses the KL divergence to measure the difference between two probability distributions, prompting the latent variable distribution to approach the normal distribution and improving the generalization ability and generation ability of the model.

[0055]

[0056] The discriminators are disjoint in structure, so they can be trained in parallel, with each discriminator focusing on a single image scale. The loss function D i of the discriminative model is used to train whether the classified input image is from the real image distribution or the fake image distribution, using the cross-entropy loss as the adversarial loss.

[0057]

[0058] Among them, I i and I i ' represent the real image and the generated image, and S represents the text condition. I i ~P data denotes sampling the sample I data from the real data distribution P i , I i '~PGi Denote sampling the sample I Gi ' from the generated image distribution P i '. The unconditional loss function identifies whether the currently input image is from the real image distribution or the generated image distribution; the conditional loss function identifies whether the currently input image semantically matches the semantics of the text features.

[0059] In the training step of the discriminative model, both the distribution of real images and the feature distribution of generated images are normalized. The loss of the discriminator consists of adversarial loss and distribution regularization loss, and the formula is defined as follows:

[0060]

[0061] λ2 is a hyperparameter used to adjust the loss to avoid overfitting of the model.

[0062] Finally, the optimized model is trained using a public dataset to obtain a text-to-image generation model (DRM-GAN) with better performance, which can capture semantic information in the text more accurately and enhance the ability to discriminate between real and fake images.

[0063] The Inception Score (IS) and the Frechet Inception Distance (FID) are used as evaluation metrics to measure the quality, diversity, and similarity of the generated images to real images. The structure diagram of DRM-GAN is as shown in Figure 2 Figure. The alignment between words and subgraphs uses the attention mechanism F attn , enabling the generation process to focus on the details of each word in the text. The variational autoencoder regularizes and denoises the image distribution, effectively reducing the complexity of the images.

[0064] The IS metric calculates the Kullback-Leibler (KL) divergence between the conditional distribution and the marginal distribution. The larger the IS metric, the better the quality and diversity of the generated images. The formula for calculating IS is shown in Equation (6).

[0065]

[0066] where x is the image generated by the generator; y is the predicted label, p(y|x) is the conditional distribution of the generated image in each category; p(y) is the marginal distribution; D KL represents the KL divergence of the two distributions. is the expectation operation, and x is sampled from the distribution p g .

[0067] The FID metric calculates the distance between the feature vectors of real images and generated images. The smaller the FID, the higher the similarity between the generated images and real images, and the closer the generated images are to real images. The formula for calculating FID is shown in Equation (7).

[0068] FID = ||μ r -μ g || 2 + Tr(ν r + ν g - 2(ν r ν g ) 1 / 2 ) (7)

[0069] where μ r , μ g are the mean values of the features of the real image and the generated image respectively; ν r , ν g are the feature covariance matrices of the real image and the generated image respectively.

[0070] The description of the above - mentioned embodiments is relatively specific and detailed, but it should not be understood as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.

[0071] As shown in Table 1, it is a comparison of the specific indicators of this invention patent and other algorithms. Compared with the basic network AttnGAN, the IS score of the COCO dataset has increased from 25.89 to 27.32, an increase of 5.52%. The FID score has decreased from 35.49 to 30.84, a decrease of 13.1%, indicating that the model in this paper performs better in terms of image quality and diversity.

[0072] Table 1 Data comparison between the invention and other algorithms

[0073]

[0074] The above is an explanation of the specific implementation of the present invention, rather than a limitation on the present invention. Those skilled in the relevant technical field can also make various equivalent technical solutions without departing from the scope of the present invention. Therefore, all equivalent technical solutions should be included in the patent protection scope of the present invention.

Claims

1. A cross-modal text-to-image generation method based on distribution regularization, characterized in that: The specific steps are as follows: Step 1: The text encoder preprocesses the text corresponding to the public dataset through a bidirectional long short-term memory recurrent network (BiLSTM). The preprocessing refers to encoding the text to obtain word feature vectors and global sentence feature vectors; Step 2: The feature vector and randomly sampled noise z are sent as input to the generator module, and are processed and transformed by the upsampling convolution layer and the residual layer to generate a series of hidden layer features h i (i=0,1,2), and then the generator generates images of different resolutions through three stages, and finally generates an image x corresponding to the text i (i=0,1,2); Step 3: The image x given by the generator module i is first fed to the encoder to obtain the latent embedding vector v. A variational autoencoder (VAE) is introduced in the discriminator module. The discriminator is divided into two sub-modules: the discriminator module and the variational autoencoder module. The discriminator module consists of a discriminative encoder and a logistic classifier. By combining text information through the logistic classifier, it discriminates whether the image is generated by the generator or belongs to the original dataset. An image generated by the generator that conforms to the semantics is determined to be a real image. The variational autoencoder module consists of a variational encoder, a variational sampling module, and a decoder. The obtained latent embedding vector v is regularized through the variational sampling module to obtain z * , which is then combined with the sentence vector and fed into the decoder to obtain the generated image; Step 4: Use the real image and the generated image as the input of the distribution regularization loss function, and jointly train with the adversarial loss of the discriminator generative adversarial network. Update the parameters in reverse to obtain the optimal parameter model.

2. The cross-modal text-to-image generation method based on distribution regularization according to claim 1, characterized in that: In the third step, the variational sampling module regularizes to obtain z * The process is as follows: First, a Gaussian distribution is constructed Then, through operations, the normalization of the latent distribution is achieved. Finally, feature resampling is adopted to obtain 3. A cross-modal text-to-image generation method based on distribution regularization according to claim 1, characterized in that: In Step 4, the loss function in the generative adversarial network is divided into two parts: the conditional loss function and the distribution regularization loss function. Among them, the conditional loss function has two types: unconditional loss and conditional loss.

4. A cross-modal text-to-image generation method based on distribution regularization according to claim 1, characterized in that: In the fourth step, based on the design of the variational lower bound of the variational autoencoder network, the loss function of the variational autoencoder network in the distribution regularization module DRM i in the i-th stage is as follows: The loss of the variational autoencoder (VAE) uses the L1 norm to measure the distance between the input image and the image after encoding, transformation, and decoding; the image reconstruction loss uses the KL divergence to measure the difference between two probability distributions; the loss function D of the discriminative model i , uses the cross-entropy loss as the adversarial loss, denoted as 5. A cross-modal text-to-image generation method based on distribution regularization according to claim 1, characterized in that: In Step 4, the loss of the discriminator consists of the adversarial loss and the distribution regularization loss, and the definition formula is: λ2 is a hyperparameter used to adjust the loss to avoid overfitting of the model.

Citation Information

Patent Citations

  • Text generation image learning model based on deep learning

    CN113869007A

  • Cross-modal text-to-image generation method based on generative adversarial network

    CN114329025A