Text-to-image generation method and device, medium and product

By aligning the feature space of the diffusion model and self-supervised model, combining the acceleration module and the relationship alignment module, the problems of low efficiency and unstable quality of diffusion model generation are solved, and efficient and high-quality text-to-image generation is achieved.

CN120339460APending Publication Date: 2025-07-18CHINA UNITED NETWORK COMM GRP CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510370258.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the diffusion model is inefficient and the image quality is unstable when generating images, which is mainly due to the differences in the feature spatial distribution of self-supervised learning model and diffusion model and the constraints of distillation technology, which leads to the convergence of the model.

Method used

By combining the diffusion model, self-supervised model and mapper of the text-to-image generation task, an alignment system is established, and the acceleration module and the relational alignment module are used to reduce the difference in the spatial distribution of features to achieve rapid convergence of the model.

Benefits of technology

While maintaining the generation efficiency, the image quality is significantly improved, ensuring the stability and performance of the generated image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339460A_ABST
    Figure CN120339460A_ABST
Patent Text Reader

Abstract

The invention provides a text-to-image generation method and device, a medium and a product, and the method comprises the steps: obtaining a first diffusion model and a first self-supervision model corresponding to a text-to-image generation task, and obtaining a first mapper; obtaining preprocessed sample data; establishing a first stage alignment system according to the first diffusion model, the first self-supervised model and the first mapper; according to the preprocessed sample data, performing first-stage alignment processing on the first-stage alignment system to obtain a second self-supervised model and a second mapper; mounting an acceleration module on the first diffusion model, and establishing a relation alignment module; establishing a second-stage alignment system according to the first diffusion model on which the acceleration module is mounted, a second self-supervised model, a second mapper and a relation alignment module; and according to the preprocessed sample data, performing second-stage alignment processing on the second-stage alignment system to obtain a target diffusion model, and realizing text-to-image generation. The efficiency and the quality of generating images by texts are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimodal artificial intelligence technology, and in particular, to a method, device, medium, and product for text-to-image generation. Background Art

[0002] Diffusion models are a type of generative model that convert random noise into clear images by simulating the physical diffusion process. However, diffusion models need to perform multiple rounds of iterative sampling steps, which reduces the efficiency of generating images.

[0003] In the prior art, distillation technology is usually used to compress the steps of image generation, and self-supervised learning models are used to improve the effect of generated images. However, there are differences in the feature space distributions between self-supervised learning models and diffusion models, and the constraints of distillation technology further exacerbate the differences in feature space distributions, making it difficult for the model to converge, thereby reducing the model performance and resulting in unstable image quality.

[0004] Based on this, in the prior art, there is a technical problem of unstable image quality in the generated images. Summary of the Invention

[0005] Embodiments of this application provide a method, device, medium, and product for text-to-image generation to achieve the technical effect of improving image quality.

[0006] In a first aspect, embodiments of this application provide a method for text-to-image generation, including:

[0007] Obtain a first diffusion model and a first self-supervised model corresponding to a text-to-image generation task, and obtain a first mapper;

[0008] Obtain preprocessed sample data;

[0009] Establish a first-stage alignment system according to the first diffusion model, the first self-supervised model, and the first mapper;

[0010] Perform first-stage alignment processing on the first-stage alignment system according to the preprocessed sample data to obtain a second self-supervised model and a second mapper;

[0011] Mount an acceleration module on the first diffusion model and establish a relationship alignment module;

[0012] Establish a second-stage alignment system according to the first diffusion model with the acceleration module mounted, the second self-supervised model, the second mapper, and the relationship alignment module;

[0013] Perform second-stage alignment processing on the second-stage alignment system according to the preprocessed sample data to obtain a target diffusion model, where the target diffusion model is used to implement text-to-image generation.

[0014] In a possible implementation, the preprocessed sample data includes preprocessed image data and preprocessed text data corresponding to the preprocessed image data.

[0015] In a possible implementation, according to the preprocessed sample data, perform first-stage alignment processing on the first-stage alignment system to obtain a second self-supervised model and a second mapper, including:

[0016] In the first-stage alignment system, freeze the parameters of the first diffusion model;

[0017] Input the preprocessed text data into the first diffusion model to obtain a first result through the first diffusion model;

[0018] Input the first result into the first mapper for training to obtain first characterization data;

[0019] Input the preprocessed image data into the first self-supervised model to obtain second characterization data through the first self-supervised model, and obtain the self-supervised loss function of the first self-supervised model;

[0020] Determine the first feature alignment loss function according to the first characterization data and the second characterization data;

[0021] Determine the second self-supervised model and the second mapper according to the self-supervised loss function and the first feature alignment loss function.

[0022] In a possible implementation, determine the second self-supervised model and the second mapper according to the self-supervised loss function and the first feature alignment loss function, including:

[0023] Determine the first-stage total loss function according to the self-supervised loss function and the first feature alignment loss function;

[0024] According to the first-stage total loss function, update the first self-supervised model and the first mapper through backpropagation to obtain the second self-supervised model and the second mapper.

[0025] In a possible implementation, according to the preprocessed sample data, perform second-stage alignment processing on the second-stage alignment system to obtain a target diffusion model, including:

[0026] In the second-stage alignment system, freeze the parameters of the second self-supervised model and the first diffusion model;

[0027] Input the preprocessed text data into the first diffusion model equipped with an acceleration module to obtain a second result through the first diffusion model equipped with an acceleration module;

[0028] Input the second result into the second mapper for training to obtain third characterization data;

[0029] Input the preprocessed image data into the second self-supervised model to obtain fourth representation data through the second self-supervised model, and obtain the self-supervised loss function of the second self-supervised model;

[0030] Obtain the distillation loss function of the first diffusion model with an acceleration module mounted;

[0031] Determine the second feature alignment loss function according to the third representation data and the fourth representation data;

[0032] Obtain the relationship alignment loss function according to the relationship alignment module;

[0033] Determine the target diffusion model according to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function.

[0034] In a possible implementation manner, determining the target diffusion model according to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function includes:

[0035] According to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function, perform gradient backpropagation to update the parameters of the acceleration module;

[0036] Train to obtain the target diffusion model according to the acceleration module with updated parameters.

[0037] In a possible implementation manner, after performing the second-stage alignment process on the second-stage alignment system according to the preprocessed sample data to obtain the target diffusion model, it further includes:

[0038] Obtain the text data to be processed;

[0039] Input the text data into the target diffusion model to determine the corresponding image data according to the output result of the target diffusion model.

[0040] In a second aspect, an embodiment of the present application provides a text-to-image generation device, including: a memory, a processor;

[0041] The memory stores computer execution instructions;

[0042] The processor executes the computer execution instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.

[0043] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.

[0044] Fourthly, an embodiment of the present application provides a computer program product, including a computer program, which when executed by a processor, implements the above first aspect and / or various possible implementation manners of the first aspect.

[0045] The text-to-image generation method, device, medium, and product provided by the embodiments of the present application obtain a first diffusion model, a first self-supervised model, a first mapper, and preprocessed sample data corresponding to a text-to-image generation task; establish a first-stage alignment system based on the first diffusion model, the first self-supervised model, and the first mapper; perform first-stage alignment processing on the first-stage alignment system according to the preprocessed sample data to obtain a second self-supervised model and a second mapper; thereby reducing the difference in the distribution of the feature space and making the basic features initially consistent; mount an acceleration module on the first diffusion model and establish a relationship alignment module; establish a second-stage alignment system based on the first diffusion model with the acceleration module mounted, the second self-supervised model, the second mapper, and the relationship alignment module; perform second-stage alignment processing on the second-stage alignment system according to the preprocessed sample data to obtain a target diffusion model, so as to realize text-to-image generation, thereby improving the image quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0047] Figure 1 It is a schematic diagram of a text-to-image generation system architecture provided by the present application;

[0048] Figure 2 It is a flowchart of a text-to-image generation method provided by the present application Figure 1 ;

[0049] Figure 3 It is a schematic diagram of the structure of the first-stage alignment system provided by the present application;

[0050] Figure 4 It is a schematic diagram of the structure of the second-stage alignment system provided by the present application;

[0051] Figure 5 It is a flowchart of a text-to-image generation method provided by the present application Figure 2 ;

[0052] Figure 6 It is a flowchart of a text-to-image generation method provided by the present application Figure 3 ;

[0053] Figure 7 It is a flowchart of a text-to-image generation method provided by the present application Figure 4 ;

[0054] Figure 8 Flow schematic of a text - to - image generation method provided for this application Figure 5 ;

[0055] Figure 9 Flow schematic of a text - to - image generation method provided for this application Figure 6 ;

[0056] Figure 10 Structural schematic diagram of a text - to - image generation device provided for this application;

[0057] Figure 11 Structural schematic diagram of a text - to - image generation device provided for this application

[0058] Through the above - mentioned drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0059] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0060] It should be noted that the data involved in this application (including but not limited to text data, image data, and pre - processed sample data) are all information and data authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0061] Diffusion models are a class of generative models that generate high - quality image data by gradually adding and removing noise. However, since diffusion models need to perform multiple rounds of iterative sampling steps, the efficiency of generating images decreases.

[0062] In the prior art, methods such as model pruning, quantization, cache mechanism, and distillation technology can be used to improve the generation efficiency of images. Among them, the distillation technology is used to compress the steps of image generation, which significantly improves the generation efficiency of images. However, this method also leads to a decrease in the quality of the generated images. Further, a self-supervised learning model can be used to improve the effect of the generated images. However, there are differences in the feature space distributions between the self-supervised learning model and the diffusion model, and the constraints of the distillation technology further exacerbate the differences in the feature space distributions, making it difficult for the model to converge, thus reducing the model performance and resulting in unstable quality of the generated images.

[0063] To solve the above problems, the core concept of this application is as follows: By combining a first diffusion model, a first self-supervised model, and a first mapper corresponding to the text-to-image generation task, a first-stage alignment system is established; using preprocessed sample data to perform first-stage alignment processing on the first-stage alignment system to obtain a second self-supervised model and a second mapper; thereby reducing the differences in the feature space distributions of the two models and making the basic features initially consistent; mounting an acceleration module on the first diffusion model and establishing a relationship alignment module; based on the first diffusion model with the acceleration module mounted, the second self-supervised model, the second mapper, and the relationship alignment module, a second-stage alignment system is established; according to the preprocessed sample data, performing second-stage alignment processing on the second-stage alignment system to accelerate the model convergence and obtain a target diffusion model to achieve text-to-image generation, thereby improving the quality of the images while maintaining the efficiency of text-generated images.

[0064] Optionally, Figure 1 is a schematic diagram of a text-to-image generation system architecture provided by this application. As Figure 1 shown, the text-to-image generation system architecture includes at least one of a data acquisition device 101, a processing device 102, and a display device 103.

[0065] It can be understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the above architecture. In other feasible embodiments of this application, the above architecture may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements, which can be specifically determined according to the actual application scenario and will not be limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0066] In the specific implementation process, the data acquisition device 101 may include an input / output interface or a communication interface. The data acquisition device 101 can be connected to the processing device through the input / output interface or the communication interface.

[0067] The processing device 102 can be used to establish a first-stage alignment system according to a first diffusion model, a first self-supervised model, and a first mapper; perform first-stage alignment processing on the first-stage alignment system according to preprocessed sample data to obtain a second self-supervised model and a second mapper; mount an acceleration module on the first diffusion model and establish a relationship alignment module; establish a second-stage alignment system according to the first diffusion model with the acceleration module mounted, the second self-supervised model, the second mapper, and the relationship alignment module; perform second-stage alignment processing on the second-stage alignment system according to preprocessed sample data to obtain a target diffusion model, thereby realizing text-to-image generation.

[0068] The display device 103 can also be a touch display screen or the screen of a terminal device, which is used to receive user instructions while displaying the above content to achieve interaction with the user.

[0069] The following uses specific embodiments to elaborate in detail on the technical solutions of this application and how the technical solutions of this application solve the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0070] Figure 2 Flow schematic of a text-to-image generation method provided by this application Figure 1 , as Figure 2 shown, this method includes:

[0071] S201. Obtain a first diffusion model and a first self-supervised model corresponding to a text-to-image generation task, and obtain a first mapper.

[0072] In this embodiment, the first diffusion model corresponding to the text-to-image generation task includes SDXL (stable-diffusion-xl-base-1.0, an open-source image generation model for stable diffusion expansion); the first self-supervised model corresponding to the text-to-image generation task includes DinoV2 (Distillation with NO labels-V2, robust visual features for unsupervised learning), MoCo (Momentum Contrast for Unsuper Visual Representation Learning, momentum contrast for unsupervised visual representation learning), or MAE (Masked Autoencoder, masked autoencoder); the first mapper includes a multi-layer perceptron.

[0073] S202. Obtain preprocessed sample data.

[0074] Optionally, the preprocessed sample data includes preprocessed image data and preprocessed text data corresponding to the preprocessed image data.

[0075] In this embodiment, the image data and the text data corresponding to the image data are synchronously processed to ensure the alignment of the image data and the text data at the sample level, and the synchronously processed sample data is batch-processed to obtain preprocessed sample data; so as to synchronously input the model during training and improve the training efficiency.

[0076] S203. Establish a first-stage alignment system according to the first diffusion model, the first self-supervised model, and the first mapper.

[0077] In this embodiment, a first-stage alignment system is established according to the first diffusion model, the first self-supervised model, and the first mapper to improve the consistency of the feature distributions between the first diffusion model and the first self-supervised model.

[0078] S204. Perform first-stage alignment processing on the first-stage alignment system according to the preprocessed sample data to obtain a second self-supervised model and a second mapper.

[0079] In this embodiment, as Figure 3 shown, the preprocessed sample data is input into the first-stage alignment system including the first diffusion model, the first mapper, and the first self-supervised model for first-stage alignment processing to obtain a second self-supervised model and a second mapper; after the first-stage alignment processing, the difference in the feature spaces output by the first diffusion model and the second self-supervised learning model is minimized, thereby reducing the difficulty of feature matching.

[0080] S205. Mount an acceleration module on the first diffusion model and establish a relationship alignment module.

[0081] In this embodiment, for example, the acceleration module is LCM-Lora (Low-Rank Adaptation of Large Language Models, a low-rank adaptation technology for large language models). By mounting the acceleration module on the first diffusion model, the computational amount and storage requirements can be reduced through the decomposition of the low-rank matrix, and the model parameters can be reduced, thereby accelerating the training process.

[0082] The relationship alignment module is a preset alignment method that transfers knowledge by learning the relationship between the first diffusion model with an acceleration module and the second self-supervised model. For example, if the sample representations of the first diffusion model with an acceleration module are abstracted as S1, S2, and S3, and the sample representations of the second self-supervised model are abstracted as Z1, Z2, and Z3, and R is the relationship between samples, then the relationship alignment module matches the relationships of S1, S2, S3, Z1, Z2, and Z3 as R(S1, S2, S3) vs R(Z1, Z2, Z3); compared with (S1 vs Z1, S2 vs Z 2, S3 vs Z3) in feature alignment, by strong alignment to reduce the distance between corresponding samples, the relationship alignment module is used to weakly match the samples, realizing the transition from strong alignment to weak match, eliminating the limitation of only being able to focus on the features of a single sample in feature alignment, and reducing the difficulty of feature alignment.

[0083] S206. Establish a second-stage alignment system according to the first diffusion model with an acceleration module, the second self-supervised model, the second mapper, and the relationship alignment module.

[0084] In this embodiment, by establishing a second-stage alignment system according to the first diffusion model with an acceleration module, the second self-supervised model, the second mapper, and the relationship alignment module, the alignment of the overall relationship of the samples is realized, avoiding the influence of the sample representation dimension during relationship alignment, and improving the flexibility of relationship alignment.

[0085] S207. Perform second-stage alignment processing on the second-stage alignment system according to the preprocessed sample data to obtain a target diffusion model, where the target diffusion model is used to implement text-to-image generation.

[0086] In this embodiment, as Figure 4 shown, input the preprocessed sample data into the second-stage alignment system including the first diffusion model with an acceleration module, the second self-supervised model, the second mapper, and the relationship alignment module for second-stage alignment processing to obtain a target diffusion model; after the second-stage alignment processing, the alignment method of the target diffusion model is more flexible, not limited by the sample representation dimension, and the performance and applicability of the target diffusion model are improved.

[0087] Optionally, as Figure 5 shown, after performing second-stage alignment processing on the second-stage alignment system according to the preprocessed sample data to obtain a target diffusion model, it further includes:

[0088] S208. Obtain the text data to be processed.

[0089] In this embodiment, the text data to be processed can be descriptive text or prompt statements generated in response to user needs.

[0090] S209. Input the text data into the target diffusion model to determine the corresponding image data based on the output result of the target diffusion model.

[0091] In this embodiment, the text data is input into the target diffusion model, and the target diffusion model is used to generate the image data corresponding to the text data.

[0092] The text-to-image generation method provided by the embodiments of the present application includes obtaining a first diffusion model, a first self-supervised model, a first mapper, and preprocessed sample data corresponding to a text-to-image generation task; establishing a first-stage alignment system based on the first diffusion model, the first self-supervised model, and the first mapper; performing first-stage alignment processing on the first-stage alignment system according to the preprocessed sample data to obtain a second self-supervised model and a second mapper, thereby preliminarily reducing the difference in the distribution of the feature space; mounting an acceleration module on the first diffusion model and establishing a relationship alignment module; establishing a second-stage alignment system based on the first diffusion model with the acceleration module mounted, the second self-supervised model, the second mapper, and the relationship alignment module; performing second-stage alignment processing on the second-stage alignment system according to the preprocessed sample data to obtain a target diffusion model, realizing text-to-image generation, accelerating the convergence speed of the model, improving the efficiency of text-generated images, and ensuring the quality of the generated images.

[0093] Figure 6 Flow schematic of a text-to-image generation method provided by the present application Figure 3 , such as Figure 6 shown. Based on the Figure 2 embodiment, the step of performing first-stage alignment processing on the first-stage alignment system according to the preprocessed sample data in step S204 above to obtain a second self-supervised model and a second mapper is described in detail. The method includes:

[0094] S601. Freeze the parameters of the first diffusion model in the first-stage alignment system.

[0095] In this embodiment, for example, freezing the parameters of the SDXL model in the first-stage alignment system can effectively reduce the number of parameters to be updated, thereby reducing the training time, accelerating the convergence speed of the first diffusion model, and improving the generalization ability of the first diffusion model.

[0096] S602. Input the preprocessed text data into the first diffusion model to obtain a first result through the first diffusion model.

[0097] In this embodiment, the first result is that the preprocessed text data is trained by the first diffusion model to generate corresponding feature vectors; the preprocessed text data is input into the SDXL model to obtain the first result corresponding to the preprocessed text data.

[0098] S603. Input the first result into the first mapper for training to obtain the first characterization data.

[0099] In this embodiment, the first characterization data is a new feature vector converted by the first mapper based on the first result. For example, through the first mapper, the first result is converted into the first characterization data p1(x) so that the feature vectors output by the first diffusion model and the first self-supervised model are consistent in dimension, achieving the effect of unified form.

[0100] S604. Input the preprocessed image data into the first self-supervised model to obtain the second characterization data through the first self-supervised model, and obtain the self-supervised loss function of the first self-supervised model.

[0101] In this embodiment, for example, the preprocessed image data is input into the Dinov2 model, and after being trained by the Dinov2 model, the second characterization data q1(x) is obtained.

[0102] Further, the first self-supervised model includes a student network and a teacher network. Among them, the parameters of the teacher network are updated by the exponential moving average of the student network. The self-supervised loss function of the first self-supervised model is a cross-entropy loss, which is used to measure the difference between the feature distributions output by the student network and the teacher network.

[0103] The calculation formula of the self-supervised loss function of the first self-supervised model is as follows:

[0104]

[0105] In the formula, loss1 is the self-supervised loss function of the first self-supervised model; N is the number of preprocessed image data in the batch; K is the feature dimension; i is the identifier of the preprocessed image data; j is the identifier of the feature dimension; x i is the preprocessed image data; is the output probability of the teacher network for the j-th feature dimension of the sample x i ; is the output probability of the student network for the j-th feature dimension of the sample x i ;

[0106] S605. Determine the first feature alignment loss function according to the first characterization data and the second characterization data.

[0107] In this embodiment, by calculating the Kullback-Leibler Divergence, the difference between the distributions of the first characterization data and the second characterization data is quantified. The calculation formula of the first feature alignment loss function is as follows:

[0108]

[0109] In the formula, loss2 is the first feature alignment loss function; x is the preprocessed sample data; p1(x) is the first characterization data; q1(x) is the second characterization data.

[0110] S606. Determine the second self-supervised model and the second mapper according to the self-supervised loss function and the first feature alignment loss function.

[0111] In this embodiment, the parameters of the first diffusion model are kept frozen, and the first self-supervised model is updated according to the self-supervised loss function and the first feature alignment loss function to obtain the second self-supervised model, and the parameters of the first mapper are updated to obtain the second mapper.

[0112] Optionally, as Figure 7 shown, determining the second self-supervised model and the second mapper according to the self-supervised loss function and the first feature alignment loss function includes:

[0113] S6061. Determine the total loss function of the first stage according to the self-supervised loss function and the first feature alignment loss function.

[0114] In this embodiment, the total loss function L reverse of the first stage has the following calculation formula:

[0115]

[0116] In the formula, loss1 is the self-supervised loss function of the first self-supervised model; loss2 is the first feature alignment loss function; is the scale factor.

[0117] S6062. Update the first self-supervised model and the first mapper through backpropagation according to the total loss function of the first stage to obtain the second self-supervised model and the second mapper.

[0118] In this embodiment, according to the total loss function of the first stage, the gradients of the total loss function of the first stage with respect to each parameter in the first self-supervised model and the first mapper are calculated through the chain rule; in the way of backpropagation, starting from the output layer, the gradients are propagated layer by layer towards the input layer, and a gradient multiplication operation is performed to obtain the global gradient; then, based on the optimization algorithm of stochastic gradient descent, the parameters of the first self-supervised model and the first mapper are updated to obtain the second self-supervised model and the second mapper.

[0119] The text-to-image generation method provided in the embodiment of the present application freezes the parameters of the first diffusion model; trains the preprocessed text data through the first diffusion model and the first mapper to obtain the first characterization data; trains the preprocessed image data through the first self-supervised model to obtain the second characterization data; calculates the first feature alignment loss function based on the first characterization data and the second characterization data; and calculates the self-supervised loss function of the first self-supervised model; updates the parameters of the first self-supervised model and the first mapper to obtain the second self-supervised model and the second mapper; thereby minimizing the difference in the feature spaces output by the diffusion model and the self-supervised learning model and reducing the difficulty of feature matching.

[0120] Figure 8 Schematic flow of a text-to-image generation method provided by the present application Figure 4 , as Figure 8 shown, based on the Figure 2 embodiment, the step S207 of performing second-stage alignment processing on the second-stage alignment system according to the preprocessed sample data to obtain the target diffusion model is described in detail. The method includes:

[0121] S801. In the second-stage alignment system, freeze the parameters of the second self-supervised model and the first diffusion model.

[0122] In this embodiment, in the second-stage alignment system, freezing the parameters of the second self-supervised model and the first diffusion model can effectively reduce the number of parameters to be updated in the second-stage alignment processing, thereby reducing the training time, accelerating the convergence speed of the first diffusion model, and improving the generalization ability of the second self-supervised model and the first diffusion model.

[0123] S802. Input the preprocessed text data into the first diffusion model equipped with an acceleration module to obtain a second result through the first diffusion model equipped with an acceleration module.

[0124] In this embodiment, the second result is the corresponding feature vector generated after the preprocessed text data is trained by the first diffusion model equipped with an acceleration module; for example, input the preprocessed text data into the SDXL model equipped with LCM-Lora to obtain the second result corresponding to the preprocessed text data.

[0125] S803. Input the second result into the second mapper for training to obtain the third characterization data.

[0126] In this embodiment, the third characterization data is a new feature vector converted by the second mapper based on the second result. For example, through the second mapper, the second result is converted into the third characterization data P2(x), so that the feature vectors output by the first diffusion model and the second self-supervised model equipped with the acceleration module are consistent in dimension, achieving the effect of unified form.

[0127] S804. Input the preprocessed image data into the second self-supervised model to obtain the fourth characterization data through the second self-supervised model, and obtain the self-supervised loss function of the second self-supervised model.

[0128] In this embodiment, for example, the preprocessed image data is input into the second self-supervised model, and after being trained by the second self-supervised model, the fourth characterization data q2(x) is obtained.

[0129] Furthermore, the calculation formula of the self-supervised loss function of the second self-supervised model is as follows:

[0130]

[0131] In the formula, loss3 is the self-supervised loss function of the second self-supervised model; N is the number of preprocessed image data in the batch; K is the feature dimension; i is the identifier of the preprocessed image data; j is the identifier of the feature dimension; x i is the preprocessed image data; is the output probability of the second self-supervised model for the j-th feature dimension of the sample x i ; is the output probability of the first self-supervised model for the j-th feature dimension of the sample x i .

[0132] S805. Obtain the distillation loss function of the first diffusion model equipped with the acceleration module.

[0133] In this embodiment, the first diffusion model equipped with the acceleration module includes a student diffusion model and a teacher diffusion model, where the teacher diffusion model is obtained by updating the parameters of the student diffusion model through exponential moving average.

[0134] Furthermore, the calculation formula of the distillation loss function is as follows:

[0135]

[0136] In the formula, loss4 is the distillation loss function; x is the preprocessed sample data; t n+1 and tn is the time step; is the target number of steps; is the student diffusion model; is the student diffusion model the input at time step t and the target number of steps the output under; is the teacher diffusion model; is the ODE (Ordinary Differential Equation Solver) solver; is the result predicted by the ODE solver for time step t n+1 to the previous time step t n the result; is the teacher diffusion model at time step t n the input and the target number of steps the output under; represents the square of the L2 norm, used to measure the difference between the outputs of the student diffusion model and the teacher diffusion model.

[0137] S806. Determine the second feature alignment loss function according to the third characterization data and the fourth characterization data.

[0138] In this embodiment, the difference between the distributions of the third characterization data and the fourth characterization data is quantified by calculating the KL divergence. The calculation formula of the second feature alignment loss function is as follows:

[0139]

[0140] In the formula, loss5 is the second feature alignment loss function; x is the preprocessed sample data; p2(x) is the third characterization data; q2(x) is the fourth characterization data.

[0141] S807. Obtain the relationship alignment loss function according to the relationship alignment module.

[0142] In this embodiment, the feature representations output by the first diffusion model with an acceleration module and the second self-supervised model are aligned by calculating the similarity matrix between samples within a batch, so that the relationship alignment module is not affected by the dimension during relationship alignment.

[0143] For example, if the dimension of the third characterization data p2(x) is (B, N p , D p ), and the dimension of the fourth characterization data q2(x) is (B, N q , D q ), where B is the batch size, N pand N q is the number of patches (also known as image blocks) for each image, D p and D q are the embedding dimensions for each patch; after flattening p2(x) and q2(x) into two-dimensional vectors, we obtain (B, N p D p ) and (B, N q D q ), calculate the similarity matrix of p2(x) and q2(x) in the batch direction, and obtain the first similarity matrix m corresponding to p2(x) p(x) and the second similarity matrix m corresponding to q2(x) q(x) , and the dimensions of the first similarity matrix m p(x) and the second similarity matrix m q(x) are finally both (B, B); and optimize the distance between the first similarity matrix m p(x) and the second similarity matrix m q(x) through the Huber loss (also known as the smooth absolute value loss) to obtain the relationship alignment loss function .

[0144] S808. Determine the target diffusion model according to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function.

[0145] Optionally, as Figure 9 shown, determine the target diffusion model according to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function, including:

[0146] S8081. Update the parameters of the acceleration module by gradient backpropagation according to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function.

[0147] In this embodiment, calculate the total loss function of the second stage through the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function. Among them, the calculation formula of the total loss function of the second stage is as follows:

[0148]

[0149] In the formula, Loss is the total loss function of the second stage; , , is the scaling factor; loss4 is the distillation loss function; loss5 is the second feature alignment loss function; is the relationship alignment loss function.

[0150] Using the backpropagation algorithm, calculate the gradients of the total loss function in the second stage with respect to the parameters of the acceleration module; using the optimization algorithm, update the parameters of the acceleration module according to the gradients of the total loss function in the second stage with respect to the parameters of the acceleration module.

[0151] S8082. Train the target diffusion model based on the acceleration module with updated parameters.

[0152] In this embodiment, use the first diffusion model mounted with the acceleration module with updated parameters to perform training for text-to-image generation, so as to obtain the target diffusion model.

[0153] The text-to-image generation method provided by the embodiments of the present application freezes the parameters of the second self-supervised model and the first diffusion model, and mounts the acceleration module on the first diffusion model; inputs the preprocessed text data into the first diffusion model mounted with the acceleration module for training to obtain a second result; inputs the second result into the second mapper for training to obtain third characterization data; inputs the preprocessed image data into the second self-supervised model to obtain fourth characterization data, and obtain the self-supervised loss function of the second self-supervised model; obtain the distillation loss function of the first diffusion model mounted with the acceleration module; determine the second feature alignment loss function according to the third characterization data and the fourth characterization data; obtain the relationship alignment loss function according to the relationship alignment module; determine the target diffusion model according to the distillation loss function, the second feature alignment loss function and the relationship alignment loss function, so as to improve the flexibility and integrity of feature alignment, accelerate the convergence of the model, and thus improve the rate and generation effect of text-to-image generation.

[0154] Figure 10 It is a schematic structural diagram of the text-to-image generation device provided by the present application, as Figure 10 shown, the text-to-image generation device provided by this embodiment includes:

[0155] The first acquisition module 1001 is used to acquire the first diffusion model and the first self-supervised model corresponding to the text-to-image generation task, and acquire the first mapper.

[0156] The second acquisition module 1002 is used to acquire the preprocessed sample data.

[0157] Optionally, the preprocessed sample data includes preprocessed image data and preprocessed text data corresponding to the preprocessed image data.

[0158] The first establishment module 1003 is used to establish the first-stage alignment system according to the first diffusion model, the first self-supervised model and the first mapper;

[0159] The first obtaining module 1004 is configured to perform first-stage alignment processing on the first-stage alignment system according to the preprocessed sample data, so as to obtain a second self-supervised model and a second mapper;

[0160] The second establishing module 1005 is configured to mount an acceleration module on the first diffusion model and establish a relationship alignment module;

[0161] The third establishing module 1006 is configured to establish a second-stage alignment system according to the first diffusion model mounted with the acceleration module, the second self-supervised model, the second mapper, and the relationship alignment module;

[0162] The second obtaining module 1007 is configured to perform second-stage alignment processing on the second-stage alignment system according to the preprocessed sample data, so as to obtain a target diffusion model, where the target diffusion model is used to implement text-to-image generation.

[0163] Optionally, the first obtaining module 1004 may specifically further be configured to:

[0164] In the first-stage alignment system, freeze the parameters of the first diffusion model;

[0165] Input the preprocessed text data into the first diffusion model to obtain a first result through the first diffusion model;

[0166] Input the first result into the first mapper for training to obtain first characterization data;

[0167] Input the preprocessed image data into the first self-supervised model to obtain second characterization data through the first self-supervised model, and obtain the self-supervised loss function of the first self-supervised model;

[0168] Determine a first feature alignment loss function according to the first characterization data and the second characterization data;

[0169] Determine the second self-supervised model and the second mapper according to the self-supervised loss function and the first feature alignment loss function.

[0170] Optionally, the first obtaining module 1004 may specifically further be configured to:

[0171] Determine a first-stage total loss function according to the self-supervised loss function and the first feature alignment loss function;

[0172] Update the first self-supervised model and the first mapper through backpropagation according to the first-stage total loss function to obtain the second self-supervised model and the second mapper.

[0173] Optionally, the second obtaining module 1007 may specifically further be configured to:

[0174] In the second-stage alignment system, freeze the parameters of the second self-supervised model and the first diffusion model;

[0175] Input the preprocessed text data into the first diffusion model equipped with an acceleration module to obtain a second result through the first diffusion model equipped with an acceleration module;

[0176] Input the second result into the second mapper for training to obtain third representation data;

[0177] Input the preprocessed image data into the second self-supervised model to obtain fourth representation data through the second self-supervised model, and obtain the self-supervised loss function of the second self-supervised model;

[0178] Obtain the distillation loss function of the first diffusion model equipped with an acceleration module;

[0179] Determine the second feature alignment loss function according to the third representation data and the fourth representation data;

[0180] Obtain the relationship alignment loss function according to the relationship alignment module;

[0181] Determine the target diffusion model according to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function.

[0182] Optionally, the second obtaining module 1007 can specifically also be used for:

[0183] Backpropagate the gradient to update the parameters of the acceleration module according to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function;

[0184] Train to obtain the target diffusion model according to the acceleration module with updated parameters.

[0185] Optionally, after performing the second-stage alignment process on the second-stage alignment system according to the preprocessed sample data to obtain the target diffusion model, it further includes:

[0186] A third obtaining module, configured to obtain the text data to be processed;

[0187] A generation module, configured to input the text data into the target diffusion model to determine the corresponding image data according to the output result of the target diffusion model.

[0188] The text-to-image generation device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0189] Figure 11 It is a schematic structural diagram of the text-to-image generation device provided in this application. As Figure 11As shown in the figure, the text-to-image generation device provided in this embodiment includes at least one processor 1101 and a memory 1102. Optionally, the text-to-image generation device further includes a communication component 1103. Among them, the processor 1101, the memory 1102, and the communication component 1103 are connected through a bus 1104.

[0190] In a specific implementation process, at least one processor 1101 executes the computer-executable instructions stored in the memory 1102, so that at least one processor 1101 executes the above-mentioned method.

[0191] For the specific implementation process of the processor 1101, reference can be made to the above method embodiment. The implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.

[0192] In the above embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), and may also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application-specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0193] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0194] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0195] This application also provides a computer program product, including a computer program, which implements the above-mentioned method when executed by a processor.

[0196] The present application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above method.

[0197] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0198] An exemplary readable storage medium is coupled to the processor so that the processor can read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0199] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0200] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0201] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0202] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, and other various media that can store program codes.

[0203] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: ROMs, RAMs, magnetic disks, or optical discs, and other various media that can store program codes.

[0204] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will easily think of other implementation manners of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, and these variations, uses, or adaptations follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A text-to-image generation method, characterized in that, Including: Obtain a first diffusion model and a first self-supervised model corresponding to a text-to-image generation task, and obtain a first mapper; Obtain preprocessed sample data; Establish a first-stage alignment system according to the first diffusion model, the first self-supervised model, and the first mapper; Perform first-stage alignment processing on the first-stage alignment system according to the preprocessed sample data to obtain a second self-supervised model and a second mapper; Mount an acceleration module on the first diffusion model and establish a relationship alignment module; Establish a second-stage alignment system according to the first diffusion model with the acceleration module mounted, the second self-supervised model, the second mapper, and the relationship alignment module; Perform second-stage alignment processing on the second-stage alignment system according to the preprocessed sample data to obtain a target diffusion model, where the target diffusion model is used to implement text-to-image generation.

2. The method according to claim 1, wherein The preprocessed sample data includes preprocessed image data and preprocessed text data corresponding to the preprocessed image data.

3. The method according to claim 2, characterized in that, The performing first-stage alignment processing on the first-stage alignment system according to the preprocessed sample data to obtain a second self-supervised model and a second mapper includes: In the first-stage alignment system, freeze the parameters of the first diffusion model; Input the preprocessed text data into the first diffusion model to obtain a first result through the first diffusion model; Input the first result into the first mapper for training to obtain first characterization data; Input the preprocessed image data into the first self-supervised model to obtain second characterization data through the first self-supervised model, and obtain the self-supervised loss function of the first self-supervised model; Determine a first feature alignment loss function according to the first characterization data and the second characterization data; Determine a second self-supervised model and a second mapper according to the self-supervised loss function and the first feature alignment loss function.

4. The method according to claim 3, wherein The determining a second self-supervised model and a second mapper according to the self-supervised loss function and the first feature alignment loss function includes: Determine a first-stage total loss function according to the self-supervised loss function and the first feature alignment loss function; Update the first self-supervised model and the first mapper through backpropagation according to the first-stage total loss function to obtain a second self-supervised model and the second mapper.

5. The method according to any one of claims 2 to 4, characterized in that, The performing second-stage alignment processing on the second-stage alignment system according to the preprocessed sample data to obtain a target diffusion model includes: In the second-stage alignment system, freeze the parameters of the second self-supervised model and the first diffusion model; Input the preprocessed text data into the first diffusion model with the acceleration module mounted to obtain a second result through the first diffusion model with the acceleration module mounted; Input the second result into the second mapper for training to obtain third characterization data; Input the preprocessed image data into the second self-supervised model to obtain fourth representation data through the second self-supervised model, and obtain the self-supervised loss function of the second self-supervised model; Obtain the distillation loss function of the first diffusion model with the acceleration module mounted; Determine a second feature alignment loss function according to the third representation data and the fourth representation data; Obtain a relationship alignment loss function according to the relationship alignment module; Determine a target diffusion model according to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function.

6. The method according to claim 5, characterized in that, The determining the target diffusion model according to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function includes: Update the parameters of the acceleration module by backpropagating the gradient according to the distillation loss function, the second feature alignment loss function, and the relationship alignment loss function; Train to obtain the target diffusion model according to the acceleration module with updated parameters.

7. The method according to any one of claims 1 to 4, characterized in that, After obtaining the target diffusion model by performing second-stage alignment processing on the second-stage alignment system according to the preprocessed sample data, it further includes: Obtain text data to be processed; Input the text data into the target diffusion model to determine corresponding image data through the output result of the target diffusion model.

8. A text-to-image generation device, characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the text-to-image generation method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the text-to-image generation method according to any one of claims 1 to 7.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the text-to-image generation method according to any one of claims 1 to 7.