Figure graph and model training method and device, electronic equipment and storage medium

By introducing the expert LoRA network and expert control network, using single-concept and multi-concept graph-text training sets, predicting expert activation information and noise parameters, the accuracy problem of the text-based graph model in multi-concept generation is solved, and more efficient and flexible image generation is achieved.

CN120807674APending Publication Date: 2025-10-17NINGXIA WISDOM HOUSE CULTURE & MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510749155.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing text-based graph models have problems such as mismatch between generated results and concepts, insufficient concept learning, and confusion when generating multiple concepts, resulting in insufficient generation accuracy.

Method used

By introducing the expert LoRA network and expert control network, and using the single-concept and multi-concept text-graph training sets, the expert activation information and noise are predicted respectively, and the parameters are adjusted to generate the target text-graph model.

Benefits of technology

The accuracy of the text graph model in multi-concept generation is improved, supporting flexible generation of single and multi-concepts, adapting to various application scenarios, and achieving more efficient training and automated activation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807674A_ABST
    Figure CN120807674A_ABST
Patent Text Reader

Abstract

The invention relates to a text graph and model training method and device, electronic equipment and a storage medium. In order to solve the problems that in the prior art, when multiple concepts are generated, a generation result is mismatched with the concepts, and concept learning is insufficient, the invention provides a new text graph model training scheme and a corresponding device, electronic equipment and a storage medium. A plurality of expert LoRA networks are introduced, each network corresponds to one concept, the activation probability of each expert network is predicted in combination with an expert control network, and accurate generation of a multi-concept image is achieved. During training, a single-concept image-text pair and a multi-concept image-text pair are respectively input, prediction noise and expert activation information are calculated, and model parameters are adjusted. During generation, expert activation information is predicted according to an input text, noise reduction processing is performed by combining expert network output and noise reduction network output, and a target image is generated. According to the method, the generation accuracy and flexibility of the text graph model in the multi-concept scene are improved, and a new thought and method are provided for the development of the text graph technology.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a text-to-image method and model training method and device, electronic equipment and storage medium. BACKGROUND

[0002] Text-to-image refers to generating a corresponding picture according to a text. A latent diffusion model (SD) is one of the important models for text-to-image in related technologies, and can perform an image generation task based on a textual description. In a multi-concept scenario, the SD model is required to support single generation and joint generation of different new concepts. Currently, a common method is to fine-tune the SD model using Low Rank Adaptation (LoRA), and fine-tune new concepts into LoRA through an auxiliary branch. However, the existing method has problems such as mismatch between generation results and concepts, insufficient concept learning, and confusion in multi-concept generation.

[0003] Patent CN118155023B proposes a text-to-image model training method, which realizes the generation of multi-concept images by introducing an expert LoRA network and an expert control network, but there is still room for improvement.

[0004] It can be seen that the existing text-to-image method and model training method, device, electronic equipment and storage medium have the above-mentioned problems, and there is still room for improvement. In order to solve the problems of the existing text-to-image method and model training method, device, electronic equipment and storage medium, relevant manufacturers have made great efforts to seek solutions, but for a long time no suitable design has been developed. The general method, manufacturing method, processing method and text-to-image method, device, electronic equipment and storage medium do not have a suitable method to solve the above problems, which is obviously a problem that relevant manufacturers are eager to solve.

[0005] In view of the defects of the existing text-to-image method and model training method, device, electronic equipment and storage medium, the present application is based on the rich practical experience and professional knowledge accumulated from many years of designing and manufacturing such products, and cooperates with the application of theory, actively researches and innovates, in order to create a new text-to-image method and model training method, device, electronic equipment and storage medium, which can improve the general existing text-to-image method and model training method, device, electronic equipment and storage medium, and make it more practical. After continuous research, design, and repeated trial and improvement, the present application with practical value is finally created. SUMMARY

[0006] The main purpose of the present application is to overcome the defects existing in the prior art, provide a new text-to-image and model training method, device, electronic equipment and storage medium, and solve the technical problem of further improving the accuracy of the text-to-image model output, realizing more accurate generation of multi-concept images, thereby being more suitable for practical use and having industrial utilization value.

[0007] Another purpose of the present application is to provide a text-to-image and model training device, which solves the technical problem of more effectively training the text-to-image model, thereby being more suitable for practical use.

[0008] Still another purpose of the present application is to provide an electronic equipment, which solves the technical problem of more effectively running the text-to-image model training method, thereby being more suitable for practical use.

[0009] Still another purpose of the present application is to provide a storage medium, which solves the technical problem of more effectively storing the program code of the text-to-image model training method, thereby being more suitable for practical use.

[0010] The purpose of the present application and the solution to the technical problem are realized by adopting the following technical scheme. The text-to-image model training method according to the present application comprises the following steps:

[0011] Select single-concept image-text pairs and multi-concept image-text pairs from the image-text training set; wherein the single-concept image-text pairs comprise single-concept sample images containing one object and corresponding first description texts; the multi-concept image-text pairs comprise multi-concept sample images containing multiple objects and corresponding second description texts.

[0012] Input the single-concept image-text pairs and the multi-concept image-text pairs into a to-be-trained text-to-image model comprising multiple expert LoRA networks, obtain first expert activation information predicted based on the first description texts, and determine first prediction noise of the single-concept sample images according to the first expert activation information; and obtain second expert activation information predicted based on the second description texts, and determine second prediction noise of the multi-concept sample images according to the second expert activation information; wherein each expert LoRA network corresponds to one object; the expert activation information comprises the probability of each expert LoRA network being activated in the text-to-image process.

[0013] Based on the first prediction noise, the second prediction noise, the first expert activation information and the second expert activation information, the to-be-trained text-to-image model is adjusted in parameters to obtain a trained target text-to-image model; the target text-to-image model is used to generate a target image.

[0014] The purposes and technical problems of the present application can also be further achieved by the following technical measures. The text-to-image model training method described above, wherein the text-to-image method comprises:

[0015] Based on the preset mapping relationship, each object vocabulary in the to-be-generated text is replaced with an associated uniform marker, and then the initial noise graph and the to-be-generated text are input into a target text-to-image model comprising a plurality of expert LoRA networks; wherein the to-be-generated text is a description text for at least one object; each expert LoRA network corresponds to an object; and the target text-to-image model is obtained by training according to any one of the text-to-image model training methods described above.

[0016] Based on the to-be-generated text, the target expert activation information corresponding to each of the at least one object in the to-be-generated text is predicted; the target expert activation information comprises the probability of each expert LoRA network being activated in the current text-to-image process.

[0017] According to the target expert activation information, the initial noise graph is denoised to obtain a target image generated based on the to-be-generated text.

[0018] The purposes and technical problems of the present application are also achieved by the following technical solutions. The text-to-image model training device according to the present application comprises:

[0019] A sample selection unit is configured to select single-concept image-text pairs and multi-concept image-text pairs from a text-image pair training set; wherein the single-concept image-text pairs comprise single-concept sample images containing one object and corresponding first description texts; and the multi-concept image-text pairs comprise multi-concept sample images containing multiple objects and corresponding second description texts.

[0020] A prediction unit is configured to input the single-concept image-text pairs and the multi-concept image-text pairs into a to-be-trained text-to-image model comprising a plurality of expert LoRA networks, to obtain first expert activation information predicted based on the first description texts and to determine first predicted noise of the single-concept sample images based on the first expert activation information; and to obtain second expert activation information predicted based on the second description texts and to determine second predicted noise of the multi-concept sample images based on the second expert activation information; wherein each expert LoRA network corresponds to an object; and the expert activation information comprises the probability of each expert LoRA network being activated in the text-to-image process.

[0021] A parameter adjustment unit is configured to adjust parameters of the to-be-trained text-to-image model based on the first predicted noise, the second predicted noise, the first expert activation information and the second expert activation information, to obtain a trained target text-to-image model; and the target text-to-image model is used to generate a target image.

[0022] The purposes and solutions of the present application can also be further realized by the following technical measures.

[0023] The input unit is configured to receive a text to be generated input by a user.

[0024] The output unit is configured to output the generated target image.

[0025] The present application has obvious advantages and beneficial effects compared with the prior art.

[0026] A text-to-image model training method is proposed, which selects single concept image-text pairs and multiple concept image-text pairs, respectively inputs a to-be-trained text-to-image model including multiple expert LoRA networks, obtains expert activation information and prediction noise based on description text prediction, adjusts the parameters of the model, and thus obtains a trained target text-to-image model.

[0027] A text-to-image model training device is proposed, which includes a sample selection unit, a prediction unit and a parameter adjustment unit, and can effectively train a text-to-image model.

[0028] An electronic device is proposed, which can effectively run the text-to-image model training method.

[0029] A storage medium is proposed, which can effectively store the program code of the text-to-image model training method.

[0030] As known from the above, the present application proposes a text-to-image model training method, device, electronic device and storage medium, which improves the accuracy of the output of the text-to-image model through a unique training method and device, and realizes more accurate generation of multiple concept images.

[0031] A text-to-image model training method, the present application is realized through the following scheme, the method comprises:

[0032] Selecting single-concept image-text pairs and multi-concept image-text pairs from an image-text pair training set; wherein the single-concept image-text pairs include: a single-concept sample image containing one object and a corresponding first description text; the multi-concept image-text pairs include: a multi-concept sample image containing multiple objects and a corresponding second description text;

[0033] The single-concept image-text pair and the multi-concept image-text pair are respectively input into a to-be-trained text-based graph model comprising a plurality of expert LoRA networks, first expert activation information based on the prediction of the first descriptive text is obtained, and a first prediction noise of the single-concept sample image is determined based on the first expert activation information; and second expert activation information based on the prediction of the second descriptive text is obtained, and a second prediction noise of the multi-concept sample image is determined based on the second expert activation information; wherein each expert LoRA network corresponds to an object; and the expert activation information includes: the probability of each expert LoRA network being activated during the text-based graph process;

[0034] Based on the first predicted noise, the second predicted noise, the first expert activation information, and the second expert activation information, parameters of the to-be-trained culture graph model are adjusted to obtain a trained target culture graph model; the target culture graph model is used to generate a target image.

[0035] Preferably, the infrastructure of the text-graph model to be trained adopts a pre-trained diffusion model, and the text-graph model to be trained further includes an expert control network; for each selected single-concept graph-text pair, the corresponding first prediction noise is determined by the following method:

[0036] After replacing the object vocabulary contained in the first description text of the single concept image-text pair with a unified identifier associated with the corresponding object, the unified identifier is input into the text-image model to be trained;

[0037] Based on the text encoder in the diffusion model, extracting text features of the first descriptive text and inputting the extracted features into the expert control network to obtain the first expert activation information; and inputting the text features into each expert LoRA network and the noise reduction network in the diffusion model respectively;

[0038] According to the comprehensive output features obtained by weighted summing the output features of the expert LoRA networks based on the first expert activation information, and the output features of the downsampling attention network in the denoising network, the single-concept sample image in the single-concept image-text pair is denoised to obtain the first predicted noise.

[0039] Preferably, the base architecture of the text-to-graph model to be trained adopts a pre-trained diffusion model, and the text-to-graph model to be trained further comprises an expert control network; for each selected multi-concept graph-text pair, the corresponding second predicted noise is determined by the following manner:

[0040] After each object vocabulary contained in the second description text in the multi-concept graph-text pair is replaced by a corresponding object-associated uniform identifier, the text-to-graph model to be trained is inputted;

[0041] Based on the text encoder in the diffusion model, the text features of the second description text are extracted and inputted into the expert control network to obtain the second expert activation information; and the text features are respectively inputted into the expert LoRA network and the denoising network in the diffusion model;

[0042] According to the comprehensive output features obtained by weighting and summing the output features of the expert LoRA network based on the second expert activation information, and the output features of the down-sampling attention network in the denoising network, the multi-concept sample image in the multi-concept graph-text pair is denoised to obtain the second predicted noise.

[0043] Preferably, the text-to-graph model to be trained further comprises an expert control network for predicting expert activation information; and the parameter adjustment of the text-to-graph model to be trained based on the first predicted noise, the second predicted noise, the first expert activation information and the second expert activation information comprises:

[0044] For each selected single-concept graph-text pair, the parameters of the expert LoRA network and the expert control network are adjusted based on the difference between the first predicted noise corresponding to the single-concept graph-text pair and the actual added noise, and the difference between the first expert activation information and the corresponding actual expert activation information; the actual expert activation information corresponding to the single-concept graph-text pair indicates that only the expert network corresponding to the object in the single-concept graph-text pair is activated;

[0045] For each selected multi-concept graph-text pair, the parameters of the expert LoRA network and the expert control network are adjusted based on the difference between the second predicted noise corresponding to the multi-concept graph-text pair and the actual added noise, and the difference between the second expert activation information and the corresponding actual expert activation information; the actual expert activation information corresponding to the multi-concept graph-text pair indicates that the expert network corresponding to each object in the multi-concept graph-text pair is activated.

[0046] Preferably, the base architecture of the text-to-image model to be trained adopts a pre-trained diffusion model; the expert LoRA network is identical in structure to a down-sampling attention network in the diffusion model, and the parameters contained in the expert LoRA network are initialized based on the parameters of the corresponding down-sampling attention network.

[0047] Preferably, the text-image pair training set includes a plurality of single-concept text-image pair training subsets and multi-concept text-image pair training subsets; each single-concept text-image pair training subset corresponds to an object; and the selecting of the single-concept text-image pairs and the multi-concept text-image pairs corresponding to the plurality of objects respectively from the text-image pair training set comprises:

[0048] selecting at least one single-concept text-image pair from each single-concept text-image pair training subset, and selecting a plurality of multi-concept text-image pairs from the multi-concept text-image pair training subsets; the number of the selected multi-concept text-image pairs is positively correlated with the number of the expert LoRA networks.

[0049] The technical scheme of the present application also provides a text-to-image method, characterized in that the method comprises:

[0050] based on a preset mapping relationship, replacing each object vocabulary in the text to be generated with an associated uniform marker, inputting an initial noise graph and the text to be generated into a target text-to-image model containing a plurality of expert LoRA networks; wherein the text to be generated is a description text for at least one object; each expert LoRA network corresponds to an object; the target text-to-image model is trained according to the text-to-image model training method of any one of claims 1 to 6;

[0051] based on the text to be generated, predicting target expert activation information corresponding to the at least one object respectively in the text to be generated; the target expert activation information includes the probability of each expert LoRA network being activated in the current text-to-image process;

[0052] performing noise reduction processing on the initial noise graph according to the target expert activation information to obtain a target image generated based on the text to be generated.

[0053] Preferably, the base architecture of the target text-to-image model adopts a pre-trained diffusion model, and the target text-to-image model further includes an expert control network; the predicting of the target expert activation information corresponding to the at least one object respectively in the text to be generated based on the text to be generated comprises:

[0054] inputting the text features of the text to be generated into the expert control network to obtain the target expert activation information based on the text encoder in the diffusion model;

[0055] The denoising processing of the initial noise map according to the target expert activation information to obtain a target image generated based on the text to be generated comprises:

[0056] The text features are respectively input into the expert LoRA network and the denoising network in the diffusion model.

[0057] The denoising processing of the initial noise map according to the comprehensive output feature obtained by weighting and summing the output features of the expert LoRA networks based on the target expert activation information and the output feature of the down-sampling attention network in the denoising network obtains the target image.

[0058] Preferably, the expert LoRA network and the down-sampling attention network in the diffusion model have the same structure; the down-sampling attention network comprises a down-sampling attention first sub-network and a down-sampling attention second sub-network; each expert LoRA network comprises an expert LoRA first sub-network corresponding to the down-sampling attention first sub-network and an expert LoRA second sub-network corresponding to the down-sampling attention second sub-network.

[0059] Preferably, the denoising processing of the initial noise map according to the comprehensive output feature obtained by weighting and summing the output features of the expert LoRA networks based on the target expert activation information and the output feature of the down-sampling attention network in the denoising network obtains the target image, comprising:

[0060] Based on the target expert activation information, the output features of the expert LoRA first sub-network in each expert LoRA network are weighted and summed to obtain a first comprehensive output feature.

[0061] Based on the target expert activation information, the expert LoRA second sub-network in each expert LoRA network is weighted and summed to obtain a second comprehensive output feature.

[0062] The denoising processing of the initial noise map according to the first comprehensive output feature and the output feature of the down-sampling attention first sub-network and the second comprehensive output feature and the output feature of the down-sampling attention second sub-network obtains the target image.

[0063] In the technical scheme provided by the application, each expert LoRA network is responsible for processing the features of a specific object in the text-to-image model. The expert control network outputs the activation probability of each expert LoRA network, which reflects the importance of each expert network in the current text-to-image task. The purpose of weighting and summing is to weight the output features of the expert LoRA networks according to the activation probabilities to generate a comprehensive output feature that can better reflect the requirements of the current task.

[0064] Assume we have N expert LoRA networks, each with output features F i (i = 1, 2, …, N), and the activation probability of each expert LoRA network controlling the output is F i . The integrated output feature F 综合 after weighted summation can be represented as:

[0065]

[0066] Where:

[0067] F i is the output feature of the i-th expert LoRA network.

[0068] F i is the activation probability of the i-th expert LoRA network, and 0 ≤ p i ≤ 1.

[0069] F 综合 is the integrated output feature after weighted summation.

[0070] The actual operation steps are:

[0071] Step 1: Obtain the output features of expert LoRA networks

[0072] For each expert LoRA network, input the text features and obtain its output feature F i . These features are usually multi-dimensional, such as a feature vector or a feature map.

[0073] Step 2: Obtain the activation probability of the expert control network

[0074] Input the text features into the expert control network to obtain the activation probability p i of each expert LoRA network. These probability values reflect the importance of each expert network in the current task.

[0075] Step 3: Perform weighted summation

[0076] Multiply the output feature F i of each expert LoRA network by its corresponding activation probability p i , then add all the weighted features to obtain the integrated output feature F 综合 .

[0077] Assuming we use the PyTorch framework, the following is an example code to implement weighted summation:import torch# Assume there are 3 expert LoRA networks N = 3# Output features of each expert LoRA network (assuming they are two-dimensional feature vectors)

[0078] F1 = torch.tensor([1.0, 2.0])

[0079] F2 = torch.tensor([3.0, 4.0])

[0080] F3 = torch.tensor([5.0, 6.0]) # expert control network output activation probability p1 = 0.2

[0081] p2 = 0.3

[0082] p3 = 0.5 # store activation probabilities and features into list features = [F1, F2, F3]

[0083] probabilities = [p1, p2, p3] # perform weighted sum F_composite = sum(p*F for p, F in zip(probabilities, features)) print("composite output feature:", F_composite)

[0084] The specific method of weighted sum is to multiply the output feature of each expert LoRA network by its corresponding activation probability, and then add all the weighted features to get the composite output feature. This method can adjust the contribution of each expert LoRA network according to the prediction of the expert control network, so as to generate a more comprehensive feature that meets the current task requirements.

[0085] 11. A text-to-graph model training device, comprising:

[0086] a sample selection unit configured to select single-concept graph-text pairs and multi-concept graph-text pairs from a graph-text pair training set, wherein the single-concept graph-text pairs comprise single-concept sample images containing one object and corresponding first description texts, and the multi-concept graph-text pairs comprise multi-concept sample images containing multiple objects and corresponding second description texts;

[0087] a prediction unit configured to input the single-concept graph-text pairs and the multi-concept graph-text pairs into a to-be-trained text-to-graph model comprising a plurality of expert LoRA networks, respectively, to obtain first expert activation information predicted based on the first description texts and determine first prediction noise of the single-concept sample images according to the first expert activation information, and to obtain second expert activation information predicted based on the second description texts and determine second prediction noise of the multi-concept sample images according to the second expert activation information, wherein each expert LoRA network corresponds to an object, and the expert activation information comprises a probability that each expert LoRA network is activated during text-to-graph conversion;

[0088] a parameter adjustment unit configured to perform parameter adjustment on the to-be-trained text-to-image model based on the first predicted noise, the second predicted noise, the first expert activation information, and the second expert activation information, to obtain a trained target text-to-image model, wherein the target text-to-image model is configured to generate a target image.

[0089] The application further provides a text-to-image model training device, characterized in that the base architecture of the to-be-trained text-to-image model adopts a pre-trained diffusion model, and the to-be-trained text-to-image model further comprises an expert control network.

[0090] Preferably, the device further comprises an input unit configured to receive a user-inputted text to be generated, and an output unit configured to output a generated target image.

[0091] The application further provides an electronic device, characterized in that the electronic device comprises:

[0092] a processor;

[0093] a memory configured to store program instructions;

[0094] When the processor executes the program instructions, the text-to-image model training method according to any one of claims 1 to 6 is implemented.

[0095] Preferably, the electronic device further comprises an input device configured to receive a user-inputted text to be generated, and an output device configured to output a generated target image.

[0096] Preferably, the storage medium stores program codes, and the program codes, when executed by the processor, implement the text-to-image model training method according to any one of claims 1 to 6.

[0097] Preferably, the storage medium is a non-volatile storage medium, such as a flash memory, a hard disk, an optical disk, etc.

[0098] Advantages and beneficial effects of the technical scheme of the application:

[0099] 1. Higher accuracy: by introducing the expert LoRA network and the expert control network, the model can more accurately generate multi-concept images.

[0100] 2. Automatic activation: the model can automatically activate the corresponding expert LoRA network according to the input text without human intervention.

[0101] 3. Flexibility: supports single-concept and multi-concept generation, and is suitable for various application scenarios.

[0102] 4. High efficiency training: adopt phased training method, first single concept training, then multi-concept training, improve training efficiency.

[0103] The above description is only a summary of the technical solutions of the present application. In order to make the technical means of the present application clearer and can be implemented according to the content of the specification, the following will be described in detail with the preferred embodiments of the present application and with the help of the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0104] Figure 1 : flowchart of the text-to-image model training method;

[0105] Figure 2 : detailed flowchart of single concept text-to-image pair determining first predicted noise;

[0106] Figure 3 : detailed flowchart of multi-concept text-to-image pair determining second predicted noise;

[0107] Figure 4 : flowchart of the text-to-image method;

[0108] Figure 5 : structure diagram of the target text-to-image model. DETAILED DESCRIPTION

[0109] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined invention purpose, the following will be described in detail with the help of the accompanying drawings and the preferred embodiments, the specific implementation, method, steps, structure, features and effects of the text-to-image and model training method, device, electronic equipment and storage medium according to the present application, as follows.

[0110] Please refer to Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 The text-to-image and model training method, device, electronic equipment and storage medium of the preferred embodiment of the present application mainly include the following steps:

[0111] Embodiment 1

[0112] The present embodiment provides a text-to-image model training method of animal main graph:

[0113] 1.1 Data preparation

[0114] Text-to-image pair training set: prepare a training set containing single concept text-to-image pair and multi-concept text-to-image pair. Single concept text-to-image pair contains sample image of an object and its description text (first description text), and multi-concept text-to-image pair contains sample image of multiple objects and its description text (second description text).

[0115] Examples:

[0116] Single-concept image-text pair: an image containing a "cat", with the description text "a cute cat on the grass".

[0117] Multi-concept image-text pair: an image containing "cat" and "dog", with the description text "a cat and a dog playing in the park".

[0118] 1.2 Model Architecture

[0119] Base architecture: use a pre-trained diffusion model.

[0120] Expert LoRA network: set up an expert LoRA network for each object, with the same structure as the down-sampling attention network in the diffusion model, and the parameters are initialized based on the parameters of the down-sampling attention network.

[0121] Expert control network: used to predict expert activation information, i.e. the probability of each expert LoRA network being activated during the text-to-image process.

[0122] 1.3 Training Process

[0123] Single-concept image-text pair processing:

[0124] Replace the object words in the first description text with uniform symbols (e.g. "cat" replaced by "OBJ1").

[0125] Input the text-to-image model to be trained, and extract the text features through the text encoder of the diffusion model.

[0126] Input the text features into the expert control network to obtain the first expert activation information.

[0127] Input the text features into each expert LoRA network and the denoising network respectively.

[0128] According to the first expert activation information, weight the output features of each expert LoRA network and sum them up to get the comprehensive output features.

[0129] Combine the comprehensive output features with the output features of the down-sampling attention network in the denoising network to denoise the single-concept sample image and obtain the first predicted noise.

[0130] Multi-concept image-text pair processing:

[0131] Replace each object word in the second description text with a uniform symbol (e.g. "cat" replaced by "OBJ1" and "dog" replaced by "OBJ2").

[0132] Input the text-to-image model to be trained, and extract the text features through the text encoder of the diffusion model.

[0133] Input the text features into the expert control network to obtain second expert activation information.

[0134] Input the text features into each expert LoRA network and the denoising network respectively.

[0135] Weighted sum the output features of each expert LoRA network according to the second expert activation information to obtain comprehensive output features.

[0136] Combine the comprehensive output features with the output features of the down-sampling attention network in the denoising network to perform denoising processing on the multi-concept sample image to obtain second predicted noise.

[0137] Parameter adjustment:

[0138] For single-concept image-text pairs, based on the difference between the first predicted noise and the actual added noise, and the difference between the first expert activation information and the corresponding actual expert activation information, the parameters of each expert LoRA network and the expert control network are adjusted.

[0139] For multi-concept image-text pairs, based on the difference between the second predicted noise and the actual added noise, and the difference between the second expert activation information and the corresponding actual expert activation information, the parameters of each expert LoRA network and the expert control network are adjusted.

[0140] 1.4 Training results

[0141] After the above training process, a trained target text-to-image model is obtained, which can generate high-quality target images according to input text.

[0142] 2. Text-to-image method

[0143] 2.1 Input preparation

[0144] Text to be generated: Prepare a text describing the target image, such as "a cat on the grass".

[0145] Initial noise map: Generate a random initial noise map as the starting point for image generation.

[0146] 2.2 Text preprocessing

[0147] Based on the preset mapping relationship, replace the object words in the text to be generated with uniform symbols (e.g., "cat" is replaced with "OBJ1").

[0148] 2.3 Model input

[0149] Input the preprocessed text and initial noise map into the target text-to-image model.

[0150] 2.4 Generation process

[0151] Text feature extraction: Extract text features through the text encoder of the diffusion model.

[0152] Expert activation information prediction: Input the text features into the expert control network to obtain the target expert activation information.

[0153] Feature processing:

[0154] Input the text features into each expert LoRA network and the denoising network.

[0155] According to the target expert activation information, the output features of each expert LoRA network are weighted and summed to obtain the comprehensive output features.

[0156] Combine the comprehensive output features with the output features of the down-sampling attention network in the denoising network.

[0157] 2.5 Image generation

[0158] Denoising the initial noise map to generate the target image.

[0159] 2.6 Output results

[0160] Output the generated target image, for example, generate an image containing a cat on the grass according to the input text "a cat on the grass".

[0161] The training process of the above embodiment includes:

[0162] Single concept image-text pair processing:

[0163] Weighted sum: Assuming that the expert control network predicts that the expert LoRA network corresponding to "cat" has an activation probability of p1=0.9. The expert LoRA network output feature is F1.

[0164] Comprehensive output feature F 综合 The calculation is as follows:

[0165] F 综合 = p1·F1 = 0.9·F1

[0166] This comprehensive output feature will be combined with the output feature of the down-sampling attention network in the denoising network to denoise the single concept sample image.

[0167] Multi-concept image-text pair processing:

[0168] Weighted sum:

[0169] Assuming that the expert control network predicts that the expert LoRA networks corresponding to "cat" and "dog" have activation probabilities of p1=0.8 and p2=0.7, respectively.

[0170] The expert LoRA network output features are F1 and F2, respectively.

[0171] The comprehensive output feature F is calculated as follows:

[0172] F 综合 =p1·F1+p2·F2=0.8·F1+0.7·F2

[0173] This comprehensive output feature will be combined with the output feature of the down-sampling attention network in the denoising network to perform denoising on multi-concept sample images.

[0174] 4. Wenshengtu Method

[0175] Generation process:

[0176] Weighted sum:

[0177] Assume that the expert control network predicts that the activation probability of the expert LoRA network corresponding to "cat" is p1 = 0.9. The output feature of the expert LoRA network is F1.

[0178] Comprehensive output feature F 综合 The calculation is as follows:

[0179] F 综合 =p1·F1=0.9·F1

[0180] This comprehensive output feature will be combined with the output feature of the downsampled attention network in the denoising network to denoise the initial noise map and finally generate the target image.

[0181] Example 2:

[0182] This embodiment provides a natural landscape theme cultural image model training and application:

[0183] 1. Data Preparation

[0184] Image-text pair training set:

[0185] Single-concept image-text pair: contains an image of a mountain with the description text "a majestic mountain in the clouds".

[0186] Multi-concept image-text pair: contains an image of a mountain and a lake, and the description text is "A mountain and a lake in the morning sun."

[0187] 2. Model Architecture

[0188] Infrastructure: Use pre-trained diffusion model.

[0189] Expert LoRA network: Set up one expert LoRA network for "mountain" and one for "lake".

[0190] Expert control network: used to predict the activation probability of each expert LoRA network.

[0191] 3. Training process

[0192] Single concept image-text pair processing:

[0193] 1. Replace "mountain" with "OBJ1".

[0194] 2. Input model, extract text features.

[0195] 3. Expert control network predicts activation probability, assuming that the activation probability of the expert LoRA network corresponding to "OBJ1" is 0.85.

[0196] 4. Weighted sum to get integrated output features.

[0197] 5. Combine the output of the denoising network to denoise the image and get the first predicted noise.

[0198] Multi-concept image-text pair processing:

[0199] 1. Replace "mountain" with "OBJ1" and "lake" with "OBJ2".

[0200] 2. Input model, extract text features.

[0201] 3. Expert control network predicts activation probability, assuming that the activation probability of the expert LoRA network corresponding to "OBJ1" and "OBJ2" is 0.8 and 0.75 respectively.

[0202] 4. Weighted sum to get integrated output features.

[0203] 5. Combine the output of the denoising network to denoise the image and get the second predicted noise.

[0204] Parameter adjustment:

[0205] 1. According to the difference between the predicted noise and the actual noise, adjust the parameters of the expert LoRA network and the expert control network.

[0206] 4. Text-to-image method

[0207] Input:

[0208] Text to be generated: "A mountain and a lake in the morning sun".

[0209] Initial noise map: randomly generated.

[0210] Preprocessing:

[0211] Replace "mountain" with "OBJ1" and "lake" with "OBJ2".

[0212] Model input:

[0213] Input pre-processed text and initial noise map.

[0214] Generation process:

[0215] Extract text features, predict expert activation information.

[0216] Weighted sum to get integrated output features, combined with denoising network output.

[0217] Output:

[0218] Generate an image containing a mountain and a lake under the morning sun.

[0219] The training process of the above example includes:

[0220] Single concept image-text pair processing:

[0221] Weighted sum:

[0222] Assume that the expert control network predicts the expert LoRA network activation probability corresponding to "mountain" as p1 = 0.85.

[0223] Expert LoRA network output feature F1.

[0224] Integrated output feature calculation:

[0225] F 综合 = p1·F1 = 0.85·F1

[0226] This integrated output feature will be combined with the output feature of the down-sampling attention network in the denoising network to perform denoising processing on the single concept sample image.

[0227] Multi-concept image-text pair processing:

[0228] Weighted sum:

[0229] Assume that the expert control network predicts the expert LoRA network activation probability corresponding to "mountain" and "lake" as p1 = 0.8 and p2 = 0.75, respectively.

[0230] Expert LoRA network output features F1 and F2, respectively.

[0231] Integrated output feature F 综合 Calculation:

[0232] F 综合 = p1·F1 + p2·F2 = 0.8·F1 + 0.75·F2

[0233] This integrated output feature will be combined with the output feature of the down-sampling attention network in the denoising network to perform denoising processing on the multi-concept sample image.

[0234] 4. Image-to-text method

[0235] Generation process:

[0236] Weighted sum:

[0237] Assume that the expert control network predicts the expert LoRA network activation probabilities for "mountain" and "lake" to be p1 = 0.8 and p2 = 0.75, respectively.

[0238] The expert LoRA network output features are F1 and F2, respectively.

[0239] Integrated output feature F 综合 The calculation is as follows:

[0240] F 综合 = p1·F1 + p2·F2 = 0.8·F1 + 0.75·F2

[0241] This integrated output feature will be combined with the output feature of the down-sampling attention network in the denoising network to perform denoising processing on the initial noise map, and finally generate the target image.

[0242] Example 3:

[0243] This example provides a city landscape theme image-to-text model training and application:

[0244] 1. Data preparation

[0245] Image-text pair training set:

[0246] Single-concept image-text pair: contains an image with only high-rise buildings, and the description text is "a modern high-rise building in the city center."

[0247] Multi-concept image-text pair: contains an image with high-rise buildings and bridges, and the description text is "a high-rise building and a bridge by the river."

[0248] 2. Model architecture

[0249] Basic architecture: use a pre-trained diffusion model.

[0250] Expert LoRA network: set up one expert LoRA network for "high-rise building" and "bridge" respectively.

[0251] Expert control network: used to predict the activation probability of each expert LoRA network.

[0252] 3. Training process

[0253] Single-concept image-text pair processing:

[0254] Replace "high-rise building" with "OBJ1".

[0255] Input the model and extract text features.

[0256] The expert control network predicts the activation probability, assuming that the activation probability of the expert LoRA network corresponding to "OBJ1" is 0.9.

[0257] The weighted sum is obtained as the comprehensive output feature.

[0258] Combine the output of the denoising network to denoise the image and obtain the first predicted noise.

[0259] Multi-concept image-text pair processing:

[0260] Replace "high-rise building" with "OBJ1" and "bridge" with "OBJ2".

[0261] Input the model and extract text features.

[0262] The expert control network predicts the activation probability, assuming that the activation probability of the expert LoRA network corresponding to "OBJ1" and "OBJ2" is 0.85 and 0.8, respectively.

[0263] The weighted sum is obtained as the comprehensive output feature.

[0264] Combine the output of the denoising network to denoise the image and obtain the second predicted noise.

[0265] Parameter adjustment:

[0266] Adjust the parameters of the expert LoRA network and the expert control network according to the difference between the predicted noise and the actual noise.

[0267] 4. Text-to-image method

[0268] Input:

[0269] Text to be generated: "A high-rise building and a bridge by the river."

[0270] Initial noise map: randomly generated.

[0271] Preprocessing:

[0272] Replace "high-rise building" with "OBJ1" and "bridge" with "OBJ2".

[0273] Model input:

[0274] Input the preprocessed text and the initial noise map.

[0275] Generation process:

[0276] Extract text features and predict expert activation information.

[0277] The weighted sum is obtained as the comprehensive output feature, combined with the output of the denoising network.

[0278] Output:

[0279] Generate an image containing a tall building and a bridge on the riverbank.

[0280] The training process of the above example includes:

[0281] Single-concept image-text pair processing:

[0282] Weighted summation:

[0283] Assume that the expert control network predicts the expert LoRA network activation probability corresponding to "tall building" as p1 = 0.9.

[0284] The expert LoRA network output feature is F1. The integrated output feature F 综合 is calculated as follows:

[0285] F 综合 = p1 · F1 = 0.9 · F1

[0286] This integrated output feature will be combined with the output feature of the down-sampling attention network in the denoising network to perform denoising processing on the single-concept sample image.

[0287] Multi-concept image-text pair processing:

[0288] Weighted summation:

[0289] Assume that the expert control network predicts the expert LoRA network activation probability corresponding to "tall building" and "bridge" as p1 = 0.85 and p2 = 0.8, respectively.

[0290] The expert LoRA network output features are F1 and F2, respectively. The integrated output feature F 综合 is calculated as follows:

[0291] F 综合 = p1 · F1 + p2 · F2 = 0.85 · F1 + 0.8 · F2

[0292] This integrated output feature will be combined with the output feature of the down-sampling attention network in the denoising network to perform denoising processing on the multi-concept sample image.

[0293] 4. Text-to-image generation method

[0294] Generation process:

[0295] Weighted summation:

[0296] Assume that the expert control network predicts the expert LoRA network activation probability corresponding to "tall building" and "bridge" as p1 = 0.85 and p2 = 0.8, respectively.

[0297] The expert LoRA network output features are F1 and F2, respectively.

[0298] The comprehensive output feature Fcomprehensive is calculated as follows:

[0299] F 综合 = p1·F1+ p2·F2= 0.85·F1+ 0.8·F2

[0300] The comprehensive output feature will be combined with the output feature of the down-sampling attention network in the noise reduction network to perform noise reduction processing on the initial noise map, and finally generate a target image.

[0301] The above is only a preferred embodiment of the present application, and is not intended to limit the present application in any form. Although the present application has been disclosed as above with a preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed methods and technical contents to make equivalent embodiments with equivalent changes, without departing from the technical solution of the present application. Any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application shall still fall within the scope of the technical solution of the present application.

Claims

1. A method for training a cultural graph model, characterized in that: The method comprises: Selecting single-concept image-text pairs and multi-concept image-text pairs from an image-text pair training set; wherein the single-concept image-text pairs include: a single-concept sample image containing one object and a corresponding first description text; the multi-concept image-text pairs include: a multi-concept sample image containing multiple objects and a corresponding second description text; The single-concept image-text pair and the multi-concept image-text pair are respectively input into a to-be-trained text-based graph model comprising a plurality of expert LoRA networks, first expert activation information based on the prediction of the first descriptive text is obtained, and a first prediction noise of the single-concept sample image is determined based on the first expert activation information; and second expert activation information based on the prediction of the second descriptive text is obtained, and a second prediction noise of the multi-concept sample image is determined based on the second expert activation information; wherein each expert LoRA network corresponds to an object; and the expert activation information includes: the probability of each expert LoRA network being activated during the text-based graph process; Based on the first predicted noise, the second predicted noise, the first expert activation information, and the second expert activation information, parameters of the to-be-trained culture graph model are adjusted to obtain a trained target culture graph model; the target culture graph model is used to generate a target image.

2. The method for training a cultural graph model according to claim 1, wherein: The infrastructure of the text-graph model to be trained adopts a pre-trained diffusion model, and the text-graph model to be trained further includes an expert control network. For each selected single-concept graph-text pair, the corresponding first prediction noise is determined by the following method: After replacing the object vocabulary contained in the first description text of the single concept image-text pair with a unified identifier associated with the corresponding object, the unified identifier is input into the text-image model to be trained; Based on the text encoder in the diffusion model, extracting text features of the first description text, and inputting the features into the expert control network to obtain the first expert activation information; and inputting the text features into the expert LoRA networks and the denoising network in the diffusion model respectively; According to the comprehensive output features obtained by weighted summing the output features of the expert LoRA networks based on the first expert activation information, and the output features of the downsampling attention network in the denoising network, the single-concept sample image in the single-concept image-text pair is denoised to obtain the first predicted noise.

3. The method for training a cultural graph model according to claim 1, wherein: The infrastructure of the text-graph model to be trained adopts a pre-trained diffusion model, and the text-graph model to be trained further includes an expert control network. For each selected multi-concept text-graph pair, the corresponding second prediction noise is determined by the following method: After replacing each object word contained in the second description text of the multi-concept image-text pair with a unified identifier associated with the corresponding object, the resultant word is input into the to-be-trained text-image model; Based on the text encoder in the diffusion model, extracting the text features of the second description text, and inputting the features into the expert control network to obtain the second expert activation information; and inputting the text features into the expert LoRA networks and the noise reduction network in the diffusion model respectively; According to the comprehensive output features obtained by weighted summing the output features of the expert LoRA networks based on the second expert activation information, and the output features of the downsampling attention network in the denoising network, the multi-concept sample images in the multi-concept image-text pair are denoised to obtain the second predicted noise.

4. The method for training a cultural graph model according to claim 1, wherein: The to-be-trained cultural graph model further includes an expert control network for predicting expert activation information; and the parameter adjustment of the to-be-trained cultural graph model based on the first prediction noise, the second prediction noise, the first expert activation information, and the second expert activation information includes: For each selected single-concept image-text pair, parameters of the expert LoRA networks and the expert control network are adjusted based on the difference between the first predicted noise corresponding to the single-concept image-text pair and the actual added noise, and the difference between the first expert activation information and the corresponding actual expert activation information; the actual expert activation information corresponding to the single-concept image-text pair indicates that only the expert network corresponding to the object in the single-concept image-text pair is activated; For each selected multi-concept graph-text pair, parameters of the expert LoRA networks and the expert control network are adjusted based on the difference between the second predicted noise corresponding to the multi-concept graph-text pair and the actual added noise, as well as the difference between the second expert activation information and the corresponding actual expert activation information; the actual expert activation information corresponding to the multi-concept graph-text pair indicates: activating the expert network corresponding to each object in the multi-concept graph-text pair.

5. The method for training a cultural graph model according to any one of claims 1 to 4, wherein: The infrastructure of the text graph model to be trained adopts a pre-trained diffusion model; the expert LoRA network has the same structure as the downsampling attention network in the diffusion model, and the parameters contained in the expert LoRA network are initialized based on the parameters of the corresponding downsampling attention network.

6. The method for training a cultural graph model according to any one of claims 1 to 4, wherein: The image-text pair training set includes multiple single-concept image-text pair training subsets and multiple-concept image-text pair training subsets; each single-concept image-text pair training subset corresponds to one object; The step of selecting single-concept image-text pairs and multi-concept image-text pairs corresponding to multiple objects from the image-text pair training set includes: Select at least one single-concept picture-text pair from each single-concept picture-text pair training subset; And, selecting a plurality of multi-concept image-text pairs from the multi-concept image-text pair training subset; the number of the selected multi-concept image-text pairs is positively correlated with the number of expert LoRA networks.

7. A method for generating a Wensheng diagram, characterized in that: The method comprises: After replacing each object word in the to-be-generated text with an associated unified identifier based on a preset mapping relationship, the initial noise map and the to-be-generated text are input into a target text graph model comprising multiple expert LoRA networks; wherein the to-be-generated text is a descriptive text for at least one object; each expert LoRA network corresponds to one object; and the target text graph model is trained using the text graph model training method according to any one of claims 1 to 6; Based on the text to be generated, predicting target expert activation information corresponding to each of the at least one object in the text to be generated; the target expert activation information includes: the probability of each expert LoRA network being activated during this text generation process; The initial noise image is subjected to denoising processing according to the target expert activation information to obtain a target image generated based on the text to be generated.

8. The method of claim 7, wherein: The target text graph model has a pre-trained diffusion model as its infrastructure, and further includes an expert control network. The target text graph model predicts target expert activation information corresponding to at least one object in the text to be generated based on the text to be generated, including: Based on the text encoder in the diffusion model, after extracting the text features of the to-be-generated text, the features are input into the expert control network to obtain the target expert activation information; Then, performing denoising on the initial noise image according to the target expert activation information to obtain a target image generated based on the text to be generated includes: Inputting the text features into the expert LoRA networks and the noise reduction network in the diffusion model respectively; The initial noise map is denoised according to the comprehensive output features obtained by weighted summing the output features of the expert LoRA networks based on the target expert activation information and the output features of the down-sampling attention network in the denoising network to obtain the target image.

9. The method of claim 8, wherein: The expert LoRA network has the same structure as the downsampling attention network in the diffusion model; the downsampling attention network includes a downsampling attention first subnetwork and a downsampling attention second subnetwork; each expert LoRA network includes an expert LoRA first subnetwork corresponding to the downsampling attention first subnetwork, and an expert LoRA second subnetwork corresponding to the downsampling attention second subnetwork.

10. The method according to claim 9, characterized in that: The method further comprises: performing denoising on the initial noise map based on the comprehensive output feature obtained by weighted summing the output features of the expert LoRA networks based on the target expert activation information and the output feature of the down-sampling attention network in the denoising network to obtain the target image, including: Based on the target expert activation information, the output features of the expert LoRA first sub-network in each expert LoRA network are weighted summed to obtain a first comprehensive output feature; Based on the target expert activation information, weighted summing is performed on the expert LoRA second subnetwork in each expert LoRA network to obtain a second comprehensive output feature; Combining the first comprehensive output feature with the output feature of the downsampling attention first sub-network, and the second comprehensive output feature with the output feature of the downsampling attention second sub-network, the initial noise map is denoised to obtain the target image.

11. A device for training a cultural graph model, characterized in that: The device comprises: A sample selection unit is used to select single-concept image-text pairs and multi-concept image-text pairs from the image-text pair training set; wherein the single-concept image-text pairs include: a single-concept sample image containing one object and a corresponding first description text; the multi-concept image-text pairs include: a multi-concept sample image containing multiple objects and a corresponding second description text; A prediction unit is configured to input the single-concept image-text pair and the multi-concept image-text pair into a to-be-trained text-based graph model comprising a plurality of expert LoRA networks, obtain first expert activation information based on the prediction of the first descriptive text, and determine a first predicted noise for the single-concept sample image according to the first expert activation information; and obtain second expert activation information based on the prediction of the second descriptive text, and determine a second predicted noise for the multi-concept sample image according to the second expert activation information; wherein each expert LoRA network corresponds to one object; and the expert activation information includes: a probability of each expert LoRA network being activated during the text-based graph process; A parameter adjustment unit is configured to adjust parameters of the to-be-trained culture graph model based on the first predicted noise, the second predicted noise, the first expert activation information, and the second expert activation information to obtain a trained target culture graph model; the target culture graph model is used to generate a target image.

12. The cultural graph model training device according to claim 11, characterized in that: The infrastructure of the to-be-trained cultural graph model adopts a pre-trained diffusion model, and the to-be-trained cultural graph model further includes an expert control network.

13. The cultural graph model training device according to claim 11, characterized in that: The device further includes an input unit for receiving a text to be generated input by a user, and an output unit for outputting a generated target image.

14. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing program instructions; When the processor executes the program instructions, the method for training a cultural graph model according to any one of claims 1 to 6 is implemented.

15. The electronic device according to claim 14, characterized in that The electronic device further includes an input device and an output device, wherein the input device is used to receive the text to be generated input by a user, and the output device is used to output the generated target image.

16. A storage medium, characterized in that The storage medium stores program code, and when the program code is executed by the processor, it implements the text graph model training method according to any one of claims 1 to 6.

17. The storage medium according to claim 16, wherein: The storage medium is a non-volatile storage medium, such as a flash memory, a hard disk, an optical disk, etc.