Training of a multi-modal generative model, image generation method and apparatus

CN119540673BActive Publication Date: 2026-09-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311094824.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-29
Publication Date
2026-09-25
Estimated Expiration
2043-08-29

AI Technical Summary

Benefits of technology

[0033]根据本公开实施例提供的多模态生成模型的训练、图像生成方法及装置,第一对话样本包括第一样本描述、该第一样本描述对应的第二样本描述以及用于指示第二样本描述的生成方向的样本指令,基于第一样本描述通过多模态生成模型的图像生成网络,得到第一生成图像,接着基于第二样本描述和样本指令通过多模态生成模型的文本生成网络,得到第二样本描述按照样本指令指示的生成方向而生成的第一生成描述;接着,基于第一生成图像和第一生成描述,通过多模态生成模型的多模态生成网络,得到融合有多模态数据(即第一生成图像的特征和第一生成描述的特征)的第二生成描述,该第二生成描述与该多模态生成网络的输出相关;基于第二生成描述,生成第二生成图像;基于第一样本描述、多模态生成网络的输出以及第二生成图像,确定目标生成损失,并以最小化目标生成损失为目标,调整多模态生成模型的参数,以利用多模态学习的思路,调整该多模态生成模型,以提升该多模态生成模型生成质量更高的图像描述的能力,进而可以利用该多模态生成模型得到质量更高的图像描述,提升基于该图像描述所生成图像的质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540673B_ABST
    Figure CN119540673B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method and device for training a multi-modal generation model and generating an image. The method comprises: obtaining a first dialogue sample comprising a first sample description, a second sample description and a sample instruction; obtaining a first generated image based on the first sample description through an image generation network; obtaining a first generated description based on the second sample description and the sample instruction through a text generation network; obtaining a second generated description related to an output of a multi-modal generation network based on the first generated image and the first generated description; generating a second generated image based on the second generated description; determining a target generation loss based on the first sample description, the output of the multi-modal generation network and the second generated image; and adjusting parameters of the multi-modal generation model comprising the three networks to train a model capable of generating better image descriptions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method and apparatus for training a multimodal generative model and generating images. Background Technology

[0002] In the field of image generation, image descriptions (also known as text prompts) are typically used to generate images corresponding to the descriptions via content-based image generation networks (such as those based on the Diffusion Model). During image generation, improving the quality of the image description is crucial to enhancing the overall quality of the generated image.

[0003] Therefore, how to obtain better image descriptions has become an urgent problem to be solved. Summary of the Invention

[0004] This disclosure provides one or more embodiments of a multimodal generative model training, image generation method, and apparatus to train a model that can generate better image descriptions, thereby obtaining better quality image descriptions and ultimately better quality images.

[0005] According to a first aspect, a training method for a multimodal generative model is provided, wherein the multimodal generative model includes an image generation network, a text generation network, and a multimodal generative network, and the method includes:

[0006] Obtain any first dialogue sample from the training sample set. The first dialogue sample includes a first sample description and its corresponding second sample description and sample instructions. The sample instructions are used to indicate the generation direction of the second sample description.

[0007] Based on the first sample description, a first generated image is obtained through the image generation network;

[0008] Based on the second sample description and the sample instruction, a first generated description is obtained through the text generation network;

[0009] Based on the first generated image and the first generated description, a second generated description is obtained through the multimodal generation network, and the second generated description is related to the output of the multimodal generation network;

[0010] Based on the second generation description, a second generated image is generated;

[0011] Based on the first sample description, the output of the multimodal generation network, and the second generated image, the target generation loss is determined;

[0012] The parameters of the multimodal generation model are adjusted with the goal of minimizing the target generation loss.

[0013] According to the second aspect, an image generation method is provided, comprising:

[0014] Acquire the data to be processed and a first instruction for indicating the generation direction of the data to be processed;

[0015] Based on the first instruction, a third generated description is obtained through the target text generation network of the target multimodal generation model;

[0016] Based on the data to be processed and the third generated description, a fourth generated description is obtained through the target multimodal generation network of the target multimodal generation model;

[0017] Based on the fourth generation description, a generated image corresponding to the data to be processed and the first instruction is obtained through the target image generation network of the target multimodal generation model.

[0018] According to a third aspect, a training apparatus for a multimodal generative model is provided, the multimodal generative model including an image generation network, a text generation network, and a multimodal generative network, the apparatus comprising:

[0019] The first acquisition module is configured to acquire any first dialogue sample in the training sample set. The first dialogue sample includes a first sample description and its corresponding second sample description and sample instruction. The sample instruction is used to indicate the generation direction of the second sample description.

[0020] The first obtaining module is configured to obtain a first generated image based on the first sample description and through the image generation network;

[0021] The second obtaining module is configured to obtain a first generated description based on the second sample description and the sample instructions through the text generation network;

[0022] The third obtaining module is configured to obtain a second generating description based on the first generated image and the first generated description through the multimodal generation network, wherein the second generated description is related to the output of the multimodal generation network;

[0023] The generation module is configured to generate a second generated image based on the second generation description;

[0024] The first determining module is configured to determine the target generation loss based on the first sample description, the output of the multimodal generation network, and the second generated image;

[0025] The first adjustment module is configured to adjust the parameters of the multimodal generation model with the goal of minimizing the target generation loss.

[0026] According to a fourth aspect, an image generation apparatus is provided, comprising:

[0027] The second acquisition module is configured to acquire data to be processed and a first instruction for indicating the generation direction of the data to be processed.

[0028] The fourth module is configured to obtain a third generated description based on the first instruction, through the target text generation network of the target multimodal generation model.

[0029] The fifth module, based on the data to be processed and the third generation description, obtains the fourth generation description through the target multimodal generation network of the target multimodal generation model;

[0030] The sixth module, based on the fourth generation description, obtains a generated image corresponding to the data to be processed and the first instruction through the target image generation network of the target multimodal generation model.

[0031] According to a fifth aspect, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method described in the first or second aspect.

[0032] According to a sixth aspect, an electronic device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method described in the first or second aspect.

[0033] According to the multimodal generation model training, image generation method, and apparatus provided in this disclosure, a first dialogue sample includes a first sample description, a second sample description corresponding to the first sample description, and sample instructions for indicating the generation direction of the second sample description. Based on the first sample description, a first generated image is obtained through the image generation network of the multimodal generation model. Then, based on the second sample description and the sample instructions, a first generated description generated according to the generation direction indicated by the sample instructions is obtained through the text generation network of the multimodal generation model. Next, based on the first generated image and the first generated description, a fused multimodal image is obtained through the multimodal generation network of the multimodal generation model. The data (i.e., features of the first generated image and features of the first generated description) are used to generate a second generated description, which is related to the output of the multimodal generation network. Based on the second generated description, a second generated image is generated. Based on the first sample description, the output of the multimodal generation network, and the second generated image, a target generation loss is determined, and the parameters of the multimodal generation model are adjusted with the goal of minimizing the target generation loss. This utilizes the idea of ​​multimodal learning to adjust the multimodal generation model, thereby improving its ability to generate higher-quality image descriptions. As a result, higher-quality image descriptions can be obtained using the multimodal generation model, improving the quality of the image generated based on the image description. Attached Figure Description

[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0035] Figure 1 This is a schematic diagram illustrating the implementation framework of one embodiment disclosed in this disclosure;

[0036] Figure 2 A schematic flowchart illustrating a training method for a multimodal generative model provided in an embodiment;

[0037] Figure 3 A schematic flowchart of an image generation method provided in an embodiment;

[0038] Figure 4 A schematic block diagram of a training device for a multimodal generative model provided in an embodiment;

[0039] Figure 5 A schematic block diagram of an image generation apparatus provided in an embodiment;

[0040] Figure 6 This is a schematic block diagram of an electronic device provided for an embodiment. Detailed Implementation

[0041] The technical solutions of the embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0042] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0043] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0044] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0045] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0046] This disclosure presents a training method and apparatus for a multimodal generative model, as well as an image generation method and apparatus. The application scenarios and technical concepts of the method are first introduced below:

[0047] As mentioned earlier, in the field of image generation, image descriptions (also known as text prompts) are typically used to generate images corresponding to the descriptions through content-based image generation networks (such as those based on the Diffusion Model). During image generation, improving the quality of the image description is crucial to enhancing the overall quality of the generated image. Therefore, how to obtain higher-quality image descriptions is a pressing issue that needs to be addressed.

[0048] In view of this, embodiments of the present disclosure provide a method and apparatus for training a multimodal generative model and generating images, so as to obtain a better quality image description, and then obtain a better quality image based on the image description. Figure 1The illustration shows an implementation scenario according to an embodiment disclosed herein. In this implementation scenario, an exemplary multimodal generation model is shown. This multimodal generation model is a conversational multimodal generation model that can expand and optimize an image description (or image) based on user input indicating the generation direction of the image description (or image), obtaining a corresponding expanded and optimized image description, and further obtaining the corresponding image based on the expanded and optimized image description. Specifically, as shown... Figure 1 As shown, the multimodal generation model can include an image generation network, a text generation network, and a multimodal generation network.

[0049] The image generation network can be a pre-trained network using several image-text pairs. Each image-text pair includes an image and its corresponding image description. For clarity, the image in the image-text pair can be referred to as the second sample image, and the image description as the third sample description. In one implementation, the image-text pair can be obtained from any specified public dataset that provides a large number of image-text pairs or from a specified artificially created dataset. The artificially created dataset includes images obtained from specified sources and manually annotated image descriptions. The image generation network is trained using the image-text pairs so that the trained network has the ability to generate corresponding images based on the image descriptions.

[0050] This text generation network can be a pre-trained text generation network that utilizes individual texts and occluded texts generated from the occluded portions of those texts. These individual texts serve as image descriptions for generating the images; for clarity, they can be referred to as fourth-sample descriptions. This text generation network can be used to predict, expand, and complete the text descriptions of the input images.

[0051] In one implementation, the aforementioned image description can also be called a text prompt, that is, a text prompt used to generate the image.

[0052] The training process of the multimodal generative model is described below. Specifically, a training sample set is obtained for training the multimodal generative model. This training sample set may include multiple dialogue samples, where any first dialogue sample may include a first sample description, its corresponding second sample description, and sample instructions for indicating the generation direction of the second sample description. The second sample description may include a portion of the corresponding first sample description. For example, the second sample description may be generated based on the occluded portion of the corresponding first sample description, or the second sample description may include a portion of the corresponding first sample description that has been decomposed. For example, if the first sample description is "a highly detailed portrait of a man with dark green hair and green glowing eyes, high detail clothing, concept art, anime, artstation, professional", the corresponding second sample description could be "a highly detailed portrait of a man", or "a highly detailed portrait of a man with dark green hair and green glowing eyes, professional".

[0053] In one implementation, the first sample description can be a description that meets specified conditions. Specifically, the first sample description and its corresponding image (i.e., image-text pair) can be image-text pairs in a specified public dataset whose quality exceeds preset screening conditions. The preset screening conditions can include, but are not limited to, at least one of the following: the aesthetic score corresponding to the image in the image-text pair is not lower than a preset aesthetic score threshold (for example, when the aesthetic score ranges from [0, 1], the preset aesthetic score threshold is, for example, 0.75); the sharpness score corresponding to the image in the image-text pair is not lower than a preset sharpness threshold (for example, when the sharpness score ranges from [0, 1], the preset sharpness threshold is, for example, 0.80); and the matching degree value corresponding to the image-text pair is not lower than a preset matching degree threshold (for example, when the matching degree value ranges from [0, 1], the preset matching degree threshold is, for example, 0.80).

[0054] By using sample descriptions that meet specified conditions, the image description generation capability of the trained multimodal generative model can be better guaranteed, enabling it to generate higher quality image descriptions, and thus obtain higher quality images.

[0055] The system obtains arbitrary dialogue samples from the training sample set as the first dialogue sample. Based on the first sample description in the first dialogue sample, it generates a first generated image through an image generation network. Based on the second sample description and sample instructions, it generates a first generated description generated according to the generation direction indicated by the sample instructions through a text generation network. Then, based on the first generated image and the first generated description, it generates a second generated description through a multimodal generation network. This second generated model is related to the output of the multimodal generation model. Then, it generates a second generated image based on the second generated description. Finally, it combines the first sample description, the output of the multimodal generation model, and the second generated image to determine the target generation loss. The system then adjusts the parameters of the multimodal generation model with the goal of minimizing the target generation loss.

[0056] Specifically, a first generation loss can be determined based on the output of the multimodal generation model and the first sample description; and a first matching loss can be determined based on the second generated image and the first sample description through a text-image association model; then, at least the first generation loss and the first matching loss are combined to determine the target generation loss, where the target generation loss is positively correlated with both the first generation loss and the first matching loss. The parameters of the multimodal generation model are adjusted to minimize the target generation loss, i.e., to minimize both the first matching loss and the first generation loss, in order to improve the quality of the image descriptions generated by the multimodal generation model.

[0057] Minimizing the first generation loss allows the text generation network to generate higher-quality image descriptions. Minimizing the constructed first matching loss allows the second generated image, which integrates features from both the first generated image and the first generated description, to better match the first sample description. This leads to a better fusion of the image generation capabilities of the image generation network and the text generation capabilities of the text generation network, thereby better integrating image domain data and text (language) domain data. Ultimately, this improves the image description generation capability of the multimodal generation model and enhances the quality of its generated image descriptions (text).

[0058] In the above process, by using the idea of ​​multimodal learning, fine-grained loss is constructed to adjust the multimodal generation model, so as to improve the ability of the multimodal generation model to generate higher quality image descriptions. In this way, higher quality image descriptions can be obtained by using the multimodal generation model, thereby improving the quality of the images generated based on the image descriptions.

[0059] The multimodal generation model can be a conversational multimodal generation model. For a trained multimodal generation model, it can generate a high-quality image description based on the user-input data to be processed (the image to be processed and / or the image description to be processed) and the user-input instructions (dialogue content) indicating the generation direction of the data to be processed. Then, it can generate a high-quality image based on the high-quality image description.

[0060] The training method and apparatus for the multimodal generative model, as well as the image generation method and apparatus provided in this disclosure, will be described in detail below with reference to specific embodiments. For clarity, the processes of training the image generation network and the text generation network will be introduced first.

[0061] The process of training the image generation network can include: acquiring the image generation network to be trained; acquiring a training dataset S for pre-training the image generation network; and acquiring a training dataset S that can include multiple image-text pairs, where each image-text pair includes a sample image (referred to as the second sample image for clarity) T and its corresponding image description (referred to as the third sample description for clarity) X. In one scenario, the training dataset S can be obtained from any publicly available dataset that provides a large number of image-text pairs or from a specified artificially created dataset. This artificially created dataset includes images obtained from a specified source and image descriptions manually annotated for the images. In another scenario, the image description can also be called a text prompt, i.e., a text prompt used to generate the image.

[0062] From the training dataset S, obtain any image-text pair M, including a third sample description X and its corresponding second sample image T. Using the third sample description X, pass the image generation network to be trained to obtain the corresponding generated image, called the third generated image T'. Then, using the difference between the second sample image T in the image-text pair M and the third generated image T', determine the generation loss (called the fourth generation loss) corresponding to the image generation network to be trained. In one case, a preset loss function can be used to determine the fourth generation loss based on the difference between the second sample image T and the third generated image T'. This preset loss function may include, but is not limited to, the cross-entropy loss function and the MSE (mean squared error) loss function.

[0063] Subsequently, with the goal of minimizing the fourth generation loss (i.e., reducing the difference between the second sample image and the corresponding third generated image), the parameters of the image generation network to be trained are adjusted until the image generation network reaches the first convergence condition. This indicates that the training of the image generation network is complete, and the image generation network in the aforementioned multimodal generation model is obtained. The first convergence condition may include, but is not limited to, the number of parameter adjustments not being less than a first adjustment threshold, or the calculated fourth generation loss being less than a first loss threshold.

[0064] To recap the aforementioned process of training the image generation network, the above embodiment used an image-text pair (i.e., a pair of third sample descriptions and second sample images) as an example. In another embodiment, the above process can also be performed on a batch of samples, i.e., multiple image-text pairs, to obtain the third generated image corresponding to the third sample description in each image-text pair. Then, based on the difference between the third generated image corresponding to the third sample description in each image-text pair and the second sample image in that image-text pair, a fourth generation loss is determined (this fourth generation loss can be the first sum of the generation losses determined by the second sample images and the third generated images corresponding to the third sample descriptions in all image-text pairs, or the average of the first sum). The parameters of the image generation network to be trained are adjusted with the goal of minimizing the fourth generation loss. In this embodiment, the fourth generation loss is determined for a batch of samples, and then the parameters of the image generation network to be trained are adjusted. This reduces the number of parameter adjustments to the image generation network and makes the above process easier to implement.

[0065] In one implementation, the image generation network can be an image generation network based on the Diffusion Model.

[0066] The process of training a text generation network can include: constructing a sample description pair set D for pre-training the text generation network to be trained. The sample description pair set D can include several sample description pairs. The construction process of a sample description pair Di can be as follows: obtain the fourth sample description yi (for example, it can be represented as {(a)(sunny)(windless)(morning)} or {(a)(sunny)(windless)(morning)(upon)}, where the content in each parenthesis can represent a word in the fourth sample description yi. The first representation method is used as an example below for explanation), and assign it to the sample description pair Di; use a preset occlusion algorithm to occlude part of the content of the fourth sample description yi to obtain the corresponding occluded fourth sample description, which is used as the target occlusion description xi (for example, it can be represented as {(a)(?)(?)(morning)}, where the question mark "?" represents the occluded word), and assign it to the sample description pair Di. That is, a sample description pair Di includes a fourth sample description yi and its corresponding target occlusion description xi.

[0067] The preset occlusion algorithm can be any text occlusion algorithm in the relevant technology. In one implementation, the preset occlusion algorithm can be used to randomly occlude any content (or word) at any position in the fourth sample description yi, except for the first sentence (e.g., the sentence before the first punctuation mark); it can also randomly occlude words at any position in the fourth sample description yi, except for the first a words, where a is a positive integer.

[0068] Next, the text generation network to be trained is obtained. Any sample description pair Di is obtained from the sample description pair set D. Using the target occlusion description xi in the sample description pair Di, the output of the text generation network to be trained is obtained (hereinafter referred to as the first output for clarity). Then, using the first output and the fourth sample description yi in the sample description pair Di, the generation loss corresponding to the text generation network to be trained is determined (referred to as the fifth generation loss).

[0069] In one implementation, the first output may include a set of first word probabilities corresponding to each of several positions. Each set of first word probabilities includes the first probability that each word (token) in a preset word set determined by the text generation network to be trained for the target occlusion description xi appears at that position. The preset word set includes a set of words that the text generation network to be trained can predict. Each fourth sample description yi may include several words from the preset word set. The specific number of these positions can be determined based on specified parameters set for the text generation network to be trained, representing the maximum length of the output.

[0070] Then, based on the fourth sample description yi, the first probability of the word appearing at the corresponding position of the fourth sample description yi can be determined from the first output for each occluded position (or each position in the target occlusion description xi) of the text generation network to be trained (for example, the probability of the word "sunny" appearing at the first occlusion position, i.e., the position of the second word, in the aforementioned example, and the probability of the word "windless" appearing at the second occlusion position, i.e., the position of the third word, in the target occlusion description xi).

[0071] Then, the cross-entropy loss function is used to determine the fifth generation loss based on the first probability that the word at the corresponding position of the fourth sample description yi appears at each occluded position (or each position in the target occlusion description xi). The higher the first probability that the word at the corresponding position of the fourth sample description yi appears at each occluded position (or each position in the target occlusion description xi), the smaller the fifth generation loss, indicating a better image description generation ability of the text generation network to be trained. Accordingly, in one case, the fifth generation loss l can be expressed by the following formula:

[0072]

[0073] Where p(yi|xi) represents the probability that the text generation network to be trained outputs the fourth sample description yi for the target occlusion description xi. This p(yi|xi) can be determined based on the first probability of the word at the corresponding position of the fourth sample description yi appearing at each occluded position in the target occlusion description xi (or at each position in the target occlusion description xi) (e.g., the average of the sum of the first probabilities of the word at the corresponding position of the fourth sample description yi appearing at each occluded position in the target occlusion description xi, or the product of these probabilities); D represents the aforementioned set of sample description pairs D to which the fourth sample description yi and the target occlusion description xi belong. This indicates the calculation of the expected value.

[0074] In another implementation, the first output includes a set of first word probabilities corresponding to each of several positions. Then, the target program can determine the predicted description corresponding to the target occlusion description xi based on this first output, for example, by combining a specified first word determination rule and the first output. Then, based on the difference between the predicted description and the fourth sample description (e.g., using their Euclidean distance to represent the difference, or using the MSE loss function to determine the difference between the predicted description and the fourth sample description), a fifth generation loss is determined, where the smaller the difference between the predicted description and the fourth sample description, the smaller the fifth generation loss. The first word determination rule can be any rule in related technologies that determines the corresponding word based on the word probability set corresponding to each of the aforementioned positions (the words at some positions can be empty). In one implementation, the parameters in the first word determination rule can include pre-set parameters or parameters that are adjusted during the training process of the text generation network to be trained.

[0075] In another implementation, the first output can directly include the predicted description corresponding to the aforementioned target occlusion description xi. Accordingly, the fifth generation loss can be determined directly based on the difference between the predicted description and the fourth sample description (e.g., Euclidean distance).

[0076] Then, the parameters of the text generation network are adjusted to minimize the fifth generation loss, i.e., to maximize the probability that the text generation network to be trained outputs the fourth sample description yi for the target occlusion description xi, which aims to reduce the difference between the aforementioned predicted description and the fourth sample description, until the text generation network reaches the second convergence condition, thus obtaining the text generation network in the aforementioned multimodal generation model. The second convergence condition may include, but is not limited to, the number of parameter adjustments not being less than a second adjustment threshold, or the calculated fifth generation loss being less than a second loss threshold. This second adjustment threshold may be the same as or different from the aforementioned first adjustment threshold, and the second loss threshold may be the same as or different from the aforementioned first loss threshold.

[0077] To recap the aforementioned process of training the text generation network, the above embodiment used a single sample description pair (i.e., a pair of fourth sample descriptions and their corresponding target occlusion descriptions) as an example. In another embodiment, the above process can be performed on a batch of samples, i.e., multiple sample description pairs, to obtain the output of the text generation network to be trained, which processes the target occlusion description in each sample description pair. Then, based on the output corresponding to the target occlusion description in each sample description pair and the fourth sample description in each sample description pair, a fifth generation loss is determined. The parameters of the text generation network to be trained are adjusted with the goal of minimizing the fifth generation loss. In this embodiment, the fifth generation loss is determined for a batch of samples before adjusting the parameters of the text generation network. This reduces the number of parameter adjustments to the text generation network and makes the above process easier to implement.

[0078] In one implementation, the text generation network can be a text generation network based on a Transformer network.

[0079] After obtaining the image generation network and the text generation network, a multimodal generation model is constructed using the image generation network, the text generation network, and the multimodal generation network, and then the multimodal generation model is trained. Figure 2 A flowchart illustrating a training method for a multimodal generative model according to one embodiment of this disclosure is shown. This method can be implemented using a target program. This target program can be installed in an electronic device, which can be implemented using any device, equipment, platform, device cluster, etc., with computing and processing capabilities. Figure 2 As shown, the multimodal generation model includes an image generation network, a text generation network, and a multimodal generation network. The method includes the following steps S210-S270:

[0080] In step S210, an arbitrary first dialogue sample is obtained from the training sample set. The first dialogue sample includes a first sample description and its corresponding second sample description and sample instructions. The sample instructions are used to indicate the generation direction of the second sample description.

[0081] In one implementation, the target program can first construct a training sample set before training the multimodal generative model. This training sample set may include multiple dialogue samples. For any first dialogue sample in the training sample set, the acquisition process may specifically include the following steps 11-13: In step 11, obtain the description of the first sample and assign the description of the first sample to the first dialogue sample.

[0082] In one implementation, the first sample description can be a description that meets specified conditions. Specifically, the first sample description and its corresponding image (i.e., image-text pair) can be image-text pairs in a specified public dataset whose quality exceeds preset filtering conditions. The preset filtering conditions may include, but are not limited to, at least one of the following: the aesthetic score corresponding to the image in the image-text pair is not lower than a preset aesthetic score threshold (for example, if the aesthetic score range is [0, 1], the preset aesthetic score threshold is, for example, 0.75); the sharpness score corresponding to the image in the image-text pair is not lower than a preset sharpness threshold (for example, if the sharpness score range is [0, 1], the preset sharpness threshold is, for example, 0.80); and the matching degree value corresponding to the image-text pair is not lower than a preset matching degree threshold (for example, if the matching degree value range is [0, 1], the preset matching degree threshold is, for example, 0.80). Accordingly, the target program can determine the image descriptions in the image-text pairs in the specified public dataset whose quality exceeds the preset filtering conditions as the first sample description.

[0083] In one implementation, an aesthetic score can be determined for each image in an image-text pair using a trained aesthetic prediction model. This aesthetic prediction model can be any model from related technologies capable of evaluating the aesthetic appeal, i.e., the aesthetic score, of an image. A higher aesthetic score indicates a more aesthetically pleasing and higher-quality image. In another implementation, the aesthetic prediction model can be a model pre-trained based on images and their corresponding aesthetic score labels. This model can output the aesthetic score of an input image. The images and their corresponding aesthetic score labels used to train the aesthetic prediction model can be selected from any dataset containing a large number of images and their corresponding aesthetic score labels representing the aesthetic appeal of those images.

[0084] Based on the image in the image-text pair, the gradient value corresponding to the image in the image-text pair can be determined by a preset image gradient algorithm (such as the Tenengrad gradient algorithm). Then, based on the gradient value corresponding to the image, the sharpness score corresponding to the image can be determined (for example, based on the preset correspondence between gradient values ​​and sharpness scores, and the gradient value corresponding to the image, the sharpness score corresponding to the image can be determined). The larger the gradient value corresponding to the image, the larger the sharpness score corresponding to the image, and the better the image quality.

[0085] Based on the image-text pair, a trained text-image association model can be used to determine the matching degree value between the image and its corresponding image description in the image-text pair. The matching degree value can represent the matching degree between the image and its corresponding image description in the corresponding image-text pair. The larger the matching degree value, the stronger the matching degree between the image and its corresponding image description in the corresponding image-text pair.

[0086] The better the quality of the image in an image-text pair, the stronger the matching degree between the image and its corresponding image description. The better the quality of the image-text pair, the more beneficial it is for training multimodal image generation.

[0087] In one scenario, the text-image association model can be any model from the relevant technologies that can determine the matching degree between the input image and its corresponding image description. For example, this text-image association model could be the CLIP model (Contrastive Language-Image Pre-Training).

[0088] Next, in step 12, based on the first sample description, a second sample description is determined, which includes part of the content of the first sample description, and the second sample description is included in the first dialogue sample.

[0089] In one scenario, a second sample description can be obtained by partially occluding the content of the first sample description based on a preset occlusion algorithm. This preset occlusion algorithm can be any text occlusion algorithm available in related technologies. In one implementation, using this preset occlusion algorithm, any content at any position in the first sample description, excluding the first sentence (e.g., the sentence before the first punctuation mark), can be randomly occluded. Similarly, words at any position in the fourth sample description yi, excluding the first *a* words, can be randomly occluded, where *a* is a positive integer.

[0090] In another scenario, a second sample description can be obtained by semantically decomposing the first sample description. This decomposition can be performed using a specific program or manually. For example, if the first sample description is "a beautiful girl wearing a red dress and holding yellow sunflowers," the second sample description obtained after semantic decomposition could be, for example, "a girl," "a beautiful girl wearing a red dress," or "a girl holding sunflowers," etc. Understandably, a first sample description can be decomposed into multiple second sample descriptions, and these first sample descriptions can combine with the multiple second sample descriptions to form multiple dialogue samples.

[0091] Then, in step 13, for the second sample description, a sample instruction for indicating the generation direction of the second sample description is obtained. In one implementation, a preset storage space corresponding to the target program may store a dialogue instruction set in advance, wherein the dialogue instruction set includes a plurality of dialogue instructions, and the dialogue instruction may be, for example, "please add more details for this sentence", or "please add more details for this sentence", or "modify "XX (target, e.g., dog)" in the sentence to "YY (e.g., cat)", etc. The target program may select a dialogue instruction from the dialogue instruction set for the second sample description, and use it as the sample instruction for indicating the generation direction of the second sample description. In an exemplary scenario, when the selected sample instruction is "modify "XX (target, e.g., dog)" in the sentence to "YY (e.g., cat)", the second sample description may be further modified, for example, "YY" in the second sample description is modified to "XX", so as to obtain a new second sample description that can be classified into the first dialogue sample.

[0092] In yet another implementation, based on the second sample description and its corresponding first sample description, the corresponding sample instruction for the second sample description can be constructed through a specific program or manual operation. The generation direction of the second sample description indicated by the sample instruction may be a direction that instructs to generate a better (can generate an image with higher quality) image description that better matches the first sample description. For example, the sample instruction corresponding to the second sample description is constructed based on the content missing from the second sample description relative to its corresponding first sample description. For example, after the content describing image details is missing from the first sample description to obtain the second sample description, a constructor may construct a sample instruction instructing to generate more details for the second sample description. For another example, after the content describing weather is missing from the first sample description to obtain the second sample description, the constructor may construct a sample instruction instructing to generate weather content for the second sample description, which are all acceptable.

[0093] After the target program acquires any first dialogue sample in the training sample set, in step S220, a first generated image is obtained through an image generation network based on the first sample description. In this step, the target program may input the first sample description into the image generation network, and the image generation network processes the first sample description to obtain the first generated image corresponding to the first sample description, and accordingly the target program acquires the first generated image.

[0094] Next, in step S230, based on the second sample description and sample instructions, a first generated description is obtained through a text generation network. In this step, the target program inputs the second sample description and sample instructions into the text generation network. The text generation network processes the second sample description and sample instructions, processing the second sample description according to the generation direction indicated by the sample instructions to obtain the first generated description.

[0095] In one implementation, based on the second sample description and sample instructions, a text generation network is used to obtain its output (which can be referred to as the second output for clarity). In one case, this second output may include a set of second word probabilities corresponding to each of several positions. Each set of second word probabilities includes the second probability of each word (token) in the aforementioned preset word set appearing at that position, determined by the text generation network based on the second sample description and sample instructions. The preset word set is the set of words that the text generation network can predict. Based on the set of second word probabilities corresponding to each of the several positions included in the second output, the aforementioned first generation description can be determined. Specifically, by combining the set of second word probabilities corresponding to each of the several positions included in the second output with the aforementioned first word determination rule, the first generation description can be obtained. In another case, the second output may also include the first generation description.

[0096] Then in step S240, based on the first generated image and the first generated description, a second generated description is obtained through a multimodal generation network. The second generated description is related to the output of the multimodal generation network.

[0097] In this step, the target program can input the first generated image and the first generated description into the multimodal generation network, so that the multimodal generation network can process the first generated image and the first generated description to obtain the image representation corresponding to the first generated image and the text representation corresponding to the first generated description. Then, the image representation and the text representation are fused to obtain a fused representation that incorporates image domain data features and text domain data features. Based on this fused representation, the second generated description can then be obtained.

[0098] In one implementation, based on a first generated image and a first generated description, the output of a multimodal generation network (referred to as the third output for clarity) can be obtained through the multimodal generation network. In one case, the third output may include a set of third word probabilities corresponding to each of multiple positions predicted by the multimodal generation network for the first generated image and the first generated description. This set of third word probabilities may include the third probability of each word (token) in a specified word set appearing at each of the multiple positions. The specified word set includes the set of words that the multimodal generation network can predict, and this specified word set may be the same as the aforementioned preset word set. The specific number of these multiple positions can be determined based on specified parameters for the multimodal generation network, which can characterize the maximum length of the multimodal generation network's output.

[0099] The target program can determine the aforementioned second generated description based on the set of third-word probabilities corresponding to each of the multiple positions included in the third output. Specifically, by combining the set of third-word probabilities corresponding to each of the multiple positions included in the third output with the specified second-word determination rule, the second generated description is obtained. This second generated description integrates image domain data features (image representation corresponding to the first generated image) and text domain data features (text representation corresponding to the first generated description), and contains richer information for generating images. The second-word determination rule can be any rule in related technologies that determines the corresponding word based on the set of third-word probabilities corresponding to each of the aforementioned multiple positions (the word at one or more positions can be empty). In one implementation, the parameters in the second-word determination rule can include pre-set parameters or parameters that are adjusted during the training process of the multimodal generation network; both are acceptable. The specified second-word determination rule can be the same as or different from the aforementioned specified first-word determination rule.

[0100] The multimodal generation network can be any fusion network capable of fusing images and text in related technologies. In one embodiment, the multimodal generation network may include a text encoder, an image encoder, a fusion unit, and a decoder; correspondingly, step S240 may include the following steps 21-24: In step 21, using the first generated image, an image representation is obtained through the image encoder. In this step, the target program can input the first generated image into the image encoder, and the image encoder performs image encoding on the first generated image to obtain the image representation corresponding to the first generated image. The image encoder can be any encoder in related technologies capable of encoding images and extracting image representations from images.

[0101] Next, in step 22, the text representation is obtained using the first generated description and a text encoder. In this step, the target program can input the first generated description into the text encoder, which encodes the first generated description to obtain the text representation corresponding to the first generated description. This text encoder can be any encoder in the related art capable of encoding text and extracting the text representation from it.

[0102] Then, in step 23, the image representation and text representation are used to obtain a fused representation through a fusion processor. In this step, the image representation and text representation can be input into the fusion processor, which performs a fusion operation on the image representation and text representation to obtain a fused representation that incorporates both the image representation and text representation. In one implementation, this fusion operation may include, but is not limited to, a cross-attention operation, i.e., a fusion operation based on a cross-attention mechanism.

[0103] Next, in step 24, the fused representation is used to obtain a second generated description through a decoder. In this step, the fused representation is input into the decoder, which decodes the fused representation to obtain the decoding result corresponding to the fused representation. Based on the decoding result corresponding to the fused representation, a second generated description that fuses image and text representations is obtained.

[0104] After obtaining the second generated description, the target program generates a second generated image based on the second generated description in step S250. In one implementation, the target program can input the second generated description into the aforementioned image generation network to obtain the second generated image corresponding to the second generated description. In another implementation, the target program can input the second generated description into another pre-trained image generation network to obtain the second generated image corresponding to the second generated description.

[0105] Then, in step S260, the target generation loss is determined based on the first sample description, the output of the multimodal generation network, and the second generated image. In this step, the target program can determine the target generation loss based on the output of the multimodal generation network and the difference between the generation result of the multimodal generation network (i.e., the second generated description) represented by the first sample description and the first sample description, as well as the matching degree, i.e., the correlation, between the second generated image and the first sample description.

[0106] Specifically, in one implementation, the target program can use a specified loss function (e.g., cross-entropy loss function, mean squared error loss function) to determine the corresponding generation loss (i.e., the first generation loss) based on the output of the multimodal generation network and the first sample description. Based on the second generated image and the first sample description, a matching loss (called the first matching loss) is determined between the two through a text-image association model. Then, the target generation loss is determined by combining the first generation loss and the first matching loss. The target generation loss is positively correlated with both the first generation loss and the first matching loss. Minimizing the first generation loss allows the text generation network to generate higher-quality image descriptions, while minimizing the first matching loss allows the multimodal generation model to generate more accurate image descriptions.

[0107] In one implementation, the output of the multimodal generation network, i.e. the aforementioned third output, may include the set of third word probabilities corresponding to each position in multiple positions predicted by the multimodal generation network for the first generated image and the first generated description. The aforementioned process of determining the corresponding generation loss (i.e. the first generation loss) may be: using the cross-entropy loss function, i.e. according to the principle of maximizing log-likelihood, constructing the cross-entropy loss between the three positions based on the set of third word probabilities corresponding to each position in the multiple positions of the third output and the first sample description, and determining the cross-entropy loss as the first generation loss.

[0108] Specifically, this can be achieved by: determining the third probability of the word appearing at each specified first position in the third output, corresponding to the position in the first sample description, from the set of third word probabilities corresponding to each position in the multiple positions of the third output; and then determining the cross-entropy loss, i.e., the first generation loss, based on the third probability of the word appearing at each specified first position in the first sample description. The first generation loss is negatively correlated with the third probability of the word appearing at each specified first position in the third output, corresponding to the position in the first sample description. That is, the higher the third probability of the word appearing at each specified first position in the third output, the smaller the difference between the aforementioned second generated description and the first sample description, and the smaller the corresponding first generation loss.

[0109] In the above implementation, the second generated description can also be determined based on the probability set of the third word corresponding to each position in the multiple positions included in the third output. Specifically, the second generated description is determined by combining the probability set of the third word corresponding to each position in the multiple positions included in the third output and the aforementioned second word determination rule. Then, the aforementioned first generation loss is determined based on the difference between the second generated description and the first sample description. The first generation loss is positively correlated with the difference between the second generated description and the first sample description. The smaller the difference between the second generated description and the first sample description, the smaller the first generation loss.

[0110] In another implementation, the output of the multimodal generator network includes the aforementioned second generator description, and the target program can directly determine the aforementioned first generator loss based on the difference between the second generator description and the first sample description.

[0111] The aforementioned process of determining the first matching loss can be as follows: input the second generated image and the first sample description into the text-image association model to obtain the matching degree value between the second generated image and the first sample description; then, based on the matching degree value between the second generated image and the first sample description, determine the first matching loss between them. It can be understood that the matching degree value between the second generated image and the first sample description can range from [0, 1]. The larger the correlation value between them, the stronger the matching degree, i.e., the stronger the correlation, between the second generated image and the first sample description. The first matching loss can be equal to the difference between 1 and the matching degree value between the second generated image and the first sample description; the smaller this difference, the better the matching degree between the second generated image and the first sample description.

[0112] In one implementation, the text-image association model can be any model in the relevant techniques that can determine the correlation between images and text, such as the CLIP model.

[0113] In one embodiment, to ensure that the second generated description is more accurate, the target generation loss can be determined by combining the relationship between the first generated image and the second generated description. Specifically, step S260 may include: determining the target generation loss based on the first generated image, the first sample description, the output of the multimodal generation network, and the second generated image.

[0114] Specifically, step S260 may include the following steps 31-33: In step 31, a first generation loss is determined based on the output of the multimodal generation network and the first sample description. For details, please refer to the aforementioned process for determining the first generation loss, which will not be repeated here.

[0115] Next, in step 32, based on the second generated image and the first sample description, a first matching loss is determined using a text-image association model; and based on the first generated image and the second generated description, a second matching loss is determined using the text-image association model. Specifically, the process for determining the second matching loss can be referred to the aforementioned process for determining the first matching loss, and is not limited here.

[0116] Next, in step 33, the target generation loss is determined using the first generation loss, the first matching loss, and the second matching loss. In this step, the target generation loss is determined by combining the first generation loss, the first matching loss, and the second matching loss, where the target generation loss is positively correlated with the first generation loss, the first matching loss, and the second matching loss, respectively. Determining the target generation loss by combining the first generation loss can enable the text generation network to generate higher-quality image descriptions, allowing the multimodal generation network to learn the image generation content of the image generation network, thereby achieving the fusion of the image generated by the text generation network with the text (language) domain data (i.e., the first generation description). Furthermore, determining the target generation loss by combining the first matching loss and the second matching loss can better integrate the image generation capabilities of the image generation network and the text generation capabilities of the text generation network in the multimodal generation model, thereby better integrating image domain data and text (language) domain data, ensuring the accuracy of the generated image description (i.e., the output of the multimodal generation network), and thus improving the quality of the generated image (i.e., the image generated based on the second generation description) while ensuring the accuracy of the generated image.

[0117] In one embodiment, in order to ensure the stability of the image generation capability of the image generation network and to better improve the ability of the multimodal generation model to generate higher quality image descriptions, the first dialogue sample also includes a first sample image corresponding to the first sample description.

[0118] Step S260 further includes the following steps: determining a second generation loss corresponding to the image generation network based on the difference between the first sample image and the first generated image; and determining a target generation loss using the second generation loss, the first generation loss, the first matching loss, and the second matching loss. The target program can use a preset loss function to determine the second generation loss corresponding to the image generation network based on the difference between the first sample image and the first generated image, wherein the second generation loss is positively correlated with the difference between the first sample image and the first generated image. In this embodiment, by combining the first generation loss, the first matching loss, and the second matching loss with the difference between the first sample image and the first generated image (i.e., the second generation loss), the target generation loss can be determined. This allows for fine-tuning of the parameters of the image generation network using the training sample set, i.e., fine-tuning the image generation capability of the image generation network, enabling the image generation network to generate better images based on image descriptions, thereby further improving the image description generation capability of the entire multimodal generation model.

[0119] Specifically, step S260 may include the following steps 41-44: In step 41, a first generation loss is determined based on the output of the multimodal generation network and the first sample description. Step 41 is the same as step 31 described above and will not be repeated here.

[0120] In step 42, based on the difference between the first sample image and the first generated image, a second generation loss corresponding to the image generation network is determined. In this step, a preset loss function is used to determine the second generation loss corresponding to the image generation network based on the difference between the first sample image and the first generated image. This preset loss function may include, but is not limited to, the cross-entropy loss function, the MSE loss function, and the MAE loss function.

[0121] In step 43, based on the second generated image and the first sample description, a first matching loss is determined using a text-image association model; and based on the first generated image and the second generated description, a second matching loss is determined using the text-image association model. Step 43 is the same as step 32 described above, and will not be repeated here.

[0122] In step 44, the target generation loss is determined using the first generation loss, the second generation loss, the first matching loss, and the second matching loss. In this step, the target generation loss is positively correlated with the first generation loss, the second generation loss, the first matching loss, and the second matching loss, respectively.

[0123] In another embodiment, step 44 may further include: determining a third generation loss corresponding to the text generation network based on the first sample description and the output of the text generation network; and determining a target generation loss using the third generation loss, the first generation loss, the second generation loss, the first matching loss, and the second matching loss. The target generation loss is also positively correlated with the third generation loss. In this embodiment, the aforementioned specified loss function (e.g., cross-entropy loss function, mean squared error loss function) can be used to determine the third generation loss corresponding to the text generation network based on the first sample description and the output of the text generation network. In this embodiment, based on the first generation loss, the second generation loss, the first matching loss, and the second matching loss, the target generation loss is determined by combining the third generation loss determined based on the first sample description and the output of the text generation network, thereby better improving the ability of the text generation network to generate image descriptions, and further improving the ability of the multimodal generation model to generate higher-quality image descriptions.

[0124] After obtaining the target generation loss through the aforementioned method, in step S270, the parameters of the multimodal generation model are adjusted with the goal of minimizing the target generation loss. In this step, minimizing the target generation loss means reducing the difference between the output of the multimodal generation network and the generated result of the multimodal generation network (i.e., the second generated description) represented by the first sample description and the first sample description; reducing the matching loss between the second generated description and the first generated image; reducing the matching loss between the second generated image and the first sample description (and reducing the difference between the first sample image and the first generated image, and reducing the difference between the first sample description and the generated result of the text generation network (i.e., the first generated description) represented by the output of the text generation network and the first sample description). Adjusting the parameters of the multimodal generation model improves its ability to generate higher-quality image descriptions, resulting in better-quality image descriptions, and thus ensuring that better-quality images are generated using these better-quality image descriptions.

[0125] To recap the execution process of steps S210-S270, the above embodiment uses a single first dialogue sample as an example. In another embodiment, steps S210-S260 can be performed on a batch of samples, i.e., multiple first dialogue samples, to obtain a first generated image, a first sample description, the output of the multimodal generation network, and a second generated image (as well as a second generated description determined based on the output of the multimodal generation network, a first generated description determined based on the output of the text generation network, etc.) for each first dialogue sample. Then, in step S260, the target generation loss is determined based on the above data corresponding to each first dialogue sample, and in step S270, the parameters of the multimodal generation model are adjusted with the goal of minimizing the target generation loss. In this embodiment, the target generation loss is determined for a batch of samples, and then the parameters of the multimodal generation model are adjusted. This reduces the number of times the parameters of the multimodal generation model are adjusted, making the above process easier to implement.

[0126] Adjust the parameters of the multimodal generation model until it reaches a preset convergence condition, thus obtaining a trained multimodal generation model. This preset convergence condition may include, but is not limited to, the number of parameter adjustments exceeding a preset threshold, or the target generation loss falling below a preset generation loss threshold.

[0127] In this embodiment, the idea of ​​multimodal learning is used to construct various fine-grained losses and adjust the multimodal generation model to improve the ability of the multimodal generation model to generate higher quality image descriptions. In this way, higher quality image descriptions can be obtained by using the multimodal generation model, thereby improving the quality of the images generated based on the image descriptions.

[0128] The multimodal generation model can be a conversational multimodal generation model. For a trained multimodal generation model, it can be used to generate a corresponding high-quality image description based on the user-input data to be processed (the image to be processed and / or the image description to be processed) and the user-input instructions (dialogue content) indicating the generation direction of the data to be processed. Then, a high-quality image is generated based on the high-quality image description.

[0129] Corresponding to the above method embodiments, this disclosure also provides an image generation method, which can be implemented by a first program. This first program can be the same physical program as the aforementioned target program, or it can be a different physical program from the aforementioned target program, such as... Figure 3 As shown, the image generation method may include the following steps S310-S340:

[0130] In step S310, the data to be processed and a first instruction for indicating the generation direction of the data to be processed are obtained.

[0131] It is understood that the image generation method provided in this disclosure embodiment can be an image generation method based on a multimodal generation model. The target multimodal generation model can be a multimodal generation model trained using the aforementioned multimodal generation model training method.

[0132] In one implementation, the data to be processed and the first instruction are input by a user who needs to generate an image. The data to be processed can be data in any form. In one embodiment, the data to be processed can be text-based data, and correspondingly, the data to be processed includes a description of the image to be processed, i.e., a text prompt for generating the image, and the user has a need to generate an image based on the description of the image to be processed and the first instruction. In another embodiment, the data to be processed can be image-based data, and correspondingly, the data to be processed includes an image to be processed, and the user has a need to generate a higher-quality image based on the image to be processed that better matches the intent (generation direction) expressed by the first instruction.

[0133] In another implementation, the data to be processed can be text-based data (i.e., image description) output by the target multimodal generation model, or it can be image-based data (i.e., generated image) generated by the target image generation network of the target multimodal generation model based on the text-based data output by the target multimodal generation model. The first instruction is data (in text form) input by the user indicating the generation direction of the data to be processed.

[0134] In one embodiment, the data to be processed includes an image to be processed (i.e., data in image form). The image to be processed can be any image, and the first instruction can be a user input instruction indicating the generation direction of the image to be processed, such as instructing the generation of more detailed content or the generation of specified object content for the image to be processed, etc.

[0135] In yet another embodiment, the data to be processed includes an image description to be processed (i.e., data in text form). The first instruction can be a user input indicating the direction of generation of the image description to be processed, such as instructing the generation of an image description containing more details or containing a specified object, etc.

[0136] After the first program obtains the data to be processed and the first instruction, in step S320, based on the first instruction, the third generated description is obtained through the target text generation network of the target multimodal generation model.

[0137] In one implementation, when the data to be processed includes an image to be processed, the first program can input a first instruction into the target text generation network of the target multimodal generation model, and the target text generation network processes the first instruction to obtain a third generated description.

[0138] In another implementation, where the aforementioned data to be processed includes an image description to be processed, step S320 may specifically include: using the image description to be processed and the first instruction, a target text generation network is used to obtain a third generated description. In this step, the first program can input the image description to be processed and the first instruction into the target text generation network. The target text generation network processes the image description to be processed and the first instruction, that is, it processes the image description to be processed based on the generation direction indicated by the first instruction to obtain the third generated description.

[0139] Next, in step S330, based on the data to be processed and the third generation description, a fourth generation description is obtained through the target multimodal generation network of the target multimodal generation model.

[0140] In one implementation, when the data to be processed includes an image to be processed, the first program can input the image to be processed and the third generated description into the target multimodal generation network of the target multimodal generation model. The target multimodal generation network processes the image to be processed and the third generated description to obtain a fourth generated description, which integrates the features of the image to be processed and the features of the third generated description.

[0141] In another implementation, when the data to be processed includes a description of the image to be processed, step S330 may specifically include: using the description of the image to be processed, passing through a target image generation network to obtain a fourth generated image; and using the third generated description and the fourth generated image, passing through a target multimodal generation network to obtain a fourth generated description. In this implementation, the first program inputs the description of the image to be processed into the target image generation network of the target multimodal generation model to obtain the generated image corresponding to the description of the image to be processed, i.e., the fourth generated image; then, the third generated description and the fourth generated image are input into the target multimodal generation network to obtain a fourth generated description of better quality that integrates the features of the fourth generated image and the features of the third generated description.

[0142] Subsequently, in step S340, based on the fourth generation description, a generated image corresponding to the data to be processed and the first instruction is obtained through the target image generation network of the target multimodal generation model. In this step, the first program inputs the fourth generation description into the target image generation network of the target multimodal generation model to obtain a generated image corresponding to the data to be processed and the first instruction.

[0143] In the above process, at least based on the first instruction representing the generation direction of the data to be processed indicated by the user, a third generation description is obtained, and then a fourth generation description is obtained by fusing the information of the data to be processed and the third generation description. Based on the fourth generation description, a higher quality generated image can be obtained.

[0144] In one implementation, the target multimodal generation model can, for text-based data such as novels and articles, generate corresponding image descriptions for generating higher-quality illustrations based on instructions indicating the generation direction, and then generate corresponding illustrations based on these image descriptions; alternatively, it can, for the original image, generate image descriptions that can yield higher-quality images based on instructions indicating the generation direction, and then generate higher-quality images based on these images. Accordingly, the multimodal generation model trained using the training method provided in this embodiment can optimize the input image description or the input image through dialogue (i.e., user-input instructions indicating the generation direction) to obtain higher-quality image descriptions that conform to the user's intent, thereby obtaining higher-quality images.

[0145] The foregoing description describes specific embodiments of this disclosure; other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than those shown in the embodiments, and the desired result may still be achieved. Furthermore, the processes depicted in the drawings do not necessarily need to follow the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0146] Corresponding to the above method embodiments, this disclosure provides a training device 400 for a multimodal generative model. The multimodal generative model includes an image generation network, a text generation network, and a multimodal generative network, as illustrated in the schematic block diagram below. Figure 4 As shown, used to execute Figure 2 The method shown includes:

[0147] The first acquisition module 410 is configured to acquire any first dialogue sample in the training sample set. The first dialogue sample includes a first sample description and its corresponding second sample description and sample instruction. The sample instruction is used to indicate the generation direction of the second sample description.

[0148] The first obtaining module 420 is configured to obtain a first generated image based on the first sample description through the image generation network;

[0149] The second obtaining module 430 is configured to obtain a first generated description based on the second sample description and the sample instructions through the text generation network;

[0150] The third module 440 is configured to obtain a second generated description based on the first generated image and the first generated description through the multimodal generation network, wherein the second generated description is related to the output of the multimodal generation network.

[0151] The generation module 450 is configured to generate a second generated image based on the second generation description;

[0152] The first determining module 460 is configured to determine the target generation loss based on the first sample description, the output of the multimodal generation network, and the second generated image;

[0153] The first adjustment module 470 is configured to adjust the parameters of the multimodal generation model with the goal of minimizing the target generation loss.

[0154] In one alternative implementation, it further includes:

[0155] The third acquisition module (not shown in the figure) is configured to acquire the description of the first sample;

[0156] The first occlusion module (not shown in the figure) is configured to determine the second sample description based on the first sample description, wherein the second sample description includes part of the content of the first sample description;

[0157] The fourth acquisition module (not shown in the figure) is configured to acquire sample instructions for indicating the generation direction of the second sample description for the second sample description.

[0158] In one optional implementation, the first determining module 460 is specifically configured to: determine the target generation loss based on the first generated image, the first sample description, the output of the multimodal generation network, and the second generated image.

[0159] In one alternative implementation, the first determining module 460 includes:

[0160] The first determining unit (not shown in the figure) is configured to determine the first generation loss based on the output of the multimodal generation network and the first sample description;

[0161] The second determining unit (not shown in the figure) is configured to determine a first matching loss based on the second generated image and the first sample description through a text image association model; and to determine a second matching loss based on the first generated image and the second generated description through the text image association model.

[0162] The third determining unit (not shown in the figure) is configured to determine the target generation loss using the first generation loss, the first matching loss, and the second matching loss.

[0163] In one optional implementation, the first dialogue sample further includes a first sample image corresponding to the first sample description;

[0164] The third determining unit (not shown in the figure) is specifically configured to determine the second generation loss corresponding to the image generation network based on the difference between the first sample image and the first generated image.

[0165] The target generation loss is determined using the second generation loss, the first generation loss, the first matching loss, and the second matching loss.

[0166] In one optional implementation, the third determining unit (not shown in the figure) is specifically configured to determine a third generation loss corresponding to the text generation network based on the first sample description and the output of the text generation network, wherein the output of the text generation network is related to the first generation description;

[0167] The target generation loss is determined using the third generation loss, the first generation loss, the second generation loss, the first matching loss, and the second matching loss.

[0168] In one alternative implementation, the multimodal generation network includes a text encoder, an image encoder, a fusion unit, and a decoder;

[0169] The third obtaining module 440 is specifically configured to use the first generated image to obtain an image representation through the image encoder;

[0170] Using the first generated description, a text representation is obtained through the text encoder;

[0171] Using the image representation and the text representation, a fused representation is obtained through the fusion processor;

[0172] Using the fused representation, the second generated description is obtained through the decoder.

[0173] In one alternative embodiment, the device further includes:

[0174] The fifth acquisition module (not shown in the figure) is configured to acquire the third sample description and its corresponding second sample image before obtaining the first generated image;

[0175] The seventh module (not shown in the figure) is configured to obtain a third generated image based on the third sample description and through the image generation network to be trained;

[0176] The second determining module (not shown in the figure) is configured to determine the fourth generation loss corresponding to the image generation network to be trained by utilizing the difference between the second sample image and the third generated image.

[0177] The second adjustment module (not shown in the figure) is configured to adjust the parameters of the image generation network to be trained with the goal of minimizing the fourth generation loss, so as to obtain the image generation network that meets the first convergence condition.

[0178] In one alternative embodiment, the device further includes:

[0179] The sixth acquisition module (not shown in the figure) is configured to acquire the fourth sample description before the first generated description is obtained;

[0180] The second occlusion module (not shown in the figure) is configured to use a preset occlusion algorithm to occlude part of the content described by the fourth sample to obtain the target occlusion description.

[0181] The eighth module (not shown in the figure) is configured to use the target occlusion description to obtain the output of the text generation network to be trained through the text generation network to be trained;

[0182] The third determining module (not shown in the figure) is configured to determine the fifth generation loss corresponding to the text generation network to be trained by using the output of the text generation network to be trained and the fourth sample description.

[0183] The third adjustment module (not shown in the figure) is configured to adjust the parameters of the text generation network to be trained with the goal of minimizing the fifth generation loss, so as to obtain the text generation network that meets the second convergence condition.

[0184] Corresponding to the above method embodiments, this disclosure provides an image generation apparatus 500, the schematic block diagram of which is shown below. Figure 5 As shown, used to execute Figure 3 The method shown, wherein the apparatus includes:

[0185] The second acquisition module 510 is configured to acquire data to be processed and a first instruction for indicating the generation direction of the data to be processed.

[0186] The fourth module 520 is configured to obtain a third generated description based on the first instruction through the target text generation network of the target multimodal generation model;

[0187] The fifth module 530 is configured to obtain a fourth generation description based on the data to be processed and the third generation description, through the target multimodal generation network of the target multimodal generation model;

[0188] The sixth module 540 is configured to obtain a generated image corresponding to the data to be processed and the first instruction by using the image generation network of the target multimodal generation model based on the fourth generation description.

[0189] In one alternative implementation, the data to be processed includes an image to be processed.

[0190] In one alternative implementation, the data to be processed includes a description of the image to be processed;

[0191] The fourth obtaining module 520 is specifically configured to use the image description to be processed and the first instruction to obtain the third generated description through the target text generation network;

[0192] The fifth obtaining module 530 is specifically configured to use the image description to obtain a fourth generated image through the target image generation network; and use the third generated description and the fourth generated image to obtain the fourth generated description through the target multimodal generation network.

[0193] In one alternative implementation, the target multimodal generative model is a model trained using the aforementioned training device for multimodal generative models.

[0194] The above-described apparatus embodiments correspond to the method embodiments, and detailed descriptions can be found in the description of the method embodiments section, which will not be repeated here. The apparatus embodiments are derived based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments; detailed descriptions can be found in the corresponding method embodiments.

[0195] This disclosure provides an electronic device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements... Figure 2 or Figure 3 The method shown below. See below for reference. Figure 6 It shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0196] like Figure 6As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 403. The RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0197] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic devices to exchange data via wireless or wired communication with other devices. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively. Figure 6 Each box shown can represent a device or multiple devices as needed.

[0198] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by a processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0199] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the functions provided in this disclosure, such as those described in this disclosure. Figure 2 The training method for the multimodal generative model shown, or Figure 3 The image generation method shown.

[0200] It should be noted that the computer-readable medium described in the embodiments of this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the embodiments of this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiments of this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.

[0201] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire any first dialogue sample from a training sample set, the first dialogue sample including a first sample description and its corresponding second sample description and sample instructions, the sample instructions being used to indicate the generation direction of the second sample description; based on the first sample description, obtain a first generated image through an image generation network of a multimodal generation model; based on the second sample description and the sample instructions, obtain a first generated description through a text generation network of a multimodal generation model; based on the first generated image and the first generated description, obtain a second generated description through a multimodal generation network of a multimodal generation model, the second generated description being related to the output of the multimodal generation network; generate a second generated image based on the second generated description; determine a target generation loss based on the first sample description, the output of the multimodal generation network, and the second generated image; and adjust the parameters of the multimodal generation model with the goal of minimizing the target generation loss. Alternatively, the electronic device may: acquire data to be processed and a first instruction indicating the generation direction of the data to be processed; based on the first instruction, obtain a third generation description through a target text generation network of a target multimodal generation model; based on the data to be processed and the third generation description, obtain a fourth generation description through a target multimodal generation network of the target multimodal generation model; and based on the fourth generation description, obtain a generated image corresponding to the data to be processed and the first instruction through a target image generation network of the target multimodal generation model.

[0202] Computer program code for performing the operations of embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0203] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for storage media and computing devices are basically similar to the method embodiments, so they are described more simply; relevant parts can be referred to the descriptions of the method embodiments.

[0204] Those skilled in the art will recognize that the functions described in the embodiments of the present invention in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.

[0205] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, or improvements made based on the technical solutions of the present invention should be included within the scope of protection of the present invention.

Claims

1. A training method for a multimodal generative model, wherein the multimodal generative model includes an image generation network, a text generation network, and a multimodal generative network, the method comprising: Obtain any first dialogue sample from the training sample set. The first dialogue sample includes a first sample description and its corresponding second sample description and sample instructions. The sample instructions are used to indicate the generation direction of the second sample description. Based on the first sample description, a first generated image is obtained through the image generation network; Based on the second sample description and the sample instruction, a first generated description is obtained through the text generation network; Based on the first generated image and the first generated description, a second generated description is obtained through the multimodal generation network, and the second generated description is related to the output of the multimodal generation network; Based on the second generation description, a second generated image is generated; Based on the first sample description, the output of the multimodal generation network, and the second generated image, the target generation loss is determined; The parameters of the multimodal generation model are adjusted with the goal of minimizing the target generation loss.

2. The method of claim 1, further comprising: Obtain the description of the first sample; Based on the first sample description, a second sample description is determined, wherein the second sample description includes a portion of the content of the first sample description; For the second sample description, obtain sample instructions that indicate the generation direction of the second sample description.

3. The method of claim 1, wherein, The determination of the target generation loss includes: The target generation loss is determined based on the first generated image, the first sample description, the output of the multimodal generation network, and the second generated image.

4. The method of claim 3, wherein, The determination of the target generation loss includes: Based on the output of the multimodal generation network and the first sample description, a first generation loss is determined; Based on the second generated image and the first sample description, a first matching loss is determined through a text image association model; and based on the first generated image and the second generated description, a second matching loss is determined through the text image association model. The target generation loss is determined using the first generation loss, the first matching loss, and the second matching loss.

5. The method of claim 4, wherein, The first dialogue sample also includes a first sample image corresponding to the first sample description; The step of determining the target generation loss using the first generation loss, the first matching loss, and the second matching loss includes: Based on the difference between the first sample image and the first generated image, the second generation loss corresponding to the image generation network is determined; The target generation loss is determined using the second generation loss, the first generation loss, the first matching loss, and the second matching loss.

6. The method of claim 5, wherein, The step of determining the target generation loss using the second generation loss, the first generation loss, the first matching loss, and the second matching loss includes: Based on the first sample description and the output of the text generation network, a third generation loss corresponding to the text generation network is determined, wherein the output of the text generation network is related to the first generation description; The target generation loss is determined using the third generation loss, the first generation loss, the second generation loss, the first matching loss, and the second matching loss.

7. The method of claim 1, wherein, The multimodal generation network includes a text encoder, an image encoder, a fusion unit, and a decoder; The obtained second generated description includes Using the first generated image, an image representation is obtained through the image encoder; Using the first generated description, a text representation is obtained through the text encoder; Using the image representation and the text representation, a fused representation is obtained through the fusion processor; Using the fused representation, the second generated description is obtained through the decoder.

8. The method of claim 1, further comprising, before obtaining the first generated image: Obtain the description of the third sample and its corresponding image of the second sample; Based on the third sample description, a third generated image is obtained through the image generation network to be trained; The fourth generation loss corresponding to the image generation network to be trained is determined by using the difference between the second sample image and the third generated image; With the goal of minimizing the fourth generation loss, the parameters of the image generation network to be trained are adjusted to obtain the image generation network that meets the first convergence condition.

9. The method of claim 1, further comprising, before obtaining the first generated description: Obtain the description of the fourth sample; Using a preset occlusion algorithm, a portion of the content described by the fourth sample is occluded to obtain the target occlusion description; Using the target occlusion description, the output of the text generation network to be trained is obtained through the text generation network to be trained; Using the output of the text generation network to be trained and the fourth sample description, the fifth generation loss corresponding to the text generation network to be trained is determined; With the goal of minimizing the fifth generation loss, the parameters of the text generation network to be trained are adjusted to obtain the text generation network that meets the second convergence condition.

10. An image generation method, comprising: Acquire the data to be processed and a first instruction for indicating the generation direction of the data to be processed; Based on the first instruction, a third generated description is obtained through the target text generation network of the target multimodal generation model, wherein the target multimodal generation model is a model trained using the training method of the multimodal generation model according to any one of claims 1-9; Based on the data to be processed and the third generated description, a fourth generated description is obtained through the target multimodal generation network of the target multimodal generation model; Based on the fourth generation description, a generated image corresponding to the data to be processed and the first instruction is obtained through the target image generation network of the target multimodal generation model.

11. The method of claim 10, wherein, The data to be processed includes images to be processed.

12. The method of claim 10, wherein, The data to be processed includes a description of the image to be processed; The obtained third generation description includes: Using the image description to be processed and the first instruction, the third generated description is obtained through the target text generation network; The fourth generated description includes: Using the description of the image to be processed, a fourth generated image is obtained through the target image generation network; The fourth generated description is obtained by using the third generated description and the fourth generated image through the target multimodal generation network.

13. A training apparatus for a multimodal generative model, the multimodal generative model comprising an image generation network, a text generation network, and a multimodal generative network, the apparatus comprising: The first acquisition module is configured to acquire any first dialogue sample in the training sample set. The first dialogue sample includes a first sample description and its corresponding second sample description and sample instruction. The sample instruction is used to indicate the generation direction of the second sample description. The first obtaining module is configured to obtain a first generated image based on the first sample description and through the image generation network; The second obtaining module is configured to obtain a first generated description based on the second sample description and the sample instructions through the text generation network; The third obtaining module is configured to obtain a second generating description based on the first generated image and the first generated description through the multimodal generation network, wherein the second generated description is related to the output of the multimodal generation network; The generation module is configured to generate a second generated image based on the second generation description; The first determining module is configured to determine the target generation loss based on the first sample description, the output of the multimodal generation network, and the second generated image; The first adjustment module is configured to adjust the parameters of the multimodal generation model with the goal of minimizing the target generation loss.

14. An image generation apparatus, comprising: The second acquisition module is configured to acquire data to be processed and a first instruction for indicating the generation direction of the data to be processed. The fourth module is configured to obtain a third generated description based on the first instruction through the target text generation network of the target multimodal generation model, wherein the target multimodal generation model is a model trained using the training method of the multimodal generation model according to any one of claims 1-9. The fifth module, based on the data to be processed and the third generation description, obtains the fourth generation description through the target multimodal generation network of the target multimodal generation model; The sixth module, based on the fourth generation description, obtains a generated image corresponding to the data to be processed and the first instruction through the target image generation network of the target multimodal generation model.

15. An electronic device comprising a memory and a processor, wherein, The memory stores executable code, and when the processor executes the executable code, it implements the method of any one of claims 1-9, or the method of any one of claims 10-12.

16. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-9, or the method of any one of claims 10-12.

Citation Information

Patent Citations

  • Power transmission inspection image generation method and device based on multi-modal data

    CN114937181A

  • Image generation method and device, storage medium and electronic equipment

    CN116188632A

  • Image generation method and device of multi-modal model based on double-contrast learning

    CN116612205A