Visual Question Answering Model Training, Application Method, Device and Equipment
By using visual cues instructions to train the diffusion model, the problem of insufficient understanding of text cues is solved, and the model training efficiency and accuracy of generating images are improved.
Patent Information
- Application Number
- CN202510332093.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In the prior art, the model lacks the ability to understand text prompts, resulting in low model training efficiency.
By using visual prompt instructions to train the diffusion model, the features of problem text and input images are directly extracted from the matrix parameters of the visual prompt instructions, reducing the model's requirements for complex text understanding.
The training efficiency of the diffusion model is improved, allowing the model to obtain problem text features and input image features more intuitively, and generate output images that are more in line with the problem description.
Smart Images

Figure CN119848221B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and in particular, to a method, device, and equipment for training and applying a visual question answering model. Background Art
[0002] With the rapid development of information technology, the question-to-image generation task, as a typical cross-modal generation task, has been gradually applied to multiple fields. The question-to-image generation task model can generate a result image for output display by inputting an image and a proposed question to answer the proposed question.
[0003] In related technologies, questions are usually proposed in the form of text prompts. Then, for the input image and the proposed question, a multimodal large model analyzes the text prompt of the question and generates a result image based on the input image through the multimodal large model. However, since the text prompt is not intuitive and vivid enough, and some scenarios cannot be described in language, a large amount of training is required during model training to enable the model to understand the text prompt, resulting in low training efficiency. Summary of the Invention
[0004] The present disclosure provides a method, device, and equipment for training and applying a visual question answering model to at least solve the problem that the model in related technologies is not easy to understand text prompts, resulting in low model training efficiency.
[0005] The first aspect embodiment of the present disclosure provides a method for training a visual question answering model, including: determining a first output image by using a diffusion model based on a first input image and a first visual prompt instruction of the first input image; determining an instruction learning target of the diffusion model based on the first input image and the first output image; and training the diffusion model based on the first input image, the first output image, the first visual prompt instruction, and the instruction learning target to obtain a pre-trained diffusion model.
[0006] The second aspect embodiment of the present disclosure provides a method for applying a visual question answering model, including: obtaining an image to be processed and a second question text corresponding to the image to be processed; determining a second visual prompt instruction corresponding to the second question text based on the image to be processed and the second question text; and generating an answer image by using the pre-trained diffusion model based on the image to be processed and the second visual prompt instruction.
[0007] A third aspect embodiment of the present disclosure provides a visual question answering model training device, including: a first determination unit configured to determine a first output image by using a diffusion model based on a first input image and a first visual prompt instruction of the first input image; a second determination unit configured to determine an instruction learning objective of the diffusion model based on the first input image and the first output image; and a training unit configured to train the diffusion model based on the first input image, the first output image, the first visual prompt instruction, and the instruction learning objective to obtain a pre-trained diffusion model.
[0008] A fourth aspect embodiment of the present disclosure provides a visual question answering device, including: an acquisition unit configured to acquire an image to be processed and a second question text corresponding to the image to be processed; a processing unit configured to determine a second visual prompt instruction corresponding to the second question text based on the image to be processed and the second question text; and a generation unit configured to generate an answer image by using the pre-trained diffusion model based on the image to be processed and the second visual prompt instruction.
[0009] A fifth aspect embodiment of the present disclosure provides an electronic device, including: a processor and a memory configured to store a computer program that can run on the processor, wherein when the processor is configured to run the computer program, it executes the method described in the first aspect embodiment of the present disclosure, or executes the method described in the second aspect embodiment of the present disclosure.
[0010] A sixth aspect embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to cause a computer to execute the method described in the first aspect embodiment of the present disclosure, or execute the method described in the second aspect embodiment of the present disclosure.
[0011] A seventh aspect embodiment of the present disclosure provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, it implements the method described in the first aspect embodiment of the present disclosure, or implements the method described in the second aspect embodiment of the present disclosure.
[0012] By training the diffusion model by using visual prompt instructions, the present disclosure can convey key information in the input image and the question to the diffusion model more accurately. Compared with using text prompts in the prior art, there is no need to perform operations such as text splitting and semantic analysis on the input content, and the features of the question text and the input image can be directly extracted from the matrix parameters of the visual prompt instructions, making it more intuitive for the diffusion model to obtain the features of the question text and the input image, thereby reducing the requirement of the diffusion model for complex text understanding and making the model training process more stable and efficient.
[0013] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. Description of the Drawings
[0014] The drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:
[0015] Figure 1 is a schematic flowchart of a method for training a visual question answering model provided by an embodiment of the present disclosure;
[0016] Figure 2 is a schematic flowchart of another method for training a visual question answering model provided by an embodiment of the present disclosure;
[0017] Figure 3 is a schematic flowchart of yet another method for training a visual question answering model provided by an embodiment of the present disclosure;
[0018] Figure 4 is a schematic flowchart of still another method for training a visual question answering model provided by an embodiment of the present disclosure;
[0019] Figure 5 is a schematic flowchart of a method for applying a visual question answering model provided by an embodiment of the present disclosure;
[0020] Figure 6 is an example diagram of the flowchart of a method for training a visual question answering model provided by an embodiment of the present disclosure;
[0021] Figure 7 is a schematic structural diagram of a device for training a visual question answering model provided by an embodiment of the present disclosure;
[0022] Figure 8 is a schematic structural diagram of a device for applying a visual question answering model provided by an embodiment of the present disclosure;
[0023] Figure 9 is a schematic diagram of the hardware composition structure of an electronic device provided by an embodiment of the present disclosure. Detailed Embodiments
[0024] The following makes an explanation of the exemplary embodiments of the present disclosure, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0025] 1.1 First, the related technologies involved in the present application are described:
[0026] Prompt Learning, as an emerging natural language processing technology, shows great potential by designing specific prompts to guide pre-trained language models to complete downstream tasks. Its core lies in transforming downstream tasks into a cloze problem. By adding a specific prompt to the input text, it guides the model to generate outputs that meet the task requirements. This method cleverly utilizes the rich knowledge contained in pre-trained models and can achieve excellent performance in various natural language processing (NLP) tasks without large-scale fine-tuning of the model.
[0027] In the field of Vision-Language Models (VLMs), it has witnessed remarkable development in recent years, benefiting from the progress of deep learning technology and the wide availability of large-scale datasets. By learning the relationship between images and text, they have demonstrated excellent performance in a variety of visual and language tasks, such as image classification, image generation, image retrieval, and natural language description of images. Representative models include the Contrastive Language-Image Pre-Training (CLIP) model, the A Large-scale Image and Noisy-text embedding (ALIGN) model, and the DALL·E text-to-image generation model, etc. Although these models perform well, they still face some challenges when adapting to downstream tasks, such as the huge computational resource requirements during the fine-tuning process, the fact that some models can only be used through the Application Programming Interface (API), and the issue of low interpretability, etc.
[0028] Visual prompts based on vision-language large models are an emerging research area, that is, guiding the model by adding specific visual information to the input. This approach can not only reduce computational costs but also improve the generalization ability of the model. Visual prompts have a wide range of applications in fields such as image generation, object detection, and image segmentation.
[0029] The question-guided image generation task is a typical cross-modal generation task. Based on the input image and the corresponding question, the model can generate an image to answer the posed question. For example, given an image of a hand pulling a light cord in the on state, and the question "What will happen if the hand pulls down?", the model can generate an image of the light being off based on the input image and question, and the generated image is consistent with the input image, based on common sense or prior knowledge.
[0030] 1.2 Introduction to the models in the related art:
[0031] The Latent guided diffusion (LGD) model predicts the latent space transformation through the cross-attention mechanism of the multi-modal encoder (Large Language Model-encoder, LLM-encoder) for the input image and the question, and generates the result image based on the original image through the diffusion model.
[0032] The Stable Diffusion model encodes the original image into the latent space through the Variational Autoencoder (VAE), and generates high-quality images that highly match the text description by gradually denoising in the latent space under conditions such as text.
[0033] The InstructPix2Pix model realizes precise modification of the image by inputting the input image and the text instruction into a generative adversarial network together. The core innovation of this model is the combination of natural language understanding and image generation, making image editing more flexible and intuitive. Users can perform various image editing operations by simply providing simple text instructions.
[0034] However, the models in the related art usually use text prompts. Compared with visual prompts, text prompts are not intuitive, vivid, or specific enough, and there are some scenarios where it is simply impossible to describe them in language. Moreover, the understanding ability of the LLM-encoder for the question directly affects the generation result, and at the same time, LGD is composed of several large models, consuming huge computing resources.
[0035] To solve the technical problems existing in the related art, an embodiment of the present disclosure provides a method for training a visual question answering model.
[0036] The embodiments of the present disclosure will be introduced in detail below.
[0037] As Figure 1 shown, an embodiment of the present disclosure provides a method for training a visual question answering model, including the following steps:
[0038] Step 101, based on the first input image and the first visual prompt instruction of the first input image, use the diffusion model to determine the first output image.
[0039] In some embodiments, the first visual prompt instruction is used to assist the diffusion model in generating a first output image for answering questions; specifically, the first visual prompt instruction can indicate the visual features in the first input image that are relevant to the question, so that the diffusion model generates the first output image based on the visual features, thereby avoiding the analysis of the question text by the diffusion model.
[0040] In some embodiments, the first input image and the first visual prompt instruction are input into the diffusion model, so that the diffusion model generates a first output image according to the first input image and the first visual prompt instruction.
[0041] In some embodiments, the first input image is the image used to train the diffusion model, and it can be any type of image, such as a landscape image, a portrait image, etc.
[0042] Step 102, based on the first input image and the first output image, determine the instruction learning objective of the diffusion model.
[0043] In some embodiments, the difference between the first input image and the first output image can be determined as the instruction learning objective, where the instruction learning objective can be used to indicate the learning degree of the diffusion model for the first visual prompt instruction.
[0044] Step 103, based on the first input image, the first output image, the first visual prompt instruction, and the instruction learning objective, train the diffusion model to obtain a pre-trained diffusion model.
[0045] In some embodiments, the instruction learning objective can be used as the learning guidance of the diffusion model, and the diffusion model is iteratively trained using the first input image, the first output image, the first visual prompt instruction, and the instruction learning objective, and the diffusion model after the iterative training is completed is determined as the pre-trained diffusion model, so that the pre-trained diffusion model can generate an output image that more conforms to the question description according to the input image and the question text.
[0046] In summary, the visual question answering model training method proposed according to the present disclosure includes: based on the first input image and the first visual prompt instruction of the first input image, using the diffusion model to determine the first output image; based on the first input image and the first output image, determining the instruction learning objective of the diffusion model; based on the first input image, the first output image, the first visual prompt instruction, and the instruction learning objective, training the diffusion model to obtain a pre-trained diffusion model. The method of the present disclosure uses the instruction learning objective as the learning guidance of the diffusion model, so that the diffusion model can understand the first visual prompt instruction through training, so that the pre-trained model can generate an output image that conforms to the question description according to the first visual prompt instruction, avoiding the analysis of the question text in the diffusion model training and improving the training efficiency of the diffusion model.
[0047] As shown in Figure 2 , an embodiment of the present disclosure provides a method for training a visual question answering model, which is further described based on the embodiment shown in Figure 1 , and includes the following steps:
[0048] Step 201: Based on the first input image and the first visual prompt instruction of the first input image, use a diffusion model to determine a first output image.
[0049] Step 202: Perform a first encoding on the first input image to obtain a first input image encoding, and perform a second encoding on the first output image to obtain a first output image encoding.
[0050] In some embodiments, the encoding method of the second encoding may be the same as that of the first encoding, so as to facilitate the diffusion model to determine the difference between the first input image and the first output image.
[0051] In some embodiments, the first encoding and the second encoding may be performed through a CLIP model, but it is not limited thereto. The present disclosure does not limit the specific encoding methods of the first encoding and the second encoding.
[0052] Step 203: Based on the parameter difference between the first input image encoding and the first output image encoding, determine the instruction learning objective of the diffusion model.
[0053] Exemplarily, taking the first encoding and the second encoding using the CLIP model as an example for illustration, the instruction learning objective can be expressed by the following formula:
[0054]
[0055] Where represents the instruction learning objective, represents the first input image, represents the first output image, represents performing a first encoding on the first input image using the CLIP model and performing a second encoding on the first output image using the CLIP model.
[0056] Step 204: Sample a preset noise to obtain a first sampled noise.
[0057] In some embodiments, the preset noise may be noise with a Gaussian distribution having a mean of 0 and a variance of 1, but it is not limited thereto. The present disclosure does not limit the type of the preset noise.
[0058] Step 205: Based on the first output image encoding and the first sampled noise, determine a first noise image encoding.
[0059] In some embodiments, a first noise image can be generated by adding first sampling noise to the first output image encoding.
[0060] It should be understood that during the iterative training of the diffusion model, sampling noise can be added to the output image encoding in each iteration, or sampling noise can be added to some of the output image encodings during the iteration according to the preset number of iteration time steps. The present disclosure does not limit this.
[0061] Exemplarily, taking the set number of iteration time steps as T and the preset noise as ~N(0,1) as an example, the time variable , and the preset noise ~N(0,1) can be sampled, and at the sampling time of the time variable t, the sampling noise of the preset noise is added to the output image encoding of the diffusion model.
[0062] In some embodiments, the first noise image encoding can be determined by the following formula:
[0063]
[0064] Where represents the first noise image encoding, represents the first output image encoding, represents the first sampling noise.
[0065] Step 206, determine the predicted added noise based on the first noise image encoding, the first visual prompt instruction, and the first input image encoding.
[0066] In some embodiments, the predicted added noise can be determined according to the first noise image encoding, the first visual prompt instruction, and the first input image encoding by using a preset denoising model that has been pre-trained.
[0067] In some embodiments, the predicted added noise can be determined by the following formula:
[0068]
[0069] Where represents the predicted added noise, represents the preset denoising model that has been pre-trained, represents the first noise image encoding, represents the first visual prompt instruction, represents the first input image encoding, where can be a denoising model based on a U-shaped neural network convolution (UNet) structure or a transformer architecture. The present disclosure does not limit this.
[0070] Step 207: Multiply the norm of the vector difference between the predicted added noise and the first sampled noise by a preset first loss coefficient to determine the noise loss parameter.
[0071] In some embodiments, the noise loss parameter can be determined by the following formula:
[0072]
[0073] where represents the noise loss parameter, represents the first loss coefficient, represents the first sampled noise, represents the predicted added noise.
[0074] Step 208: Multiply the cosine similarity between the first visual prompt instruction and the instruction learning objective by a preset second loss coefficient to determine the instruction learning loss parameter.
[0075] In some embodiments, the learning loss parameter can be determined by the following formula:
[0076]
[0077] where represents the learning loss parameter, represents the second loss coefficient, represents the cosine similarity function, and the other parameters can be referred to the descriptions in the above embodiments and will not be elaborated here.
[0078] Step 209: Perform a weighted sum of the noise loss parameter and the instruction learning loss parameter to determine the fusion loss parameter.
[0079] In some embodiments, a weighted sum of the noise loss parameter and the instruction learning loss parameter can be performed to determine the fusion loss parameter, or a direct sum of the noise loss parameter and the instruction learning loss parameter can be performed to determine the fusion loss parameter.
[0080] In some embodiments, the fusion loss parameter can be determined by the following formula:
[0081]
[0082] where represents the fusion loss parameter.
[0083] Step 210: Update the first visual prompt instruction based on the first visual prompt instruction and the fusion loss parameter to determine the updated visual prompt instruction.
[0084] In some instances, the gradient of the fusion loss parameter can be calculated to obtain the model gradient parameter of the diffusion model; and then, the difference between the first visual prompt instruction and the model gradient parameter is determined as the updated visual prompt instruction.
[0085] In some embodiments, the updated visual prompt instruction can be determined by the following formula:
[0086]
[0087] where, represents the updated visual prompt instruction, represents the preset update weight, represents the gradient operation.
[0088] Step 211: Based on the updated visual prompt instruction, iteratively train the diffusion model until the iteration reaches the preset maximum number of iteration steps to obtain the pre-trained diffusion model.
[0089] In some embodiments, the updated visual prompt instruction can be used to replace the first visual prompt instruction, and the above steps 201 - 210 are repeatedly executed to achieve one iteration training of the diffusion model.
[0090] In summary, the method for training a visual question answering model proposed in the present disclosure includes: based on a first input image and a first visual prompt instruction for the first input image, using a diffusion model to determine a first output image; performing a first encoding on the first input image to obtain a first input image encoding, and performing a second encoding on the first output image to obtain a first output image encoding; determining an instruction learning objective of the diffusion model based on the parameter difference between the first input image encoding and the first output image encoding; sampling a preset noise to obtain a first sampled noise; determining a first noise image encoding based on the first output image encoding and the first sampled noise; determining a predicted added noise based on the first noise image encoding, the first visual prompt instruction, and the first input image encoding; multiplying the norm of the vector difference between the predicted added noise and the first sampled noise by a preset first loss coefficient to determine a noise loss parameter; multiplying the cosine similarity between the first visual prompt instruction and the instruction learning objective by a preset second loss coefficient to determine an instruction learning loss parameter; performing a weighted sum on the noise loss parameter and the instruction learning loss parameter to determine a fusion loss parameter; updating the first visual prompt instruction based on the first visual prompt instruction and the fusion loss parameter to determine an updated visual prompt instruction; and iteratively training the diffusion model based on the updated visual prompt instruction until the iteration reaches a preset maximum number of iteration steps to obtain a pre-trained diffusion model. The method of the present disclosure, by using the instruction learning objective as the training guidance of the diffusion model, can more directly guide the diffusion model to learn visual features related to the question text, so that the diffusion model pays more attention to the features that can distinguish between the input image and the output image, thereby enabling the pre-trained diffusion model to generate more compliant images and improving the training efficiency of the diffusion model; and this method adds noise during the training process, so that the diffusion model combines the scheduling process of the visual prompt instruction and the preset noise during the training process, thereby dynamically adjusting the intensity and distribution of the noise in the image, and further improving the accuracy of the pre-trained model when outputting images.
[0091] As Figure 3 shown, an embodiment of the present disclosure provides a method for training a visual question answering model, which can be executed before the steps of the embodiment shown in Figure 1 or Figure 2 and includes step 301:
[0092] Step 301, determining a first visual prompt instruction based on a first input image and a first question text of the first input image.
[0093] In some embodiments, the first visual prompt instruction may be generated according to the visual features related to the first question text in the first input image based on the first input image and the first question text.
[0094] In some embodiments, the specific determination method of the first visual prompt instruction may refer to Figure 4 the steps shown, such as Figure 4 shown, including the steps:
[0095] Step 401, divide the first input image equally to obtain the first input sub-images.
[0096] In some embodiments, the first input image can be divided equally by the following formula:
[0097]
[0098] where, represents the first input sub-image, represents the equal division operation function, represents the first input image. Since the first input sub-images are multiple images, the first sub-image can be represented as .
[0099] Step 402, perform third encoding on the first input image to obtain the second input image encoding, and perform fourth encoding on the first input sub-images to obtain the first input sub-image encodings.
[0100] In some embodiments, the encoding methods of the third encoding and the fourth encoding are the same.
[0101] In some embodiments, performing third encoding on the first input image can be represented as:
[0102]
[0103] where, represents the second input image encoding, represents the encoding method of the third encoding.
[0104] In some embodiments, performing fourth encoding on the first input sub-images can be represented as:
[0105]
[0106] where, represents the first input sub-image encoding of the i-th first input sub-image.
[0107] Step 403, perform fifth encoding on the first question text to obtain the first question text encoding.
[0108] In some embodiments, the first question text encoding can be determined by the following formula:
[0109]
[0110] where, Represents the first problem text encoding, Represents the encoding method of the fifth encoding, Represents the problem text.
[0111] Step 404: Use a splicing function to splice the second input image encoding and the first input sub-image encoding to determine the fused input image encoding.
[0112] In some embodiments, the fused input image encoding can be determined by the following formula:
[0113]
[0114] where E represents the fused input image encoding, represents the splicing function.
[0115] Step 405: Use the preset query information to extract the visual features of the first problem text from the fused input image encoding and the first problem text encoding.
[0116] In some embodiments, the visual features related to the first problem text can be extracted from the fused input image encoding and the first problem text encoding through the preset query information (i.e., a set of learnable preset queries).
[0117] In some embodiments, the visual features can be extracted by the following formula:
[0118]
[0119] where, represents the visual features, represents a transformer, represents the preset query information.
[0120] Step 406: Based on the diffusion model, map the visual features to obtain the first visual prompt instruction.
[0121] In some embodiments, since the extracted visual features have different dimensions from the diffusion model, it is necessary to map the visual features according to the model dimension of the diffusion model and determine the mapped visual features as the first visual prompt instruction.
[0122] In summary, the method for training a visual question answering model proposed according to the present disclosure includes: equally dividing a first input image to obtain first input sub-images; performing third encoding on the first input image to obtain a second input image encoding, and performing fourth encoding on the first input sub-images to obtain first input sub-image encodings; performing fifth encoding on the first question text to obtain a first question text encoding; using a splicing function to splice the second input image encoding and the first input sub-image encodings to determine a fused input image encoding; using preset query information to extract visual features of the first question text from the fused input image encoding and the first question text encoding; and mapping the visual features based on a diffusion model to obtain a first visual prompt instruction. The method of the present disclosure encodes the first input image and the first question text to obtain visual features related to the first question text, and then uses the visual features to generate a first visual prompt instruction, thereby realizing the generation of the first visual prompt instruction.
[0123] As Figure 5 shown, an embodiment of the present disclosure provides a method for applying a visual question answering model, including the following steps:
[0124] Step 501, obtain an image to be processed and a second question text corresponding to the image to be processed.
[0125] Step 502, based on the image to be processed and the second question text, determine a second visual prompt instruction corresponding to the second question text.
[0126] In some embodiments, the manner of determining the second visual prompt instruction is similar to Figure 4 the manner of determining the first visual prompt instruction in the Figure 4 shown embodiment, and reference may be made to the relevant description in the
[0127] shown embodiment, which will not be elaborated here.
[0128] Step 503, based on the image to be processed and the second visual prompt instruction, use a pre-trained diffusion model to generate an answer image.
[0129] In summary, the method for applying a visual question answering model proposed by the present disclosure includes: obtaining an image to be processed and a second question text corresponding to the image to be processed; determining a second visual prompt instruction corresponding to the second question text based on the image to be processed and the second question text; and generating an answer image by using a pre-trained diffusion model based on the image to be processed and the second visual prompt instruction. In the method of the present disclosure, by using the second visual prompt instruction to generate the answer image, the direct analysis of the question text by the pre-trained diffusion model is avoided, the requirement for the pre-trained model to understand complex text when generating the answer image is reduced, the efficiency of generating the answer image is improved, and the resource consumption of the pre-trained diffusion model when generating the answer image is reduced.
[0130] The following is an exemplary description of the present disclosure:
[0131] I. Visual Instruction Generator.
[0132] The Visual Instruction Generator. That is, according to the input image and the corresponding question , relevant visual cues to the question are obtained through a lightweight transformer such as Querying Transformer (Q-Former). These can be used as prior information to make the later optimization more targeted, thereby reducing the training difficulty of question-to-image generation and improving the convergence speed of the model. As Figure 6 shown, the specific steps are as follows:
[0133] 1. Encode the input image. For the question-to-image generation task, the original input image and the generated image have strong similarity, but the local details are different due to common sense reasoning and causal relationships. Therefore, in order to capture these details, without changing the encoding of the original input image, the original input image (i.e., the above-mentioned first input image) is equally divided to obtain equally divided images (i.e., the above-mentioned first input sub-images), and the original input image and the equally divided images are encoded simultaneously and these encodings are integrated. Specifically as follows:
[0134] 1.1 Encode the original input image.
[0135] The existing image encoder image_encoder is used to encode the original input image (i.e., the above-mentioned third encoding) to obtain the original input image encoding (i.e., the above-mentioned second input image encoding), that is , where the original input image encoding can be composed of multiple encoding blocks m1, m2... mn. Here, image_encoder can be img_encoder of CLIP or others.
[0136] 1. Divide the original input image into two equal parts and encode it.
[0137] First, the original input image is divided into equal parts. , Let \(I\) be the input image. Let \(f(I)\) be the function for dividing the image into equal parts. Let \(I_i\) be the result of dividing the original input image into equal parts, where ; Second, the divided image is encoded using the same encoding method as the original input image (i.e., the fourth encoding above) to obtain the encoded divided image (i.e., the first input sub-image encoding above), that is , where \(E_{I_i}\) is the encoding result of the \(i\)-th divided image of the original input image. Among them, the encoding result of the divided image \(E_{I_i}\) can be composed of multiple encoding blocks \(n_{i1}, n_{i2}, \cdots, n_{in}\).
[0138] 1.3 Integrate the encoding of the original input image and the encoding of the divided image after encoding.
[0139] , where \(E_I\) is the encoding result of the original input image. \(E_{I_i}\) is the encoding result of the \(i\)-th divided image. Let \(S\) be the splicing operation, that is, by splicing the encoding of the original input image and the encoding of the divided image to obtain the integrated image encoding \(E\) (i.e., the fused input image encoding above), where the integrated image encoding \(E\) can be composed of multiple encoding blocks \(p_1, p_2, \cdots, p_n\).
[0140] 2. Encoding problem.
[0141] Encode the problem text (i.e., the fifth encoding above), that is , \(E_{question}\) is the feature encoding of the input question (i.e., the first question encoding above). Let \(T\) be a common text encoder, such as the text encoder of CLIP.
[0142] 3. Obtain the visual features related to the question. That is, through a lightweight transformer structure, a set of learnable queries (i.e., the preset query information above) and the question are used to extract the related visual features (i.e., the visual features of the first text above) from the frozen image encoder. The following formula is used here:
[0143] Q-Former, .
[0144] Among them, \(L_Q\) is the learnable query. Encode the problem, and E is the visual encoding. Among them, the above set of learnable queries can be composed of multiple preset query information q1, q2... qn, and the above visual features can be composed of multiple feature information groups r1, r2... rn.
[0145] 4. Use a fully connected layer to map to the dimension required by the pre-trained denoising model . That is , the visual prompt instruction (that is, the above first visual prompt instruction) is the result after being processed by the fully connected layer. Among them, the above visual prompt instruction can be composed of multiple instruction information V1, V2... Vn.
[0146] II. Visual Prompt Instruction Learning
[0147] The visual prompt instruction generator can use the attention mechanism in the lightweight transformer to interact the learnable query and the problem with the visual information. Although this method can focus on the problem-related area and extract the corresponding visual information, this information is not sufficient as a visual prompt to guide the model to generate a suitable image based on the input image and the problem. Therefore, this information can be used as prior information to guide the diffusion model to learn more matching visual prompts. The following takes the difference between the input image and the output image as the goal, combines the original input image, the problem, and the learnable query, and learns the visual prompt based on the diffusion model. The specific steps are as follows:
[0148] 1. Encode the input image (that is, the above first encoding) to obtain the input image encoding (that is, the above first input image encoding), and encode the output image (that is, the above second encoding) to obtain the output image encoding (that is, the first output image encoding), that is, use the same image encoding method as to encode the output image , that is .
[0149] 2. Calculate the difference between the input image and the output image and use it as the target for visual prompt instruction learning (that is, the above instruction learning target). Here, the CLIP model is used to calculate the difference between them .
[0150] 3. Learn visual prompts Since it is not possible to better capture more precise editing directions based on text prompts, on the basis of the original latent space , integrate learnable visual prompts and the output image , learn and optimize visual prompts through the denoising process , let the pre-trained denoising model be , the number of optimization iteration steps is N, and the number of time steps is T. The specific steps are as follows:
[0151] 3.1. Execute steps 3 and 4 of the visual prompt instruction generator to obtain visual prompt instructions related to the problem and learnable parameters ;
[0152] 3.2. Sample from the time of the uniform distribution and the noise of the Gaussian distribution (i.e., the above-mentioned preset noise), that is, the time variable can take any parameter in [0, T], and the selection probability of each parameter is equal. The noise ~ N(0, 1) is a Gaussian distribution random variable with a mean parameter of 0 and a variance of 1.
[0153] 3.3. At time step t, add the noise to the encoding corresponding to the output image , that is (i.e., the above-mentioned preset noise image encoding).
[0154] 3.4. Predict the added noise through the pre-trained denoising model, that is , that is, the above-mentioned predicted added noise. Here, is a basic denoising model, which can be based on the UNet structure or the transformer architecture.
[0155] 3.5. Calculate the (i.e., the above-mentioned noise loss parameter) between the predicted noise and the real noise. Here, the MSE loss is used, that is , where is the loss coefficient.
[0156] 3.6. Calculate the loss (i.e., the above-mentioned learning loss parameter) between the difference between the input and output images and the visual prompt , that is , where is the loss coefficient.
[0157] 3.7. Calculate the total loss (i.e., the above-mentioned fusion loss parameter) and update the gradient, that is .
[0158] 3.8. Update , that is .
[0159] To implement the visual question answering model training method provided by the embodiments of the present disclosure, the embodiments of the present disclosure also provide a visual question answering model training device, as Figure 7 shown, the visual question answering model training device 700 includes:
[0160] A first determination unit 701, configured to determine a first output image by using a diffusion model based on a first input image and a first visual prompt instruction of the first input image;
[0161] A second determination unit 702, configured to determine an instruction learning target of the diffusion model based on the first input image and the first output image;
[0162] A training unit 703, configured to train the diffusion model based on the first input image, the first output image, the first visual prompt instruction, and the instruction learning target to obtain a pre-trained diffusion model.
[0163] In some embodiments, the second determination unit 702 is further configured to perform a first encoding on the first input image to obtain a first input image encoding, and perform a second encoding on the first output image to obtain a first output image encoding; determine the instruction learning target of the diffusion model based on the parameter difference between the first input image encoding and the first output image encoding.
[0164] In some embodiments, the training unit 703 is further configured to sample a preset noise to obtain a first sampled noise; determine a first noise image encoding based on the first output image encoding and the first sampled noise; determine a predicted added noise based on the first noise image encoding, the first visual prompt instruction, and the first input image encoding; perform iteration on the diffusion model based on the predicted added noise, the first sampled noise, the first visual prompt instruction, and the instruction learning target to obtain a pre-trained diffusion model.
[0165] In some embodiments, the training unit 703 is further configured to multiply the norm of the vector difference between the predicted added noise and the first sampled noise by a preset first loss coefficient to determine a noise loss parameter; multiply the cosine similarity between the first visual prompt instruction and the instruction learning target by a preset second loss coefficient to determine an instruction learning loss parameter; perform weighted summation on the noise loss parameter and the instruction learning loss parameter to determine a fusion loss parameter; update the first visual prompt instruction based on the first visual prompt instruction and the fusion loss parameter to determine an updated visual prompt instruction; perform iterative training on the diffusion model based on the updated visual prompt instruction until the iteration reaches a preset maximum number of iteration steps to obtain a pre-trained diffusion model.
[0166] In some embodiments, the training unit 703 is further configured to calculate the gradient of the fusion loss parameter to obtain the model gradient parameter of the diffusion model; and determine the updated visual prompt instruction based on the difference between the first visual prompt instruction and the model gradient parameter.
[0167] In some embodiments, the first determination unit 701 is further configured to determine the first visual prompt instruction based on the first input image and the first question text of the first input image.
[0168] In some embodiments, the first determination unit 701 is further configured to equally divide the first input image to obtain the first input sub-images; perform a third encoding on the first input image to obtain the second input image encoding, and perform a fourth encoding on the first input sub-images to obtain the first input sub-image encoding; perform a fifth encoding on the first question text to obtain the first question text encoding; and determine the first visual prompt instruction based on the second input image encoding, the first input sub-image encoding, and the first question text encoding.
[0169] In some embodiments, the first determination unit 701 is further configured to use a splicing function to splice the second input image encoding and the first input sub-image encoding to determine the fused input image encoding; use the preset query information to extract the visual features of the first question text from the fused input image encoding and the first question text encoding; and map the visual features based on the diffusion model to obtain the first visual prompt instruction.
[0170] In summary, the visual question answering model training device proposed by the present disclosure includes: a first determination unit configured to determine a first output image based on the first input image and the first visual prompt instruction of the first input image using a diffusion model; a second determination unit configured to determine the instruction learning objective of the diffusion model based on the first input image and the first output image; and a training unit configured to train the diffusion model based on the first input image, the first output image, the first visual prompt instruction, and the instruction learning objective to obtain a pre-trained diffusion model. The device of the present disclosure uses the instruction learning objective as the learning guidance of the diffusion model, enabling the diffusion model to understand the first visual prompt instruction through training, so that the pre-trained model can generate an output image that conforms to the question description according to the first visual prompt instruction, avoiding the analysis of the question text during the training of the diffusion model and improving the training efficiency of the diffusion model.
[0171] To implement the visual question answering model training method provided by the embodiments of the present disclosure, the embodiments of the present disclosure further provide a visual question answering model application device, as Figure 8 shown, the visual question answering model application device 800 includes:
[0172] An acquisition unit 801 that acquires the image to be processed and the second question text corresponding to the image to be processed;
[0173] A processing unit 802, configured to determine a second visual prompt instruction corresponding to the second question text based on the image to be processed and the second question text;
[0174] A generating unit 803, configured to generate an answer image by using a pre-trained diffusion model based on the image to be processed and the second visual prompt instruction.
[0175] In some embodiments, the generating unit 803 is further configured to determine an editing area and an editing direction of the image to be processed based on the second visual prompt instruction; and edit the image to be processed based on the editing area and the editing direction to generate an answer image.
[0176] In summary, the visual question answering model application device proposed according to the present disclosure includes: an acquisition unit that acquires an image to be processed and a second question text corresponding to the image to be processed; a processing unit that is configured to determine a second visual prompt instruction corresponding to the second question text based on the image to be processed and the second question text; and a generating unit that is configured to generate an answer image by using a pre-trained diffusion model based on the image to be processed and the second visual prompt instruction. The device of the present disclosure generates an answer image by using the second visual prompt instruction, avoiding the direct analysis of the question text by the pre-trained diffusion model, reducing the requirement for the pre-trained model to understand complex text when generating the answer image, improving the efficiency of generating the answer image, and reducing the resource consumption of the pre-trained diffusion model when generating the answer image.
[0177] It should be noted that: when the visual question answering model training device provided in the above embodiment performs visual question answering model training, or when the visual question answering model application device provided in the above embodiment performs visual question answering model application, only the above-mentioned division of each program module is used for illustration. In actual application, the above-mentioned processing can be allocated to different program modules according to needs, that is, the internal structure of the visual question answering model training device is divided into different program modules to complete all or part of the above-described processing. In addition, the visual question answering model training device provided in the above embodiment and the method embodiment of the visual question answering model training method provided in the embodiment of the present disclosure belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.
[0178] Figure 9 It is a schematic diagram of the hardware composition structure of the electronic device provided in the embodiment of the present disclosure. As Figure 9 shown, the electronic device 900 includes at least one processor 902; and a memory 901 communicatively connected to the at least one processor 902; wherein, the memory 901 stores instructions executable by the at least one processor 902, and the instructions are executed by the at least one processor 902 to implement the steps of the visual question answering model training method provided in the embodiment of the present disclosure.
[0179] Optionally, the electronic device may specifically be the visual question answering model training device according to an embodiment of the present application, and the electronic device may implement the corresponding processes in each method according to the embodiment of the present application that are implemented by the visual question answering model training device. For the sake of brevity, details are not described herein again.
[0180] It can be understood that the electronic device further includes a communication interface 903. Each component in the electronic device is coupled together through a bus system 904. It can be understood that the bus system 904 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 904 further includes a power bus, a control bus, and a status signal bus. However, for the sake of clear description, in Figure 9 all kinds of buses are labeled as the bus system 904.
[0181] It can be understood that the memory 901 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM), a synchronous static random access memory (SSRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a sync link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM).The memory 901 described in the embodiments of the present invention is intended to include but not limited to these and any other suitable types of memories.
[0182] The methods disclosed in the above embodiments of the present disclosure can be applied to or implemented by the processor 902. The processor 902 has the ability to process signals. During implementation, each step of the above method can be completed by the integrated logic circuit in hardware or the instructions in software form in the processor 902. The above-mentioned processor 902 can be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 902 can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module can be located in the storage medium, which is located in the memory 901. The processor 902 reads the information in the memory 901 and combines its hardware to complete the steps of the foregoing method.
[0183] In an exemplary embodiment, the electronic device can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components, and is used to execute the foregoing method.
[0184] The embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the steps of the visual question-answering model training method described in the embodiments of the present disclosure when executed.
[0185] The embodiments of the present disclosure also provide a computer program product, including a computer program, and the computer program implements the steps of the visual question-answering model training method described in the embodiments of the present disclosure when executed by a processor.
[0186] Optionally, the computer-readable storage medium can be applied to the visual question-answering model training device in the embodiments of the present application, and the computer instructions cause the computer to execute the corresponding processes implemented by the visual question-answering model training device in each method of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.
[0187] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be electrical, mechanical, or other forms.
[0188] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0189] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0190] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0191] Alternatively, if the above-mentioned integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention essentially or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present invention. And the foregoing storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0192] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims described above.
Claims
1. A visual question answering model training method, characterized in that: The method comprises: Determine a first output image based on a first input image and a first visual cue instruction of the first input image using a diffusion model; determining an instruction learning target of the diffusion model based on the first input image and the first output image; Training the diffusion model based on the first input image, the first output image, the first visual cue instruction, and the instruction learning target to obtain a pre-trained diffusion model; Wherein, the method further comprises: Dividing the first input image into equal parts to obtain first input sub-images; Performing a third encoding on the first input image to obtain a second input image code, and performing a fourth encoding on the first input sub-image to obtain a first input sub-image code; Performing fifth encoding on the first question text of the first input image to obtain a first question text encoding; The first visual prompt instruction is determined based on the second input image code, the first input sub-image code and the first question text code.
2. The method according to claim 1, characterized in that The step of determining the instruction learning target of the diffusion model based on the first input image and the first output image includes: Performing a first encoding on the first input image to obtain a first input image code, and performing a second encoding on the first output image to obtain a first output image code; An instruction learning target for the diffusion model is determined based on a parameter difference between the first input image encoding and the first output image encoding.
3. The method according to claim 1, characterized in that The step of training the diffusion model based on the first input image, the first output image, the first visual cue instruction, and the instruction learning target to obtain a pre-trained diffusion model includes: Sampling the preset noise to obtain a first sampled noise; determining a first noise image code based on the first output image code and the first sampled noise; Determining a predicted added noise based on the first noise image encoding, the first visual cue instruction, and the first input image encoding; The diffusion model is iterated based on the predicted added noise, the first sampled noise, the first visual cue instruction and the instruction learning target to obtain the pre-trained diffusion model.
4. The method according to claim 3, characterized in that The iterating the diffusion model based on the predicted added noise, the first sampled noise, the first visual cue instruction and the instruction learning target to obtain the pre-trained diffusion model comprises: Multiplying the norm of the vector difference between the predicted added noise and the first sampled noise by a preset first loss coefficient to determine a noise loss parameter; Multiplying the cosine similarity between the first visual cue instruction and the instruction learning target by a preset second loss coefficient to determine an instruction learning loss parameter; Performing a weighted summation on the noise loss parameter and the instruction learning loss parameter to determine a fusion loss parameter; Based on the first visual prompt instruction and the fusion loss parameter, updating the first visual prompt instruction to determine an updated visual prompt instruction; Based on the update visual prompt instruction, the diffusion model is iteratively trained until the iteration reaches a preset maximum number of iteration steps to obtain a pre-trained diffusion model.
5. The method according to claim 4, characterized in that The updating of the first visual prompt instruction based on the first visual prompt instruction and the fusion loss parameter to determine the updated visual prompt instruction comprises: Performing gradient calculation on the fusion loss parameter to obtain a model gradient parameter of the diffusion model; The updated visual prompt instruction is determined by taking the difference between the first visual prompt instruction and the model gradient parameter.
6. The method according to claim 1, characterized in that The determining the first visual prompt instruction based on the second input image code, the first input sub-image code and the first question text code comprises: Using a splicing function, splicing the second input image code and the first input sub-image code to determine a fused input image code; Extracting visual features of the first question text from the fused input image code and the first question text code using preset query information; Based on the diffusion model, the visual features are mapped to obtain the first visual prompt instruction.
7. A method for applying a visual question answering model, characterized in that: The method comprises: Acquire an image to be processed and a second question text corresponding to the image to be processed; Determining a second visual prompt instruction corresponding to the second question text based on the image to be processed and the second question text; Based on the image to be processed and the second visual prompt instruction, using a pre-trained diffusion model, generating a response image; The step of determining the second visual prompt instruction corresponding to the second question text based on the image to be processed and the second question text includes: Dividing the image to be processed into equal parts to obtain sub-images to be processed; Performing a third encoding on the image to be processed to obtain a code of the image to be processed, and performing a fourth encoding on the sub-image to be processed to obtain a code of the sub-image to be processed; Performing a fifth encoding on the second question text to obtain a second question text encoding; Determine the second visual prompt instruction based on the code of the image to be processed, the code of the sub-image to be processed and the code of the second question text; The training process of the pre-trained diffusion model is as follows: Determine a first output image based on a first input image and a first visual cue instruction of the first input image using a diffusion model; determining an instruction learning target of the diffusion model based on the first input image and the first output image; The diffusion model is trained based on the first input image, the first output image, the first visual cue instruction, and the instruction learning target to obtain a pre-trained diffusion model.
8. The method according to claim 7, characterized in that The determining, based on the image to be processed and the second question text, a second visual prompt instruction corresponding to the second question text comprises: Determining an editing area and an editing direction of the image to be processed based on the second visual prompt instruction; Based on the editing area and the editing direction, the image to be processed is edited to generate the answer image.
9. A visual question answering model training device, characterized in that: include: A first determining unit, configured to determine a first output image based on a first input image and a first visual cue instruction of the first input image by using a diffusion model; a second determining unit, configured to determine an instruction learning target of the diffusion model based on the first input image and the first output image; a training unit, configured to train the diffusion model based on the first input image, the first output image, the first visual cue instruction, and the instruction learning target to obtain a pre-trained diffusion model; Among them, the first determination unit is also used to divide the first input image into equal parts to obtain a first input sub-image; perform a third encoding on the first input image to obtain a second input image encoding, and perform a fourth encoding on the first input sub-image to obtain a first input sub-image encoding; perform a fifth encoding on the first question text of the first input image to obtain a first question text encoding; and determine the first visual prompt instruction based on the second input image encoding, the first input sub-image encoding and the first question text encoding.
10. A visual question answering model application device, characterized in that: include: An acquisition unit, which acquires an image to be processed and a second question text corresponding to the image to be processed; A processing unit, configured to determine a second visual prompt instruction corresponding to the second question text based on the image to be processed and the second question text; A generating unit, configured to generate a response image based on the image to be processed and the second visual prompt instruction by using a pre-trained diffusion model; Wherein, the processing unit comprises: A sub-image to be processed acquisition sub-unit, used for equally dividing the image to be processed to acquire sub-images to be processed; a sub-image code acquisition sub-unit to be processed, configured to perform a third encoding on the image to be processed to obtain a code of the image to be processed, and to perform a fourth encoding on the sub-image to be processed to obtain a code of the sub-image to be processed; A second question text code obtaining subunit, used for performing a fifth code on the second question text to obtain a second question text code; A second visual prompt instruction determining subunit, configured to determine the second visual prompt instruction based on the to-be-processed image code, the to-be-processed sub-image code and the second question text code; The training process of the pre-trained diffusion model is as follows: Determine a first output image based on a first input image and a first visual cue instruction of the first input image using a diffusion model; determining an instruction learning target of the diffusion model based on the first input image and the first output image; The diffusion model is trained based on the first input image, the first output image, the first visual cue instruction, and the instruction learning target to obtain a pre-trained diffusion model.
11. An electronic device, characterized in that: include: a processor and a memory for storing a computer program capable of being executed on the processor, Wherein, when the processor is used to run the computer program, it executes the visual question answering model training method according to any one of claims 1-6, or executes the visual question answering model application method according to claim 7 or 8.
12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the visual question answering model training method according to any one of claims 1 to 6, or to execute the visual question answering model application method according to claim 7 or 8.
13. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the visual question answering model training method according to any one of claims 1 to 6, or implements the visual question answering model application method according to claim 7 or 8.
Citation Information
Patent Citations
Text-to-image generation model optimization method and device, equipment and storage medium
CN116611496A
Image-text generation model training method, visual text image generation method and device
CN119251354A
Image editing method and device, storage medium and electronic equipment
CN119399327A