Image generation method and device based on diffusion implicit model, equipment and medium

By combining a diffusion implicit model and a multimodal large language model, target images that conform to textual descriptions are automatically generated, solving the problems of cumbersome and inefficient generation processes in existing technologies and achieving efficient and reliable image generation.

CN121053245BActive Publication Date: 2026-02-17XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511596163.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-02-17
Estimated Expiration
2045-11-04

AI Technical Summary

Technical Problem

In existing technologies, the process of generating target images that match the textual description information is cumbersome, relies on manual processing, resulting in low efficiency and susceptibility to human intervention.

Method used

An image generation method based on a diffusion implicit model is adopted. By obtaining the keyword location information in the text description information, the image is fused using a multimodal large language model and a diffusion implicit model, and the target image that conforms to the text description information is generated through a loss model optimization.

Benefits of technology

It eliminates the need for manual processing, significantly improving the efficiency of target image generation, reducing generation time, and enhancing the reliability and accuracy of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053245B_ABST
    Figure CN121053245B_ABST
Patent Text Reader

Abstract

This application relates to the field of image generation, and discloses an image generation method, apparatus, device, and medium based on a diffusion implicit model. The method includes: acquiring textual description information and multiple keywords within the textual description information, and determining the location information of each keyword; fusing a base image and the final image of each keyword to obtain a fused image; determining image description information based on the feature vector of the fused image, the feature vector of the question, the feature vector of the answer, and a multimodal large language model; selecting the feature vector of a representative word in the image description information as a representative token, and selecting the maximum value of the cosine similarity between each visual token and the representative token as a reward value; adding the reward value and the confidence level to obtain a loss value; modifying the fused image based on the loss value; and selecting the modified fused image as the target image that conforms to the textual description information. This application is beneficial for improving the generation efficiency of target images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image generation, and in particular to image generation methods, apparatus, devices and media based on diffusion implicit models. Background Technology

[0002] Target images provide a clear direction and framework for information delivery and understanding, ensuring the focus and depth of the content. Through intuitive and vivid visual forms, target images transform abstract concepts and complex information into easily perceptible content, enabling readers to quickly focus on the core points and avoid information redundancy and dispersion.

[0003] However, the current process for generating target images that match textual descriptions is cumbersome, hindering efficiency. This is because existing technologies ignore the positional relationships of keywords within the textual descriptions, thus failing to generate matching images. Currently, this relies on manual processing, which means that every step, from conceptualization to drawing and repeated revisions, is highly dependent on human intervention. This requires creators to possess solid professional skills and invest significant time and effort, increasing the generation time and making the process susceptible to human interference, ultimately hindering efficiency. Summary of the Invention

[0004] This application provides an image generation method, apparatus, device, and medium based on a diffusion implicit model to solve the technical problem that the existing generation process of target images that conform to textual description information is cumbersome and not conducive to improving the generation efficiency of target images.

[0005] In a first aspect, embodiments of this application provide an image generation method based on a diffusion implicit model, applied to an electronic device, the image generation method comprising:

[0006] Retrieve text description information and multiple keywords from the text description information, and determine the location information of each keyword through a predefined method;

[0007] The textual description information is input into the image generation model to obtain the base image. The base image and the final image of each keyword are then fused to obtain the fused image.

[0008] The question and answer texts are obtained, and the image description information is determined based on the feature vectors of the fused image, the feature vectors of the question and the answer, and a multimodal large language model.

[0009] The feature vectors of representative words in the image description information are selected as representative tokens, and the maximum value of the cosine similarity between each visual token and the representative token is selected as the reward value.

[0010] According to the loss model, the reward value and confidence level are added together to obtain the loss value. Based on the loss value and the preset method, the fused image is modified, and the modified fused image is selected as the target image that conforms to the text description information.

[0011] In one possible implementation of the first aspect, obtaining textual description information and multiple keywords within the textual description information, and determining the location information of each keyword through a predefined method, includes:

[0012] Retrieve text description information and multiple keywords from the text description information;

[0013] Each keyword is expanded using a multimodal large language model to obtain the expanded text of each keyword, and the position information of each keyword is obtained from the expanded text of each keyword.

[0014] In one possible implementation of the first aspect, the step of inputting textual description information into an image generation model to obtain a base image, and fusing the base image and the final image of each keyword to obtain a fused image, includes:

[0015] Input the text description information into the image generation model to obtain the base image. Input the expanded text of each keyword and the location information of each keyword into the image generation model to obtain the initial image of each keyword.

[0016] The initial image of each keyword is denoised using a diffusion implicit model to obtain the final image of each keyword. The base image and the final image of each keyword are then fused to obtain a fused image.

[0017] In one possible implementation of the first aspect, the acquisition of question and answer text, based on the feature vectors of the fused image, the feature vectors of the question and the answer, and a multimodal large language model, determines image description information, including:

[0018] Extract question and answer texts from the visual question-and-answer text provided by the user; extract features from the question text to obtain the feature vector of the question; extract features from the answer text to obtain the feature vector of the answer text.

[0019] The feature vectors of the fused image, the question, and the answer are input into a multimodal large language model. The multimodal large language model processes the feature vectors of the fused image, the question, and the answer to obtain image description information.

[0020] In one possible implementation of the first aspect, selecting the feature vector of the representative word in the image description information as the representative token, and selecting the maximum value of the cosine similarity between each visual token and the representative token as the reward value, includes:

[0021] The fused image is divided into multiple image blocks, and the feature vector of each image block is selected as each visual token. The feature vector of the representative word in the image description information is selected as the representative token.

[0022] Obtain the cosine similarity between each visual token and the representative token, and select the maximum value of the cosine similarity between each visual token and the representative token as the reward value.

[0023] In one possible implementation of the first aspect, the step of adding the reward value and the confidence level according to the loss model to obtain a loss value, modifying the fused image based on the loss value and a preset method, and selecting the modified fused image as the target image that conforms to the textual description information includes:

[0024] According to the loss model, the reward value and the confidence level are added together to obtain the loss value;

[0025] The model parameters of the diffusion implicit model are updated to obtain the updated model parameters. The diffusion implicit model is then controlled to modify the fused image using the updated model parameters until the loss value is less than the preset value. Only then is the modified fused image saved, and the modified fused image is selected as the target image that conforms to the text description information.

[0026] In one possible implementation of the first aspect, the step of adding the reward value and the confidence level according to the loss model to obtain the loss value includes:

[0027] The fused image is input into the image encoder, which processes the fused image to obtain the graph vector of the fused image. The image description information is input into the text encoder, which processes the image description information to obtain the text vector. The graph vector and the text vector are fused using a bidirectional cross-attention mechanism to obtain the fused feature.

[0028] The fused features are input into a multimodal classifier, which processes the fused features to obtain the confidence score. Based on the loss model, the reward value and the confidence score are added together to obtain the loss value.

[0029] Secondly, embodiments of this application provide an image generation apparatus based on a diffusion implicit model, applied to an electronic device, comprising:

[0030] The acquisition module is used to acquire text description information and multiple keywords in the text description information, and determine the location information of each keyword through a predefined method;

[0031] The fusion module is used to input textual description information into the image generation model to obtain a base image, and then fuse the base image with the final image of each keyword to obtain a fused image.

[0032] The determination module is used to obtain the question text and answer text, and determine the image description information based on the feature vectors of the fused image, the feature vectors of the question and the answer, and the multimodal large language model;

[0033] The selection module is used to select the feature vector of the representative word in the image description information as the representative token, and select the maximum value of the cosine similarity between each visual token and the representative token as the reward value.

[0034] The generation module is used to add the reward value and confidence level according to the loss model to obtain the loss value, modify the fused image based on the loss value and a preset method, and select the modified fused image as the target image that conforms to the text description information.

[0035] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the image generation method of any of the first aspects described above.

[0036] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the image generation method of any one of the first aspects described above.

[0037] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the image generation method described in any one of the first aspects.

[0038] The beneficial effects of this application embodiment are twofold. Firstly, based on the loss model, the reward value and confidence level are added together to obtain the loss value. The fused image is then modified based on the loss value and a preset method. The modified fused image is selected as the target image that conforms to the text description information. Since no manual processing is required, the generation time of the target image that conforms to the text description information is reduced, which is beneficial to improving the generation efficiency of the target image that conforms to the text description information. Secondly, since it is not affected by human intervention, it is beneficial to improve the reliability of the target image that conforms to the text description information. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is an application scenario diagram of the image generation method provided in the embodiments of this application;

[0041] Figure 2 This is a schematic flowchart of the image generation method provided in the embodiments of this application;

[0042] Figure 3 This is a flowchart illustrating the implementation of S205 in an embodiment of this application.

[0043] Figure 4 A schematic block diagram of an image generation apparatus provided in an embodiment of this application;

[0044] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0046] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0047] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0048] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0049] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0050] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0051] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0052] Furthermore, the technical solutions of the various embodiments can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0053] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0054] The image generation method provided in this application can be applied to electronic devices such as servers, mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application does not impose any restrictions on the specific type of electronic device.

[0055] Please see Figure 1 , Figure 1 The application scenario diagram of the image generation method provided in the embodiments of this application is described in detail below:

[0056] Electronic devices are connected to microphones to receive users' voices and use speech recognition technology to convert the speech into text descriptions.

[0057] In this embodiment of the application, the electronic device connected to the microphone uses speech recognition technology to convert speech into text description information, thereby improving the efficiency and convenience of text description information input.

[0058] Please see Figure 2 , Figure 2 This is a flowchart illustrating the image generation method provided in an embodiment of this application, which can be applied to electronic devices.

[0059] like Figure 2 As shown, the image generation method provided in this application includes the following steps, which are detailed below:

[0060] S201, Obtain text description information and multiple keywords in the text description information, and determine the position information of each keyword through a predefined method;

[0061] The step of obtaining text description information and multiple keywords within the text description information, and determining the location information of each keyword through a predefined method, includes:

[0062] Retrieve text description information and multiple keywords from the text description information;

[0063] Each keyword is expanded using a multimodal large language model to obtain the expanded text of each keyword, and the position information of each keyword is obtained from the expanded text of each keyword.

[0064] For ease of explanation, the following example is provided:

[0065] The textual description is: A nobleman stands in a castle. The keywords in the textual description are: nobleman, castle. The words "nobleman" and "castle" are expanded using a multimodal large language model.

[0066] The expanded text of the nobleman's name is: A young marquis in a deep red velvet suit is standing on the arched terrace on the second floor of the castle.

[0067] The expanded text of the castle reads: This greyish-white castle, built in the 15th century, stands majestically atop a cliff overlooking the sea, with steep rock walls on three sides and only a winding stone staircase on the east side leading down to the port town below.

[0068] For ease of explanation, the following example is provided:

[0069] The text description is: "A cat on an old street." The keywords for this text description are: "old street," "cat." The description of "old street" and "cat" is expanded using a multimodal large language model.

[0070] The expanded description of the old street is as follows: This east-west oriented old street is located in the old town area in the east of the city. On the north side is a row of six-story residential buildings built in the 1990s, with various small shops on the ground floor; on the south side is a construction site, with advertisements plastered all over the blue iron sheet fence. In the middle of the street is a crossroads. Under the old locust tree at the southeast corner, elderly people always gather to play chess, while at the bus stop at the northwest corner, the morning rush hour crowd anxiously looks for oncoming traffic.

[0071] The expanded text of the cat reads: A short-haired gray and white cat is sitting on the drain cover at the edge of the sidewalk on the old street. Its front paws are together, its tail is elegantly wrapped around its body, and its amber eyes are fixed on the barbecue stalls across the street that smell of delicious food.

[0072] For example, obtaining text description information and multiple keywords from the text description information includes:

[0073] Electronic devices are connected to microphones to receive user voice messages. Voice recognition technology is used to convert the voice messages into text descriptions, and the text descriptions and multiple keywords within them are extracted.

[0074] For example, each keyword is expanded using a multimodal large language model to obtain the expanded text of each keyword. The positional information of each keyword is then obtained from the expanded text of each keyword, including:

[0075] Each keyword is expanded using a multimodal large language model to obtain the expanded text of each keyword. The CoT inference module is then used to analyze the expanded text of each keyword to obtain the location information of each keyword.

[0076] The Chain-of-Thought Reasoning Module (CoT) is an artificial intelligence technology based on step-by-step logical deduction. Its core lies in breaking down complex problems into multiple intermediate reasoning steps, generating a coherent reasoning chain by simulating the human thinking process, and finally drawing a conclusion.

[0077] S202, Input the text description information into the image generation model to obtain the base image, and fuse the base image and the final image of each keyword to obtain the fused image;

[0078] The process of inputting textual description information into an image generation model to obtain a base image, and then fusing the base image with the final image of each keyword to obtain a fused image, includes:

[0079] Input the text description information into the image generation model to obtain the base image. Input the expanded text of each keyword and the location information of each keyword into the image generation model to obtain the initial image of each keyword.

[0080] The initial image of each keyword is denoised using a diffusion implicit model to obtain the final image of each keyword. The base image and the final image of each keyword are then fused to obtain a fused image.

[0081] The reasons why diffusion implicit models are superior to traditional diffusion models are detailed below:

[0082] Reduce computational overhead: Traditional diffusion models rely on thousands of iterations to gradually denoise, which is computationally expensive; while the diffusion implicit model compresses the number of generation steps to tens or even a single step through non-Markov sampling paths or implicit generation mechanisms, which significantly reduces computational overhead and meets the requirements of real-time generation.

[0083] Scene adaptability: In the task of generating fused images in real time, the diffusion implicit model can generate candidate frames quickly and retain the fusion accuracy of multimodal features, avoiding the blurring or distortion caused by the speed compromise of the traditional diffusion model.

[0084] For example, a denoising operation is performed on the initial image of each keyword using a diffusion implicit model to obtain the final image of each keyword. The base image and the final image of each keyword are then fused to obtain a fused image, including:

[0085] The initial image for each keyword is denoised using a diffusion implicit model. The time step number decreases from the maximum value. When the time step decreases to zero, the final image for each keyword is obtained. The base image and the final image for each keyword are then fused to obtain a fused image.

[0086] Among them, image fusion can break the limitations of a single image, bringing together key information that was originally scattered in the base image and the final image of each keyword, thus enriching the visual content.

[0087] For example, the diffusion formula used in the diffusion implicit model is as follows:

[0088] ;

[0089] in, Indicates the first Within the nth time step, the diffusion implicit model for the nth time step The first keyword's initial image, after undergoing denoising, becomes the... The final image for each keyword;

[0090] Indicates the first Within the nth time step, the diffusion implicit model for the nth time step The first keyword's initial image, after undergoing denoising, becomes the... The final image for each keyword;

[0091] Indicates the time step number; Indicates the first The initial image for each keyword;

[0092] Represented as the first Word vectors of each keyword; Represented as a denoising network; These are model parameters;

[0093] Indicates the first The percentage of original information retained within each time step;

[0094] Then it means in the first The percentage of original information retained within each time step;

[0095] Indicates the first Within each time step, the denoising network according to , and The predicted noise value. In the diffusion implicit model, during inverse denoising, the time step number decreases from the maximum value. This means that starting from a completely noisy state, the model uses the learned patterns and rules to remove some noise with each time step decrease, gradually restoring a state that is closer to the original data. When the time step number decreases to zero, the final image of each keyword is obtained, realizing the generation of meaningful data from noise.

[0096] S203, Obtain the question text and answer text, and determine the image description information based on the feature vector of the fused image, the feature vector of the question, the feature vector of the answer, and the multimodal large language model;

[0097] The acquisition of question and answer text, based on the feature vectors of the fused image, the feature vectors of the question and the answer, and a multimodal large language model, determines image description information, including:

[0098] Extract question and answer texts from the visual question-and-answer text provided by the user; extract features from the question text to obtain the feature vector of the question; extract features from the answer text to obtain the feature vector of the answer text.

[0099] The feature vectors of the fused image, the question, and the answer are input into a multimodal large language model. The multimodal large language model processes the feature vectors of the fused image, the question, and the answer to obtain image description information.

[0100] For ease of explanation, the following example is provided:

[0101] For example, the visual question-and-answer text provided by the user is: What shape are a cat's glasses, and what color are the frames? The answer is: A cat's glasses are round, and the frames are gold.

[0102] Extract the question and answer texts from the visual question-and-answer text provided by the user.

[0103] The question text is: What shape are the cat's glasses, and what color are the frames?

[0104] The answer text is: The cat's glasses are round, and the frames are gold.

[0105] Image description information is obtained by processing the feature vectors of the fused image, the feature vector of the question, and the feature vector of the answer through a multimodal large language model.

[0106] The image description reads: "A picture of a cat wearing glasses. The glasses are round and have gold frames."

[0107] For ease of explanation, the following example is provided:

[0108] For example, the visual question-and-answer text is: What color is the flower in the picture, and what is the shape of its petals? The answer is: The flower in the picture is red, and the petals are oval.

[0109] Extract the question and answer texts from the visual question-and-answer text provided by the user.

[0110] Question text: What color is the flower in the picture, and what is the shape of its petals?

[0111] Answer text: The flower in the picture is red, and the petals are oval in shape.

[0112] Image description information is obtained by processing the feature vectors of the fused image, the feature vector of the question, and the feature vector of the answer through a multimodal large language model.

[0113] The image description is as follows: The image shows a blooming flower. The flower is a vibrant red overall, with oval-shaped petals that have smooth edges, are closely arranged, and spread out around the stamen. The background is green leaves or a natural environment, highlighting the flower as the visual center of attention.

[0114] S204, Select the feature vector of the representative word in the image description information as the representative token, and select the maximum value of the cosine similarity between each visual token and the representative token as the reward value;

[0115] The step of selecting the feature vector of the representative word in the image description information as the representative token, and selecting the maximum value of the cosine similarity between each visual token and the representative token as the reward value, includes:

[0116] The fused image is divided into multiple image blocks, and the feature vector of each image block is selected as each visual token. The feature vector of the representative word in the image description information is selected as the representative token.

[0117] Obtain the cosine similarity between each visual token and the representative token, and select the maximum value of the cosine similarity between each visual token and the representative token as the reward value.

[0118] Among them, representative words are words that can highly summarize and accurately reflect the core content, theme, or key features of an image.

[0119] For ease of explanation, the following example is provided:

[0120] For example, the image description is: "There is an adorable giant panda eating bamboo," with "giant panda" being the representative word.

[0121] For example, the image description might be: "An image showing a cozy living room with a sofa." "Sofa" is the descriptive word.

[0122] For example, the image description is: a red sports car driving on the road, with "sports car" being the representative word.

[0123] For example, the image description information could be: "An image featuring a field of purple lavender flowers. Purple and lavender purple can be used as keywords to emphasize the main color tone of the image."

[0124] For example, the image description information could be: "A colorful abstract painting, with bright yellow as the main color. Yellow and bright yellow can be used as representative words to highlight the color characteristics of the painting."

[0125] For example, the image description is: "An image of a golden ginkgo forest with leaves falling in the autumn wind." The word "golden yellow" can be used as a representative word to reflect the color atmosphere of the image.

[0126] Visual tokens carry rich feature information about the local or global aspects of an image, while representative tokens typically represent the typical representation of a word. By calculating the cosine similarity between visual tokens and representative tokens, we can quantify their directional consistency in the feature space. The higher the similarity, the better the features contained in the visual tokens match the semantics or patterns represented by the representative tokens.

[0127] In this model, the maximum cosine similarity between each visual token and its representative token is selected as the reward. The loss model, built based on this reward, plays a guiding role during training. When the cosine similarity between a visual token and its representative token falls short of expectations, the loss model generates a large loss value, prompting the model to adjust its parameters and optimize the feature extraction process. This allows the features of the visual token to gradually converge towards the feature pattern represented by the representative token. In this way, the diffusion implicit model can learn more discriminative features.

[0128] S205. Based on the loss model, the reward value and confidence level are added together to obtain the loss value. The fused image is modified based on the loss value and a preset method, and the modified fused image is selected as the target image that conforms to the text description information.

[0129] Among them, the full English name of the diffusion implicit model is: Denoising Diffusion Implicit Models (DDIMs). The diffusion implicit model is an improved diffusion model that significantly improves the sampling speed while maintaining the generation quality through implicit probability distribution and more efficient sampling strategies.

[0130] The step of adding the reward value and the confidence level according to the loss model to obtain the loss value includes:

[0131] The fused image is input into the image encoder, which processes the fused image to obtain the graph vector of the fused image. The image description information is input into the text encoder, which processes the image description information to obtain the text vector. The graph vector and the text vector are fused using a bidirectional cross-attention mechanism to obtain the fused feature.

[0132] The fused features are input into a multimodal classifier, which processes the fused features to obtain the confidence score. Based on the loss model, the reward value and the confidence score are added together to obtain the loss value.

[0133] The loss model is as follows:

[0134] ;

[0135] in, The loss value. , These are the first proportional parameter and the second proportional parameter, respectively. As a reward value, , where is the confidence level.

[0136] Among them, the multimodal classifier adopts an image-text matching module. The image-text matching module maps the image and text to a shared semantic space through dual encoders, and then uses similarity calculation to directly quantify the matching degree of image-text pairs, generating a confidence value between 0 and 1.

[0137] The English name of the image and text matching module is: Image Text Matching Module, and its abbreviation is: ITM module.

[0138] In modifying the fused image to approximate the target, the design mechanism of higher reward scores and lower confidence scores with lower loss values ​​demonstrates significant advantages. The reward score directly measures the degree of fit between the modified fused image and the textual description. A high reward score means that the modified fused image is closer to the textual description in terms of content, style, and detail. Linking this score to the loss value guides the diffusion implicit model to continuously optimize towards higher rewards during the modification process, ensuring that the modified fused image remains consistent with the textual description at a macroscopic level. Confidence reflects the model's trust in the modified fused image; high confidence indicates that the diffusion implicit model has a high degree of confidence in the generation quality and feature accuracy of the modified fused image. Under this dual constraint, the diffusion implicit model must accurately match the target in content while maintaining high confidence during the modification process, avoiding local optima or bias that may result from optimizing a single metric. This allows the diffusion implicit model to more comprehensively and accurately approximate the textual description through modifications to the fused image.

[0139] The beneficial effects of this application embodiment are twofold. Firstly, based on the loss model, the reward value and confidence level are added together to obtain the loss value. The fused image is then modified based on the loss value and a preset method. The modified fused image is selected as the target image that conforms to the text description information. Since no manual processing is required, the generation time of the target image that conforms to the text description information is reduced, which is beneficial to improving the generation efficiency of the target image that conforms to the text description information. Secondly, since it is not affected by human intervention, it is beneficial to improve the reliability of the target image that conforms to the text description information.

[0140] Please see Figure 3 , Figure 3The implementation flowchart of S205 provided in the embodiments of this application is described in detail below:

[0141] S301, According to the loss model, the reward value and the confidence level are added together to obtain the loss value;

[0142] The step of adding the reward value and the confidence level according to the loss model to obtain the loss value includes:

[0143] The fused image is input into the image encoder, which processes the fused image to obtain the graph vector of the fused image. The image description information is input into the text encoder, which processes the image description information to obtain the text vector. The graph vector and the text vector are fused using a bidirectional cross-attention mechanism to obtain the fused feature.

[0144] The fused features are input into a multimodal classifier, which processes the fused features to obtain the confidence score. Based on the loss model, the reward value and the confidence score are added together to obtain the loss value.

[0145] S302, update the model parameters of the diffusion implicit model to obtain the updated model parameters, control the diffusion implicit model to modify the fused image using the updated model parameters until the loss value is less than the preset value, then save the modified fused image, and select the modified fused image as the target image that conforms to the text description information.

[0146] S302 includes:

[0147] The model parameters of the diffusion implicit model are updated using the gradient iteration formula to obtain the updated model parameters. The diffusion implicit model is then controlled to modify the fused image using the updated model parameters until the loss value is less than the preset value. Only then is the modified fused image saved and selected as the target image that matches the text description information.

[0148] The gradient iteration formula is as follows:

[0149] ;

[0150] in, Let L be the total loss with respect to the model parameters. The gradient is used for backpropagation to update model parameters, making the target image more consistent with multimodal context constraints;

[0151] X represents the target image. Let R be the gradient of the reward function R with respect to the target image;

[0152] The gradient of the target image with respect to the initial latent variables. Calculated by the VAE decoder, Reflecting latent variables How it affects the target image;

[0153] Multiplication terms The gradient of the parameters for the denoising process from arrive The product of consecutive products;

[0154] Indicates the time step number. This represents the total number of time steps. These are partial derivatives; For graph vectors; It is a text vector; Represented as a denoising network; These are model parameters;

[0155] Indicates the first Within each time step, the denoising network according to , and The predicted noise value.

[0156] The Chinese name for VAE decoder is: Variational Autoencoder Decoder;

[0157] The English name for VAE decoder is: Decoder of Variational Autoencoder.

[0158] In this embodiment, when the loss value is less than the preset value, it means that the modified fused image has approximated the text description information, and the modified fused image does not need to be modified. Therefore, the modified fused image is saved and selected as the target image that conforms to the text description information. Since no manual processing is required, the generation time of the target image that conforms to the text description information is reduced, which is beneficial to improving the generation efficiency of the target image that conforms to the text description information.

[0159] For the image generation method described in the above embodiments, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic block diagram of the image generation apparatus provided in the embodiments of this application. Figure 4 The image generation apparatus 400 shown can be applied to, for example... Figure 1 The application scenario diagram shows electronic devices. The following section uses electronic devices as an example to illustrate this. Figure 4 The image generation device 400 shown will be described in detail. The image generation device 400 may include an acquisition module 401, a fusion module 402, a determination module 403, a selection module 404, and a generation module 405.

[0160] The acquisition module 401 is used to acquire text description information and multiple keywords in the text description information, and determine the position information of each keyword through a predefined method;

[0161] The fusion module 402 is used to input textual description information into the image generation model to obtain a base image, and then fuse the base image with the final image of each keyword to obtain a fused image.

[0162] The determination module 403 is used to obtain the question text and answer text, and determine the image description information based on the feature vector of the fused image, the feature vector of the question, the feature vector of the answer, and the multimodal large language model;

[0163] The selection module 404 is used to select the feature vector of the representative word in the image description information as the representative token, and select the maximum value of the cosine similarity between each visual token and the representative token as the reward value.

[0164] The generation module 405 is used to add the reward value and confidence level according to the loss model to obtain the loss value, modify the fused image based on the loss value and a preset method, and select the modified fused image as the target image that conforms to the text description information.

[0165] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0166] The beneficial effects of this application embodiment are twofold. Firstly, based on the loss model, the reward value and confidence level are added together to obtain the loss value. The fused image is then modified based on the loss value and a preset method. The modified fused image is selected as the target image that conforms to the text description information. Since no manual processing is required, the generation time of the target image that conforms to the text description information is reduced, which is beneficial to improving the generation efficiency of the target image that conforms to the text description information. Secondly, since it is not affected by human intervention, it is beneficial to improve the reliability of the target image that conforms to the text description information.

[0167] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0168] like Figure 5 As shown, Figure 5 The electronic device 2 includes: at least one processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20, wherein the processor 20 executes the computer program 22 to implement the steps in any of the above method embodiments.

[0169] The electronic device 2 may include, but is not limited to, a processor 20 and a memory 21. Those skilled in the art will understand that... Figure 5 This is merely an example of electronic device 2 and does not constitute a limitation on electronic device 2. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0170] The processor 20 is used to run a computer program 22 stored in the memory 21, and performs the following steps when executing the computer program 22:

[0171] Retrieve text description information and multiple keywords from the text description information, and determine the location information of each keyword through a predefined method;

[0172] The textual description information is input into the image generation model to obtain the base image. The base image and the final image of each keyword are then fused to obtain the fused image.

[0173] The question and answer texts are obtained, and the image description information is determined based on the feature vectors of the fused image, the feature vectors of the question and the answer, and a multimodal large language model.

[0174] The feature vectors of representative words in the image description information are selected as representative tokens, and the maximum value of the cosine similarity between each visual token and the representative token is selected as the reward value.

[0175] According to the loss model, the reward value and confidence level are added together to obtain the loss value. Based on the loss value and the preset method, the fused image is modified, and the modified fused image is selected as the target image that conforms to the text description information.

[0176] In some embodiments, the processor 20 is configured to implement:

[0177] Retrieve text description information and multiple keywords from the text description information;

[0178] Each keyword is expanded using a multimodal large language model to obtain the expanded text of each keyword, and the position information of each keyword is obtained from the expanded text of each keyword.

[0179] In some embodiments, the processor 20 is configured to implement:

[0180] Input the text description information into the image generation model to obtain the base image. Input the expanded text of each keyword and the location information of each keyword into the image generation model to obtain the initial image of each keyword.

[0181] The initial image of each keyword is denoised using a diffusion implicit model to obtain the final image of each keyword. The base image and the final image of each keyword are then fused to obtain a fused image.

[0182] In some embodiments, the processor 20 is configured to implement:

[0183] Extract question and answer texts from the visual question-and-answer text provided by the user; extract features from the question text to obtain the feature vector of the question; extract features from the answer text to obtain the feature vector of the answer text.

[0184] The feature vectors of the fused image, the question, and the answer are input into a multimodal large language model. The multimodal large language model processes the feature vectors of the fused image, the question, and the answer to obtain image description information.

[0185] In some embodiments, the processor 20 is configured to implement:

[0186] The fused image is divided into multiple image blocks, and the feature vector of each image block is selected as each visual token. The feature vector of the representative word in the image description information is selected as the representative token.

[0187] Obtain the cosine similarity between each visual token and the representative token, and select the maximum value of the cosine similarity between each visual token and the representative token as the reward value.

[0188] In some embodiments, the processor 20 is configured to implement:

[0189] According to the loss model, the reward value and the confidence level are added together to obtain the loss value;

[0190] The model parameters of the diffusion implicit model are updated to obtain the updated model parameters. The diffusion implicit model is then controlled to modify the fused image using the updated model parameters until the loss value is less than the preset value. Only then is the modified fused image saved, and the modified fused image is selected as the target image that conforms to the text description information.

[0191] In some embodiments, the processor 20 is configured to implement:

[0192] The fused image is input into the image encoder, which processes the fused image to obtain the graph vector of the fused image. The image description information is input into the text encoder, which processes the image description information to obtain the text vector. The graph vector and the text vector are fused using a bidirectional cross-attention mechanism to obtain the fused feature.

[0193] The fused features are input into a multimodal classifier, which processes the fused features to obtain the confidence score. Based on the loss model, the reward value and the confidence score are added together to obtain the loss value.

[0194] The processor 20 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0195] In some embodiments, the memory 21 may be an internal storage unit of the electronic device 2, such as a hard disk or memory of the electronic device 2. In other embodiments, the memory 21 may be an external storage device of the electronic device 2, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 2. Furthermore, the memory 21 may include both internal and external storage units of the electronic device 2. The memory 21 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 21 can also be used to temporarily store data that has been output or will be output.

[0196] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0197] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0198] The computer-readable storage medium stores program code that can be called by a processor to execute the image generation method described in the above method embodiments.

[0199] Computer-readable storage media have storage space for program code.

[0200] The program code includes the code for any step of the image generation method described in the above method embodiments.

[0201] For example, when program code is invoked by the processor, it can perform the following steps:

[0202] Retrieve text description information and multiple keywords from the text description information, and determine the location information of each keyword through a predefined method;

[0203] The textual description information is input into the image generation model to obtain the base image. The base image and the final image of each keyword are then fused to obtain the fused image.

[0204] The question and answer texts are obtained, and the image description information is determined based on the feature vectors of the fused image, the feature vectors of the question and the answer, and a multimodal large language model.

[0205] The feature vectors of representative words in the image description information are selected as representative tokens, and the maximum value of the cosine similarity between each visual token and the representative token is selected as the reward value.

[0206] According to the loss model, the reward value and confidence level are added together to obtain the loss value. Based on the loss value and the preset method, the fused image is modified, and the modified fused image is selected as the target image that conforms to the text description information.

[0207] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0208] The computer-readable storage medium may also be an external storage device of the image generating apparatus or electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, or non-transitory computer-readable storage medium equipped on the image generating apparatus or electronic device.

[0209] Since the computer program stored in the computer-readable storage medium can execute any of the image generation methods based on the diffusion implicit model provided in the embodiments of this application, the computer-readable storage medium can achieve the beneficial effects that any of the image generation methods based on the diffusion implicit model provided in the embodiments of this application can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0210] This application provides a computer program product that, when run on an electronic device, causes the electronic device to execute the image generation method described above.

[0211] When a computer program is loaded into an electronic device, it can perform the following steps:

[0212] Retrieve text description information and multiple keywords from the text description information, and determine the location information of each keyword through a predefined method;

[0213] The textual description information is input into the image generation model to obtain the base image. The base image and the final image of each keyword are then fused to obtain the fused image.

[0214] The question and answer texts are obtained, and the image description information is determined based on the feature vectors of the fused image, the feature vectors of the question and the answer, and a multimodal large language model.

[0215] The feature vectors of representative words in the image description information are selected as representative tokens, and the maximum value of the cosine similarity between each visual token and the representative token is selected as the reward value.

[0216] According to the loss model, the reward value and confidence level are added together to obtain the loss value. Based on the loss value and the preset method, the fused image is modified, and the modified fused image is selected as the target image that conforms to the text description information.

[0217] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0218] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0219] Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0220] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0221] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. An image generation method based on a diffusion implicit model, characterized in that, The image generation method, applied to electronic devices, includes: Obtain text description information and multiple keywords from the text description information. Expand each keyword using a multimodal large language model to obtain the expanded text of each keyword. Extract the position information of each keyword from the expanded text of each keyword. The textual description information is input into the image generation model to obtain the base image. The base image and the final image of each keyword are then fused to obtain the fused image. The question and answer texts are obtained, and the image description information is determined based on the feature vectors of the fused image, the feature vectors of the question and the answer, and a multimodal large language model. The feature vectors of representative words in the image description information are selected as representative tokens, and the maximum value of the cosine similarity between each visual token and the representative token is selected as the reward value. Based on the loss model, the reward value and confidence level are added together to obtain the loss value. The model parameters of the diffusion implicit model are then updated to obtain the updated model parameters. The diffusion implicit model is controlled to use the updated model parameters to modify the fused image until the loss value is less than the preset value. Only then is the modified fused image saved, and the modified fused image is selected as the target image that conforms to the text description information.

2. The image generation method according to claim 1, characterized in that, The process of inputting textual description information into an image generation model to obtain a base image, and then fusing the base image with the final image of each keyword to obtain a fused image, includes: Input the text description information into the image generation model to obtain the base image. Input the expanded text of each keyword and the location information of each keyword into the image generation model to obtain the initial image of each keyword. The initial image of each keyword is denoised using a diffusion implicit model to obtain the final image of each keyword. The base image and the final image of each keyword are then fused to obtain a fused image.

3. The image generation method according to claim 1, characterized in that, The acquisition of question and answer text, based on the feature vectors of the fused image, the feature vectors of the question and the answer, and a multimodal large language model, determines image description information, including: Extract question and answer texts from the visual question-and-answer text provided by the user; extract features from the question text to obtain the feature vector of the question; extract features from the answer text to obtain the feature vector of the answer text. The feature vectors of the fused image, the question, and the answer are input into a multimodal large language model. The multimodal large language model processes the feature vectors of the fused image, the question, and the answer to obtain image description information.

4. The image generation method according to claim 1, characterized in that, The step of selecting the feature vector of the representative word in the image description information as the representative token, and selecting the maximum value of the cosine similarity between each visual token and the representative token as the reward value, includes: The fused image is divided into multiple image blocks, and the feature vector of each image block is selected as each visual token. The feature vector of the representative word in the image description information is selected as the representative token. Obtain the cosine similarity between each visual token and the representative token, and select the maximum value of the cosine similarity between each visual token and the representative token as the reward value.

5. The image generation method according to any one of claims 1 to 4, characterized in that, The loss value is obtained by adding the reward value and the confidence level according to the loss model, including: The fused image is input into the image encoder, which processes the fused image to obtain the graph vector of the fused image. The image description information is input into the text encoder, which processes the image description information to obtain the text vector. The graph vector and the text vector are fused using a bidirectional cross-attention mechanism to obtain the fused feature. The fused features are input into a multimodal classifier, which processes the fused features to obtain the confidence score. Based on the loss model, the reward value and the confidence score are added together to obtain the loss value.

6. An image generation device based on a diffusion implicit model, characterized in that, Applied to electronic devices, including: The acquisition module is used to acquire text description information and multiple keywords in the text description information. Each keyword is expanded using a multimodal large language model to obtain the expanded text of each keyword. The position information of each keyword is obtained from the expanded text of each keyword. The fusion module is used to input textual description information into the image generation model to obtain a base image, and then fuse the base image with the final image of each keyword to obtain a fused image. The determination module is used to obtain the question text and answer text, and determine the image description information based on the feature vectors of the fused image, the feature vectors of the question and the answer, and the multimodal large language model; The selection module is used to select the feature vector of the representative word in the image description information as the representative token, and select the maximum value of the cosine similarity between each visual token and the representative token as the reward value. The generation module is used to add the reward value and confidence level according to the loss model to obtain the loss value, update the model parameters of the diffusion implicit model to obtain the updated model parameters, control the diffusion implicit model to modify the fused image with the updated model parameters until the loss value is less than the preset value, and then save the modified fused image. The modified fused image is selected as the target image that matches the text description information.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the image generation method as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the image generation method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image processing method and device

    CN117456026A

  • Image generation method and device, electronic equipment and readable storage medium

    CN118736038A