Image generation method, automatic question answering method, and parameter generation model training method

By combining the method of large-scale language model, semantic disassembly and control the image generation process, the problem of opacity in the generation process of hidden space diffusion model and the large difference between the input text is solved, and the image generation is achieved with higher interpretability and controllability.

WO2025112948A1PCT designated stage expired Publication Date: 2025-06-05ALIBABA (CHINA) CO LTD

Patent Information

Application Number
PCT/CN2024/125023
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-10-15
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The existing hidden space diffusion model is difficult to understand during the image generation process, and the generated image is very different from the user input text, and cannot perform semantic control of visual elements.

Method used

Using an image generation method combined with a large-scale language model, the visual elements of the image are semantically disassembled through the parameter generation model, image generation parameters are obtained, and inputted to the image generation model for precise generation.

Benefits of technology

The interpretability and controllability of image generation are improved, so that the generated target image can clearly express image description text and image generation parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024125023_05062025_PF_FP_ABST
    Figure CN2024125023_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an image generation method, an automatic question answering method, and a parameter generation model training method. The image generation method comprises: acquiring an image description text; inputting the image description text and generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used for describing visual features of an image, and the parameter generation model is obtained by performing training on the basis of a plurality of sample image and text pairs and sample parameter information carried by the plurality of sample image and text pairs; and inputting the image generation parameters and the image description text into an image generation model to obtain a target image corresponding to the image description text. By using a parameter generation model to perform semantic decomposition on visual elements of an image to obtain image generation parameters, and further completing accurate image generation on the basis of the image generation parameters, a target image can clearly express an image description text and the image generation parameters, thereby improving the interpretability and controllability of image generation.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation, automatic question answering, and parameter generation model training methods

[0001] This disclosure claims priority to Chinese patent application number 2023116226401, filed with the Patent Office of China on November 28, 2023, and entitled “Image Generation, Automatic Question Answering, and Parameter Generation Model Training Method,” the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0002] The embodiments of the present disclosure relate to the field of computer technology, and in particular to image generation, automatic question answering, and parameter generation model training methods. Background Art

[0003] With the development of computer technology, text-based image processing has gradually become a core technology in the field of artificial intelligence generated content (AIGC). This technology can generate images from text descriptions and transform and adjust them according to user requirements and input content, making it easier for users to create unique works of art. It has been widely used in the field of digital art.

[0004] Currently, the most common text-generated graph architecture is the Hidden Space Diffusion Model (HSDM). However, because the HSDM operates in the latent space, the generation process is difficult to understand and the generated image is significantly different from the text originally input by the user.

[0005] Summary of the Invention

[0006] In view of this, embodiments of the present disclosure provide an image generation method. One or more embodiments of the present disclosure also include an automatic question-answering method, a parameter generation model training method, an image generation device, an automatic question-answering device, a parameter generation model training device, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.

[0007] According to a first aspect of an embodiment of the present disclosure, there is provided an image generation method, comprising:

[0008] Get image description text;

[0009] Input the image description text and generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image. The parameter generation model is trained based on multiple sample image-text pairs and the sample parameter information carried by the multiple sample image-text pairs;

[0010] The image generation parameters and image description text are input into the image generation model to obtain the target image corresponding to the image description text.

[0011] According to a second aspect of an embodiment of the present disclosure, a parameter generation model training method is provided, which is applied to a cloud-side device, including:

[0012] Acquire a sample set, wherein the sample set includes a plurality of sample image-text pairs, the sample image-text pairs include a sample image and a sample description text, and the sample image-text pairs carry sample parameter information;

[0013] Inputting multiple sample image-text pairs and prediction prompt information into the pre-trained language model to obtain image prediction parameters corresponding to the multiple sample image-text pairs;

[0014] According to the image prediction parameters and sample parameter information, the model parameters of the pre-trained language model are adjusted to obtain a parameter generation model that has completed training.

[0015] According to a third aspect of an embodiment of the present disclosure, there is provided an automatic question-answering method, comprising:

[0016] receiving an image question answering request, wherein the image question answering request carries image description text;

[0017] Input the image description text and generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image. The parameter generation model is trained based on multiple sample image-text pairs and the sample parameter information carried by the multiple sample image-text pairs;

[0018] The image generation parameters and image description text are input into the image generation model to obtain the answer image corresponding to the image question answering request.

[0019] According to a fourth aspect of the embodiments of the present disclosure, there is provided an image generating apparatus, including:

[0020] A first acquisition module is configured to acquire image description text;

[0021] A first input module is configured to input the image description text and generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and sample parameter information carried by the multiple sample image-text pairs;

[0022] The second input module is configured to input the image generation parameters and the image description text into the image generation model to obtain the target image corresponding to the image description text.

[0023] According to a fifth aspect of an embodiment of the present disclosure, a parameter generation model training apparatus is provided, which is applied to a cloud-side device and includes:

[0024] A second acquisition module is configured to acquire a sample set, wherein the sample set includes a plurality of sample image-text pairs, the sample image-text pairs include a sample image and a sample description text, and the sample image-text pairs carry sample parameter information;

[0025] A third input module is configured to input a plurality of sample image-text pairs and prediction prompt information into the pre-trained language model to obtain image prediction parameters corresponding to the plurality of sample image-text pairs;

[0026] The adjustment module is configured to adjust the model parameters of the pre-trained language model according to the image prediction parameters and sample parameter information to obtain a parameter generation model that has completed training.

[0027] According to a sixth aspect of an embodiment of the present disclosure, there is provided an automatic question-answering device, comprising:

[0028] A first receiving module is configured to receive an image question and answer request, wherein the image question and answer request carries an image description text;

[0029] a fourth input module configured to input the image description text and the generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and sample parameter information carried by the multiple sample image-text pairs;

[0030] The fifth input module is configured to input image generation parameters and image description text into the image generation model to obtain a reply image corresponding to the image question and answer request.

[0031] According to a seventh aspect of an embodiment of the present disclosure, there is provided a computing device, including:

[0032] memory and processor;

[0033] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.

[0034] According to an eighth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the method provided in the first aspect, the second aspect or the third aspect are implemented.

[0035] According to a ninth aspect of an embodiment of the present disclosure, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the method provided in the first aspect, the second aspect, or the third aspect above.

[0036] An embodiment of the present disclosure provides an image generation method that obtains image description text; inputs the image description text and generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and sample parameter information carried by the multiple sample image-text pairs; the image generation parameters and the image description text are input into the image generation model to obtain a target image corresponding to the image description text. By using the parameter generation model to semantically decompose the visual elements of the image to obtain the image generation parameters, accurate image generation is further performed based on the image generation parameters, so that the target image can clearly express the image description text and the image generation parameters, thereby improving the interpretability and controllability of image generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] FIG1 is an architecture diagram of an image generation system provided by one embodiment of the present disclosure;

[0038] FIG2 is an architecture diagram of another image generation system provided by an embodiment of the present disclosure;

[0039] FIG3 is a flow chart of an image generation method provided by one embodiment of the present disclosure;

[0040] FIG4 is a flowchart of a processing process of an image generation method provided by one embodiment of the present disclosure;

[0041] FIG5 is a flow chart of a parameter generation model training method provided by one embodiment of the present disclosure;

[0042] FIG6 is a flowchart of another parameter generation model training method provided by an embodiment of the present disclosure;

[0043] FIG7 is a flow chart of an automatic question-answering method provided by one embodiment of the present disclosure;

[0044] FIG8 is a flowchart of a processing process of another image generation method provided by an embodiment of the present disclosure;

[0045] FIG9 is a schematic diagram of a processing process of a parameter generation model provided by an embodiment of the present disclosure;

[0046] FIG10 is a schematic diagram of an image generation interface provided by an embodiment of the present disclosure;

[0047] FIG11 is a schematic structural diagram of an image generating device provided by an embodiment of the present disclosure;

[0048] FIG12 is a schematic structural diagram of a parameter generation model training device provided by one embodiment of the present disclosure;

[0049] FIG13 is a schematic structural diagram of an automatic question-answering device provided by one embodiment of the present disclosure;

[0050] FIG14 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0051] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.

[0052] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.

[0053] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0054] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0055] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, which typically contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. The large model is pre-trained using large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large-scale language model (LLM), a multi-modal pre-training model, etc.

[0056] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image description (IC), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0057] First, the terms involved in one or more embodiments of the present disclosure are explained.

[0058] Latent Space Diffusion Model: Latent Space Diffusion Model refers to an artificial intelligence technology widely used in image editing and image generation, which is used to encode images into latent space for denoising and denoising.

[0059] Attention mechanism: Attention mechanism is a machine learning method that can assign different weights according to the importance of each part of the input data.

[0060] CLIP: CLIP (Contrastive Language-Image Pre-training) is a model for measuring the relevance of text and images, and is widely used for text and image representation extraction.

[0061] Large-scale language models: Large-scale language models are trained on massive amounts of text data to predict and generate various language expressions, such as text, sentences, and paragraphs.

[0062] LoRA: Low-Rank adaptation is a large language model fine-tuning method that only requires training a small number of parameters when using large models to adapt to downstream tasks to achieve good results.

[0063] Text-based graph technology outputs images based on user-entered image descriptions, offering immense value in a wide range of scenarios, including advertising recommendations, interactive entertainment, and art design. A common text-based graph architecture is the latent space diffusion model. Because latent space diffusion models operate in a hidden space, their internal working mechanisms and decision-making processes are difficult to understand and interpret, making the generation process difficult to comprehend. Furthermore, these methods rely on a feature extraction module (CLIP encoder) to understand text and lack semantic control over visual elements during the generation phase. This results in significant discrepancies between the generated results and the user's initial input image description.

[0064] To improve the accuracy of image generation models, the disclosed embodiments propose an image generation method that combines a large-scale language model. This method leverages the powerful semantic understanding and associative capabilities of large-scale language models to semantically decompose and reconstruct the visual elements of an image. These visual elements are parameterized and then input into the image generation model to achieve precise image generation. This ensures greater interpretability and controllability of the image generation process. Furthermore, large-scale language models can align complex and long texts, paving the way for more user interaction methods. Furthermore, by combining large-scale language models with image generation models, extended applications such as convenient and refined image editing and generation of images with similar structures can be achieved.

[0065] Specifically, an embodiment of the present disclosure proposes an image generation scheme to obtain image description text; input the image description text and generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and sample parameter information carried by multiple sample image-text pairs; input the image generation parameters and image description text into the image generation model to obtain a target image corresponding to the image description text.

[0066] In the present disclosure, an image generation method is provided. The present disclosure also relates to an automatic question-answering method, a parameter generation model training method, an image generation device, an automatic question-answering device, a parameter generation model training device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0067] Referring to FIG1 , FIG1 shows an architecture diagram of an image generation system provided by an embodiment of the present disclosure. The image generation system may include a client 100 and a server 200;

[0068] The client 100 is used to send image description text to the server 200;

[0069] The server 200 is configured to input the image description text and generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and sample parameter information carried by the multiple sample image-text pairs; input the image generation parameters and the image description text into the image generation model to obtain a target image corresponding to the image description text; and send the target image to the client 100;

[0070] The client 100 is further configured to receive the target image sent by the server 200 .

[0071] By applying the solution of the embodiment of the present disclosure, the visual elements of the image are semantically decomposed by using a parameter generation model to obtain image generation parameters, and then accurate image generation is completed based on the image generation parameters, so that the target image can clearly express the image description text and image generation parameters, thereby improving the interpretability and controllability of image generation.

[0072] Referring to Figure 2, Figure 2 shows an architecture diagram of another image generation system provided by one embodiment of the present disclosure. The image generation system may include multiple clients 100 and a server 200. The clients 100 may include end-side devices, and the server 200 may include cloud-side devices. Multiple clients 100 may establish communication connections through the server 200. In an image generation scenario, the server 200 is used to provide image generation services between the multiple clients 100. The multiple clients 100 may act as senders or receivers, respectively, and communicate through the server 200.

[0073] Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the image generation scenario, the user can publish a data stream to the server 200 through the client 100, and the server 200 can generate a target image based on the data stream and push the target image to other clients with which communication has been established.

[0074] The client 100 and the server 200 are connected via a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.

[0075] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application. The client 100 can be based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device to run or certain APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0076] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that support backend training for models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with a blockchain. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0077] It is worth noting that the image generation method provided in the embodiments of the present disclosure is generally executed by the server. However, in other embodiments of the present disclosure, the client may also have similar functions to the server to execute the image generation method provided in the embodiments of the present disclosure. In other embodiments, the image generation method provided in the embodiments of the present disclosure may also be executed jointly by the client and the server.

[0078] 3 , which shows a flow chart of an image generation method provided by an embodiment of the present disclosure, specifically comprising the following steps:

[0079] Step 302: Obtain image description text.

[0080] In one or more embodiments of the present disclosure, when image generation begins, image description text may be obtained, and a target image that meets the actual needs of the user may be generated based on the image description text.

[0081] Specifically, image descriptions represent the user's image generation needs. Image descriptions can be in different languages, such as English or Chinese. Because large-scale language models are incorporated into the image generation process, image descriptions can be short or long. For example, an image description might be "Two black cats and a white dog on an orange sofa."

[0082] In practical applications, there are multiple ways to obtain image description text, and the specific method to be used depends on the actual situation. The embodiments of this disclosure do not impose any restrictions on this. In one possible implementation of this disclosure, image description text sent by a user through a client can be received. In another possible implementation of this disclosure, image description text can be read from other data acquisition devices or databases.

[0083] Step 304: Input the image description text and generation prompt information into the parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and the sample parameter information carried by the multiple sample image-text pairs.

[0084] In one or more embodiments of the present disclosure, after obtaining the image description text, the image description text and generation prompt information may be input into a parameter generation model to obtain image generation parameters corresponding to the image description text.

[0085] Specifically, the parameter generation model is obtained by training a pre-trained language model based on a sample set. The pre-trained language model can be a large-scale language model, such as a large multimodal model, or a language processing model trained using the first training set. The image generation parameters can be understood as parameterized image generation conditions. Image generation parameters include but are not limited to size parameters, position parameters, shape parameters, and the like. The generation prompt information can be understood as a parameter generation paradigm, also known as a generation condition paradigm. The generation prompt information is used to guide the parameter generation model to generate image generation parameters. The generation prompt information includes but is not limited to image generation condition elements such as shape, color, size, position, blurring, and key points. For example, the generation prompt information can be "what are the elements in the image".

[0086] It should be noted that in the embodiment of the present disclosure, the generation conditions are unified into parameters that can be described in natural language, rather than image conditions, so as to facilitate semantic understanding of large-scale language models.

[0087] Step 306: Input the image generation parameters and the image description text into the image generation model to obtain a target image corresponding to the image description text.

[0088] In one or more embodiments of the present disclosure, image description text is obtained, the image description text and generation prompt information are input into a parameter generation model, and after the image generation parameters corresponding to the image description text are obtained, the image generation parameters and the image description text can be further input into the image generation model to obtain a target image corresponding to the image description text.

[0089] Specifically, the image generation model can be a latent space diffusion model or an image generation model trained using the second training set. The image generation model is used to generate a final target image based on the image generation parameters and the image description text. The target image can be a black and white image or a color image (RGB image).

[0090] By applying the solution of the embodiment of the present disclosure, the visual elements of the image are semantically decomposed by using a parameter generation model to obtain image generation parameters, and then accurate image generation is completed based on the image generation parameters, so that the target image can clearly express the image description text and image generation parameters, thereby improving the interpretability and controllability of image generation.

[0091] In an optional embodiment of the present disclosure, image generation parameters and image description text are input into an image generation model. A side branch structure may be used in the image generation model to process the image generation parameters. That is, a parameter encoding unit is additionally added as a side branch structure on the basis of the encoding and decoding unit originally provided by the image generation model. That is, the image generation model includes a parameter encoding unit and an encoding and decoding unit. The above-mentioned inputting the image generation parameters and image description text into the image generation model to obtain a target image corresponding to the image description text may include the following steps:

[0092] The image generation parameters and the image description text are input into the image generation model, and the image generation parameters are encoded by the parameter encoding unit to obtain parameter encoding features;

[0093] The encoding and decoding unit generates a target image corresponding to the image description text based on the parameter encoding features and the image description text.

[0094] Refer to Figure 4, which shows a processing flow chart of an image generation method provided by an embodiment of the present disclosure. As shown in Figure 4, after the parameter generation model generates the image generation parameters, the image description text and the image generation parameters can be input into the image generation model together. In the image generation model, the image generation parameters are first encoded by the parameter encoding unit to obtain parameter encoding features. Secondly, the encoding and decoding unit generates a target image corresponding to the image description text based on the parameter encoding features and the image description text.

[0095] It should be noted that the encoding and decoding unit includes an encoding unit and a decoding unit. After obtaining the parameter encoding features, the parameter encoding features can be input into the encoding unit in the latent space for parameterized conditional control, and the residual is sent to the decoding unit to achieve the final parameter fusion generation.

[0096] In practical applications, the parameter encoding unit can directly encode the image generation parameters to obtain parameter encoding features. Furthermore, since the image generation parameters include parameters of different dimensions, the parameter encoding unit can also encode the image generation parameters into parameter encoding features of different latitudes to address the diversity of image generation parameter dimensions.

[0097] Using the solution of the disclosed embodiments, image generation parameters and image description text are input into the image generation model. The parameter encoding unit encodes the image generation parameters to obtain parameter encoding features. The encoding and decoding unit then generates a target image corresponding to the image description text based on the parameter encoding features and the image description text. By integrating the image generation parameters into the target image generation process for control, the accuracy of the target image is guaranteed.

[0098] In an optional embodiment of the present disclosure, the parameter encoding unit includes a one-dimensional parameter encoding unit, a two-dimensional parameter encoding unit, and a feature aggregation unit; the parameter encoding unit generates parameter encoding for the image to obtain the parameter encoding feature, which may include the following steps:

[0099] The one-dimensional parameter encoding unit encodes the one-dimensional parameter in the image generation parameter to obtain a one-dimensional encoding feature;

[0100] The two-dimensional parameter encoding unit encodes the two-dimensional parameters in the image generation parameters to obtain a two-dimensional encoding feature;

[0101] The one-dimensional coding features and the two-dimensional coding features are aggregated by the feature aggregation unit to obtain parameter coding features.

[0102] Specifically, image generation parameters may include parameters of different dimensions, such as one-dimensional parameters and two-dimensional parameters. One-dimensional parameters include, but are not limited to, color parameters and blur parameters. Two-dimensional parameters include, but are not limited to, position parameters and shape parameters. Therefore, to better encode image generation parameters, the disclosed embodiment divides the parameter encoding unit into a one-dimensional parameter encoding unit for encoding one-dimensional parameters and a two-dimensional parameter encoding unit for encoding two-dimensional parameters.

[0103] Furthermore, since the image generation parameters represent the overall generation conditions of the target image, a feature aggregation unit can be additionally added to the parameter coding unit, so that the one-dimensional coding features and the two-dimensional coding features are aggregated through the feature aggregation unit to obtain parameter coding features.

[0104] By applying the solution of the embodiment of the present disclosure, the one-dimensional parameters in the image generation parameters are encoded by the one-dimensional parameter encoding unit to obtain one-dimensional encoding features; the two-dimensional parameters in the image generation parameters are encoded by the two-dimensional parameter encoding unit to obtain two-dimensional encoding features; and the one-dimensional encoding features and the two-dimensional encoding features are aggregated by the feature aggregation unit to obtain parameter encoding features, thereby encoding parameters of different dimensions respectively, thereby ensuring the accuracy of the parameter encoding features.

[0105] In an optional embodiment of the present disclosure, after inputting image generation parameters and image description text into an image generation model to obtain a target image corresponding to the image description text, an image update parameter for the target image may be further received, and the target image may be edited based on the image update parameter, or an image having a structure similar to the target image may be generated. That is, after inputting image generation parameters and image description text into the image generation model to obtain a target image corresponding to the image description text, the following steps may also be included:

[0106] Obtain image update parameters for the target image;

[0107] determining an image editing region in a target image according to an image update parameter;

[0108] Mask the image editing area to obtain a mask generation sequence;

[0109] The mask generation sequence and the target image are input into the image generation model to obtain the updated target image.

[0110] Specifically, the image update parameter is used to modify or adjust the target image to obtain an updated target image. The image update parameter can be an image generation condition independent of the image generation parameter, such as if the image generation parameter is "the number of black cats is 2", and the image update parameter is "the number of black cats is 3"; for example, if the image generation parameter is "the number of black cats is 2", the image update parameter is "the number of black cats plus one". There are many ways to obtain the image update parameters for the target image, and the specific selection is based on the actual situation. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the image update parameters sent by the user through the client can be received. In another possible implementation of the present disclosure, the image update parameters for the target image can be read from other data acquisition devices or databases.

[0111] It should be noted that after obtaining the image update parameters for the target image and determining whether the image update parameters have changed compared to the image generation parameters, the image editing region in the target image can be determined. The image editing region refers to the area in the target image where the parameters need to be updated. When determining the image editing region in the target image based on the image update parameters, the image editing region can be determined in the attention map output by the attention mechanism, and the image editing region in the attention map can be masked to generate a mask generation sequence.

[0112] Applying the solution of the disclosed embodiments, image update parameters for a target image are obtained; based on the image update parameters, an image editing region within the target image is determined; the image editing region is masked to obtain a mask generation sequence; and the mask generation sequence and the target image are input into an image generation model to obtain an updated target image. By regenerating the image editing region in the attention map based on the mask generation sequence, without affecting other regions whose parameters are not updated, identity-preserving image editing and similar structure image generation are achieved.

[0113] In an optional embodiment of the present disclosure, taking the diffusion model as an example, the above-mentioned inputting the mask generation sequence and the target image into the image generation model to obtain the updated target image may include the following steps:

[0114] For the current diffusion step of the image generation model, the diffusion result of the current diffusion step is generated according to the diffusion result of the completed diffusion step and the mask generation sequence until the iteration stop condition is reached to obtain the updated target image.

[0115] Specifically, the diffusion generative model process includes multiple diffusion steps. Each diffusion step can be viewed as a transformation of the image. These transformations help the diffusion model better simulate image degradation and restore a noise-free image from a noisy image. Generally speaking, the multiple diffusion steps of the diffusion model can be divided into two phases: the diffusion phase and the denoising phase. In the diffusion phase, the diffusion model simulates the image degradation process by gradually adding noise to the input image. In the denoising phase, the diffusion model identifies and removes noise from the noisy image to restore the original noise-free image.

[0116] In the embodiment of the present disclosure, in the image generation model, the subsequent editable image generation can be performed based on the diffusion results of each diffusion step generated in the previous time. During the editable image generation process, the diffusion result and the mask generation sequence can be multiplied to obtain the diffusion result of the current diffusion step until the iteration stop condition is reached to obtain the updated target image.

[0117] By applying the solution of the embodiment of the present disclosure, for the current diffusion step of the image generation model, the diffusion result of the current diffusion step is generated according to the diffusion result of the completed diffusion step and the mask generation sequence until the iteration stop condition is reached, and the updated target image is obtained, thereby realizing identity-preserving image editing and similar structure image generation.

[0118] In an optional embodiment of the present disclosure, after inputting the mask generation sequence and the target image into the image generation model to obtain the updated target image, the following steps may be further included:

[0119] The target image and the updated target image are sent to the client, so that the client displays the target image and the updated target image to the user.

[0120] It should be noted that after obtaining the updated target image, the target image and the updated target image can be sent to the client at the same time, and the user can select an image that meets actual needs from the target image and the updated target image displayed on the client.

[0121] In actual applications, there are many ways to send the target image and the updated target image to the client, and the specific selection is based on the actual situation. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the target image and the updated target image can be directly sent to the client so that the client can display the target image and the updated target image to the user. In another possible implementation of the present disclosure, the target image and the updated target image can be sent to the client based on the user's display demand information. The display demand information represents the user's demand for viewing the target image. The display demand information includes but is not limited to displaying only the target image and the updated target image, displaying image description information, the target image and the updated target image. The display demand information is specifically set according to the actual needs of the user, and the embodiments of the present disclosure do not impose any restrictions on this.

[0122] By applying the solution of the embodiment of the present disclosure, the target image and the updated target image are sent to the client, so that the client displays the target image and the updated target image to the user, providing the user with a variety of options. In addition, the user can intuitively compare the target image and the updated target image to determine whether the updated target image meets his or her needs.

[0123] In an optional embodiment of the present disclosure, after sending the target image and the updated target image to the client, the following steps may be further included:

[0124] Receive image selection information sent by the user through the client, and adjust model parameters of the image generation model based on the image selection information.

[0125] Specifically, the image selection information is used to identify the image selected by the user. The image selection information includes but is not limited to the image number and the image position. The selection is made according to the actual situation and the embodiment of the present disclosure does not impose any limitation on this.

[0126] It should be noted that the user can send image selection information in various ways, including but not limited to voice input, touch, and text input. Based on the image selection information, an image that meets the user's needs can be determined, and the image corresponding to the image selection information can be used as the real image to adjust the parameters of the image generation model.

[0127] By applying the solution of the embodiment of the present disclosure, image selection information sent by the user through the client is received, and the model parameters of the image generation model are adjusted based on the image selection information, so that the image generation model has the ability to interact with the user, so that the user can adjust the image generation results in an interactive manner, thereby improving user satisfaction.

[0128] In an optional embodiment of the present disclosure, after inputting the image generation parameters and the image description text into the image generation model and obtaining the target image corresponding to the image description text, the following steps may be further included:

[0129] Receive adjustment sample data sent by the user through the client, and adjust the model parameters of the image generation model according to the adjustment sample data, wherein the adjustment sample data is constructed based on the target image.

[0130] In actual applications, the image generation parameters and the image description text are input into the image generation model. After obtaining the target image corresponding to the image description text, the target image can be sent to the client. Furthermore, the user can process the target image by himself. If he is not satisfied with the target image, he can also send the adjustment sample data to train the model again. Specifically, after receiving the adjustment sample data sent by the user through the client, in a first possible implementation method, the model parameters of the image generation model can be adjusted according to the adjustment sample data; in a second possible implementation method, the model parameters of the parameter generation model can be adjusted according to the adjustment sample data; in a third possible implementation method, the model parameters of the image generation model and the parameter generation model can be adjusted according to the adjustment sample data. Among them, the adjustment sample data can be obtained by the user adding labels to the target image (such as positive and negative sample labels), or by adjusting the parameters of the target image (such as adjusting the size parameters). The embodiment of the present disclosure does not impose any restrictions on the generation method of the adjustment sample data.

[0131] It should be noted that the implementation method of "adjusting the model parameters of the image generation model according to the adjusted sample data" is the same as the training method of the image generation model, and will not be repeated in the embodiment of the present disclosure.

[0132] By applying the solution of the embodiment of the present disclosure, the target image is sent to the client so that the client displays the target image to the user; the adjustment sample data sent by the user through the client is received, and the model parameters of the image generation model are adjusted according to the adjustment sample data, so that the image generation model has the ability to interact with the user, so that the user can adjust the image generation result in an interactive manner, thereby improving user satisfaction.

[0133] In an optional embodiment of the present disclosure, before inputting the image description text and the generation prompt information into the parameter generation model to obtain the image generation parameters corresponding to the image description text, the following steps may be further included:

[0134] Acquire a sample set, wherein the sample set includes a plurality of sample image-text pairs, the sample image-text pairs include a sample image and a sample description text, and the sample image-text pairs carry sample parameter information;

[0135] Inputting multiple sample image-text pairs and prediction prompt information into the pre-trained language model to obtain image prediction parameters corresponding to the multiple sample image-text pairs;

[0136] According to the image prediction parameters and sample parameter information, the model parameters of the pre-trained language model are adjusted to obtain a parameter generation model that has completed training.

[0137] Specifically, the training method of the parameter generation model is supervised training based on prompt learning, that is, each sample image-text pair in the sample set carries real sample parameter information. The sample parameter information is the generation target of the parameter generation model, which is used to guide the training process of the parameter generation model.

[0138] The prediction prompt information is used to prompt the prediction process of the pre-trained language model. The prediction prompt information is set according to the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this. For example, the prediction prompt information can be "You are an intelligent bounding box generation model. I will provide you with images and image description texts. The format of each bounding box should be (object name, center point horizontal coordinate, center point vertical coordinate, bounding box width, bounding box height)". The prediction prompt information can enable the pre-trained language model to output prediction results in a fixed data format.

[0139] It should be noted that the implementation method of "inputting multiple sample image-text pairs and prediction prompt information into the pre-trained language model to obtain image prediction parameters corresponding to the multiple sample image-text pairs" can refer to the implementation method of "inputting image description text and generation prompt information into the parameter generation model to obtain image generation parameters corresponding to the image description text" mentioned above, and this disclosure will not repeat it. When adjusting the model parameters of the pre-trained language model based on the image prediction parameters and sample parameter information, the LoRA method can be used, and rank can be set to 8 to achieve model adjustment without forgetting the semantic understanding ability of the basic model.

[0140] In practical applications, there are multiple ways to obtain a sample set, and the specific method to be used depends on the actual situation. The embodiments of this disclosure do not impose any restrictions on this. In one possible implementation of this disclosure, a large number of sample image-text pairs carrying sample parameter information input by the user can be received to form a sample set. In another possible implementation of this disclosure, a large number of sample image-text pairs carrying sample parameter information can be read from other data acquisition devices or databases to form a sample set.

[0141] It is worth noting that the sample set may include at least one of an original sample subset, a generative sample subset, and a constructed sample subset. Specifically, a plurality of sample images carrying sample description texts may be obtained to form an original sample subset. The image understanding capability of the large model may be utilized to perform sample description text on the sample images to form a generative sample subset. Moreover, when forming the generative sample subset, a collaborative detection and segmentation model may be employed to extract a variety of sample parameter information. Parameter configuration conditions may be set based on prior knowledge, such as output rules, object rules, element number rules, etc., to construct pseudo data to form a constructed sample subset, thereby utilizing the constructed sample subset to enable the parameter generation model to have a purposeful learning capability.

[0142] Referring to Figure 5, a flowchart of a parameter generation model training method provided by one embodiment of the present disclosure is shown. The method comprises mixing an original sample subset, a generative sample subset, and a constructed sample subset to obtain a sample set. Multiple sample image-text pairs and prediction prompt information in the sample set are input into a pre-trained language model to obtain image prediction parameters corresponding to each of the multiple sample image-text pairs. Based on the image prediction parameters and sample parameter information, the model parameters of the pre-trained language model are adjusted to obtain a fully trained parameter generation model. By constructing a sample set including three sample subsets, the image generation parameters output by the parameter generation model can be made more reasonable.

[0143] By applying the solution of the embodiment of the present disclosure, the model parameters of the pre-trained language model are adjusted according to the image prediction parameters and sample parameter information to obtain a parameter generation model that has completed training. By continuously adjusting the model parameters of the pre-trained language model, the parameter generation model finally obtained can be made more accurate.

[0144] In an optional embodiment of the present disclosure, taking the generative sample subset included in the sample set as an example, the above-mentioned acquisition of the sample set may include the following steps:

[0145] Acquire multiple sample images;

[0146] For the first sample image, input the first sample image and the constructed prompt information into the pre-trained language model to obtain a first sample description text of the first sample image;

[0147] generating first sample parameter information of the first sample image according to the first sample image and the first sample description text;

[0148] A sample set is constructed according to a plurality of sample images, sample description texts of the plurality of sample images, and sample parameter information.

[0149] Specifically, the prompt information is constructed to prompt the pre-trained language model to generate sample description text for the sample image. The specific configuration of the prompt information is based on the actual situation and is not limited in this embodiment of the present disclosure. For example, the predicted prompt information may be "Please describe the image."

[0150] It should be noted that there are various ways to obtain multiple sample images, and the method to be used depends on the actual situation. The embodiments of the present disclosure do not impose any restrictions on this method. In one possible implementation of the present disclosure, multiple sample images can be received from a user. In another possible implementation of the present disclosure, multiple sample images can be read from other data acquisition devices or databases.

[0151] Applying the solution of the disclosed embodiment, multiple sample images are obtained; for a first sample image, the first sample image and construction prompt information are input into a pre-trained language model to obtain a first sample description text for the first sample image; based on the first sample image and the first sample description text, first sample parameter information for the first sample image is generated; and based on the multiple sample images, the sample description texts for the multiple sample images, and the sample parameter information, a sample set is constructed. The sample description text is generated using the image understanding capabilities of the pre-trained language model, ensuring the readiness of the sample description text.

[0152] In practical applications, there are various ways to generate first sample parameter information for the first sample image based on the first sample image and the first sample description text. The method of selecting the method depends on the actual situation and is not limited in the present embodiment. In one possible implementation of the present disclosure, relevant keywords, sentence structure, and other information can be obtained from the first sample description text through methods such as text mining and image analysis, and this information can be used as the first sample parameter information.

[0153] In another possible implementation of the present disclosure, a detection segmentation model may be coordinated to extract multiple sample parameter information. That is, the above-mentioned step of generating first sample parameter information of the first sample image based on the first sample image and the first sample description text may include the following steps:

[0154] Performing region detection on the first sample image according to the first sample description text to determine key region position information of the first sample image;

[0155] Visual segmentation is performed on the first sample image according to the key area position information to obtain first sample parameter information of the first sample image.

[0156] Specifically, the key region location information refers to the coordinates of the center point of the key region in the first sample image. The key region can be understood as a bounding box. For example, the first sample image shows a little girl and a dog running. The key regions in the first sample image are the area where the little girl and the dog are located. The first sample parameter information includes, but is not limited to, the height and width of the key regions.

[0157] It should be noted that when performing region detection on the first sample image based on the first sample description text, the first sample description text and the first sample image can be input into a detection and segmentation model. After region detection by the detection and segmentation model, key region position information of the first sample image can be obtained. When performing visual segmentation on the first sample image based on the key region position information, the key region position information and the first sample image can be input into a visual segmentation model (SAM, Segmentation As A Module) to obtain first sample parameter information of the first sample image.

[0158] Using the solution of the disclosed embodiment, region detection is performed on the first sample image based on the first sample description text to determine the location information of the key regions of the first sample image. Based on this key region location information, visual segmentation is performed on the first sample image to obtain first sample parameter information for the first sample image. Region detection and visual segmentation improve the accuracy of the first sample parameter information.

[0159] In an optional embodiment of the present disclosure, a method for constructing a constructed sample subset is described by taking a constructed sample subset included in a sample set as an example. After generating the first sample parameter information of the first sample image based on the first sample image and the first sample description text, the following steps may also be included:

[0160] Get parameter configuration conditions;

[0161] When the first sample parameter information does not meet the parameter configuration condition, adjusting the first sample parameter information to obtain adjusted first sample parameter information;

[0162] A sample set is constructed based on multiple sample images, sample description texts of the multiple sample images, and sample parameter information, including:

[0163] A sample set is constructed according to the plurality of sample images, the sample description texts of the plurality of sample images and the adjusted sample parameter information.

[0164] Specifically, the parameter configuration conditions include but are not limited to brightness greater than a preset brightness threshold, clarity greater than a preset clarity threshold, etc., which are selected according to actual conditions, and the embodiments of the present disclosure do not impose any restrictions on this.

[0165] In practical applications, there are multiple ways to obtain parameter configuration conditions, and the specific method to be selected depends on the actual situation. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the parameter configuration conditions can be received from a user input. In another possible implementation of the present disclosure, the parameter configuration conditions can be read from other data acquisition devices or databases.

[0166] It should be noted that after obtaining the parameter configuration conditions, it is further possible to determine whether the first sample parameter information satisfies the parameter configuration conditions. If the first sample parameter information satisfies the parameter configuration conditions, the first sample parameters are not adjusted. If the first sample parameter information does not satisfy the parameter configuration conditions, the first sample parameter information is adjusted to obtain adjusted first sample parameter information.

[0167] Applying the solution of the disclosed embodiment, parameter configuration conditions are obtained; if first sample parameter information does not meet the parameter configuration conditions, the first sample parameter information is adjusted to obtain adjusted first sample parameter information; and a sample set is constructed based on multiple sample images, sample description texts of the multiple sample images, and the adjusted sample parameter information. The reliability of the sample parameter information is monitored using the parameter configuration conditions, thereby improving the accuracy of the sample parameter information.

[0168] In an optional embodiment of the present disclosure, before inputting the image generation parameters and the image description text into the image generation model and obtaining the target image corresponding to the image description text, the following steps may be further included:

[0169] Acquire a sample set, wherein the sample set includes a plurality of sample image-text pairs, the sample image-text pairs include a sample image and a sample description text, and the sample image-text pairs carry sample parameter information;

[0170] Inputting a plurality of sample description texts and sample parameter information corresponding to the plurality of sample description texts into an initial generation model to obtain prediction images corresponding to the plurality of sample description texts respectively;

[0171] According to the predicted image and the sample image, the model parameters of the initial generation model are adjusted to obtain a trained image generation model.

[0172] Specifically, the training method of the image generation model is supervised training, that is, the sample description text in the sample set carries real sample images, and the sample images are the generation targets of the image generation model, which are used to guide the training process of the image generation model.

[0173] It should be noted that the method for "obtaining a sample set" can refer to the method for obtaining a sample set during the training of the parameter generation model described above. The implementation method of "inputting multiple sample description texts and sample parameter information corresponding to the multiple sample description texts into the initial generation model to obtain predicted images corresponding to the multiple sample description texts" can refer to the implementation method of "inputting image generation parameters and image description texts into the image generation model to obtain target images corresponding to the image description texts" described above, and will not be further described in this embodiment of the disclosure.

[0174] In actual applications, when adjusting the model parameters of the initial generation model according to the predicted image and the sample image, the loss value can be calculated according to the predicted image and the sample image, and the model parameters of the initial generation model can be adjusted according to the loss value until the preset stopping condition is reached to obtain a trained image generation model. Among them, there are many functions for calculating the loss value, such as the cross entropy loss function, the L1 norm loss function, the maximum loss function, the mean square error loss function, the logarithmic loss function, etc. The specific selection is based on the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this.

[0175] In a possible implementation of the present disclosure, the preset stopping condition includes a loss value being less than or equal to a preset threshold value. After calculating the loss value based on the predicted image and the sample image, the loss value is compared with the preset threshold value.

[0176] Specifically, if the loss value is greater than the preset threshold, it means that the difference between the predicted image and the sample image is large, and the initial generation model has poor prediction ability for the predicted image. At this time, the model parameters of the initial generation model can be adjusted, and the initial generation model can continue to be trained until the loss value is less than or equal to the preset threshold, indicating that the difference between the predicted image and the sample image is small, and the preset stopping condition is met, and the trained image generation model is obtained.

[0177] In another possible implementation of the present disclosure, in addition to comparing the loss value and the preset threshold, the number of iterations may be combined to determine whether the current initial generation model has been trained.

[0178] Specifically, if the loss value is greater than a preset threshold, the model parameters of the initial generation model are adjusted, and the initial information extraction model is continued to be trained until the preset number of iterations is reached, and the iteration is stopped to obtain a trained image generation model. The preset threshold and the preset number of iterations are selected according to actual conditions, and the embodiments of the present disclosure do not impose any restrictions on this.

[0179] By applying the solution of the embodiment of the present disclosure, the model parameters of the initial generation model are adjusted according to the predicted image and the sample image to obtain a trained image generation model. By continuously adjusting the model parameters of the initial generation model, the final image generation model can be made more accurate.

[0180] 6 , which shows a flow chart of another parameter generation model training method provided by an embodiment of the present disclosure. The parameter generation model training is applied to a cloud-side device and specifically includes the following steps:

[0181] Step 602: Acquire a sample set, wherein the sample set includes a plurality of sample image-text pairs, the sample image-text pairs include a sample image and a sample description text, and the sample image-text pairs carry sample parameter information.

[0182] Step 604: Input the plurality of sample image-text pairs and prediction prompt information into the pre-trained language model to obtain image prediction parameters corresponding to the plurality of sample image-text pairs.

[0183] Step 606: Adjust the model parameters of the pre-trained language model according to the image prediction parameters and the sample parameter information to obtain a parameter generation model that has completed training.

[0184] It should be noted that the implementation method of steps 602 to 606 is detailed in the training method of the parameter generation model in the above-mentioned image generation method, and the present embodiment does not impose any limitation on this.

[0185] In actual applications, after obtaining the trained parameter generation model, the model parameters of the trained parameter generation model can be sent to the terminal device, so that the user can locally build the parameter generation model based on the model parameters to generate image generation parameters.

[0186] By applying the solution of the embodiment of the present disclosure, the model parameters of the pre-trained language model are adjusted according to the image prediction parameters and sample parameter information to obtain a parameter generation model that has completed training. By continuously adjusting the model parameters of the pre-trained language model, the parameter generation model finally obtained can be made more accurate.

[0187] Referring to FIG. 7 , FIG. 7 shows a flow chart of an automatic question-answering method provided by an embodiment of the present disclosure, which specifically includes the following steps:

[0188] Step 702: Receive an image question and answer request, wherein the image question and answer request carries image description text.

[0189] Step 704: Input the image description text and generation prompt information into the parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and sample parameter information carried by the multiple sample image-text pairs.

[0190] Step 706: Input the image generation parameters and the image description text into the image generation model to obtain a response image corresponding to the image question and answer request.

[0191] It should be noted that the implementation of steps 702 to 706 is detailed in steps 302 to 306 above, and the present embodiment does not impose any limitation on this.

[0192] By applying the solution of the embodiment of the present disclosure, the visual elements of the image are semantically decomposed by using a parameter generation model to obtain image generation parameters, and then accurate image generation is completed based on the image generation parameters, so that the target image can clearly express the image description text and image generation parameters, thereby improving the interpretability and controllability of the automatic question-answering process.

[0193] Referring to Figure 8, Figure 8 shows a flowchart of the processing process of another image generation method provided by one embodiment of the present disclosure. During the image generation process, the powerful semantic understanding and associative capabilities of the parameter generation model are utilized to complete the semantic decomposition and reorganization of the image's visual elements. After parameterization, these elements are input into the image generation model to achieve accurate image generation. The image generation stage can be divided into two stages: the parameter generation model processing stage and the image generation model processing stage. The parameter generation model processing stage includes fine-tuning the pre-trained language model, constructing a conditional generation sample set, and designing generated prompt information.

[0194] As shown in FIG8 , a parameter generation model is obtained by training a pre-trained language model based on multiple sample image-text pairs (composed of sample description texts and sample images) and sample parameter information carried by multiple sample image-text pairs; obtaining image description text; inputting the image description text and generation prompt information into the parameter generation model to obtain image generation parameters corresponding to the image description text; inputting the image generation parameters and the image description text into the image generation model to obtain a target image corresponding to the image description text.

[0195] By applying the solution of the embodiment of the present disclosure, a two-stage diffusion generation architecture is used to implement assisted generation based on a pre-trained language model. In the parameter generation model processing stage, the emergence ability and appropriate fine-tuning of the pre-trained language model are used to output the image generation parameters required for the image generation model processing stage. Subsequently, in the image generation model processing stage, image generation based on the image generation parameter control is performed to output an image with precise semantics. Moreover, since the generated image can clearly express the generation prompt information or image generation parameters used in the generation process, the user can understand and verify the accuracy of the generation result.

[0196] Referring to FIG9 , FIG9 shows a schematic diagram of a processing process of a parameter generation model provided by an embodiment of the present disclosure. As shown in FIG9 , the image description text “Two black cats and a white dog on an orange sofa” is obtained, and the image description text and the generated prompt information 1 “My title is “Two black cats and a white dog on an orange sofa” and “What are the elements in the image?” are input into the parameter generation model to obtain the image generation parameter 1 “[(black cat, 2), (white dog, 1), (sofa, 1)]” corresponding to the image description text; the image description text and the generated prompt information 2 “What are their bounding boxes?” are input into the parameter generation model to obtain the image generation parameter 2 “[(black cat, [30, 171 ,212,286]),(black cat,[40,200,123,412]),(white dog,[24,543,231,332]),(sofa,[264,173,222,221])]”; input the image description text and the generated prompt information 3 “What is the main color of each element?” into the parameter generation model to obtain the image generation parameters 3 corresponding to the image description text “[(black cat,[black, gray, dark gray]),...,(white dog,[white, light gray, brown]),(sofa,[orange, brown, gray])]”; finally, aggregate the image generation parameters 1, image generation parameters 2 and image generation parameters 3 to obtain the final image generation parameters.

[0197] See Figure 10, which shows a schematic diagram of an image generation interface provided by one embodiment of the present disclosure. The image generation interface is divided into a request input interface and a result display interface. The request input interface includes a request input box, an "OK" control, and a "Cancel" control. The result display interface includes a result display box.

[0198] The user enters an image generation request through the request input box displayed on the client. The image generation request carries an image description. The user clicks the "OK" control. The server receives the image description sent by the client and inputs the image description and generation prompt information into the parameter generation model to obtain image generation parameters corresponding to the image description. The image generation parameters are used to describe the visual features of the image. The parameter generation model is trained based on multiple sample image-text pairs and the sample parameter information carried by these sample image-text pairs. The image generation parameters and image description are input into the image generation model to obtain the target image corresponding to the image description. The target image is then sent to the client. The client displays the target image in the result display box.

[0199] In actual applications, users can operate controls by clicking, double-clicking, touching, hovering the mouse, sliding, long pressing, voice control, or shaking, etc. The specific selection is based on the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this.

[0200] Corresponding to the above-mentioned image generation method embodiment, the present disclosure also provides an image generation device embodiment. FIG11 shows a schematic structural diagram of an image generation device provided by an embodiment of the present disclosure. As shown in FIG11 , the device includes:

[0201] A first acquisition module 1102 is configured to acquire image description text;

[0202] A first input module 1104 is configured to input the image description text and the generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and sample parameter information carried by the multiple sample image-text pairs;

[0203] The second input module 1106 is configured to input the image generation parameters and the image description text into the image generation model to obtain a target image corresponding to the image description text.

[0204] Optionally, the image generation model includes a parameter encoding unit and a coding unit; the second input module 1106 is further configured to input the image generation parameters and the image description text into the image generation model, and encode the image generation parameters through the parameter encoding unit to obtain parameter encoding features; and generate a target image corresponding to the image description text according to the parameter encoding features and the image description text through the coding unit.

[0205] Optionally, the parameter coding unit includes a one-dimensional parameter coding unit, a two-dimensional parameter coding unit and a feature aggregation unit; the second input module 1106 is further configured to encode the one-dimensional parameters in the image generation parameters through the one-dimensional parameter coding unit to obtain one-dimensional coding features; encode the two-dimensional parameters in the image generation parameters through the two-dimensional parameter coding unit to obtain two-dimensional coding features; and aggregate the one-dimensional coding features and the two-dimensional coding features through the feature aggregation unit to obtain parameter coding features.

[0206] Optionally, the device also includes: a third acquisition module, configured to obtain image update parameters for the target image; determine the image editing area in the target image based on the image update parameters; mask the image editing area to obtain a mask generation sequence; input the mask generation sequence and the target image into the image generation model to obtain an updated target image.

[0207] Optionally, the device further includes: a first sending module configured to send the target image and the updated target image to the client, so that the client displays the target image and the updated target image to the user.

[0208] Optionally, the device further includes: a second receiving module configured to receive image selection information sent by the user through the client, and adjust model parameters of the image generation model based on the image selection information.

[0209] Optionally, the device further includes: a third receiving module configured to receive adjustment sample data sent by the user through the client, and adjust model parameters of the image generation model according to the adjustment sample data, wherein the adjustment sample data is constructed based on the target image.

[0210] Optionally, the device also includes: a parameter generation model training module, configured to obtain a sample set, wherein the sample set includes multiple sample image-text pairs, the sample image-text pairs include sample images and sample description texts, and the sample image-text pairs carry sample parameter information; input the multiple sample image-text pairs and prediction prompt information into a pre-trained language model to obtain image prediction parameters corresponding to the multiple sample image-text pairs; adjust the model parameters of the pre-trained language model according to the image prediction parameters and sample parameter information to obtain a parameter generation model that has completed training.

[0211] Optionally, the parameter generation model training module is further configured to obtain a plurality of sample images;

[0212] For the first sample image, the first sample image and the construction prompt information are input into the pre-trained language model to obtain the first sample description text of the first sample image; based on the first sample image and the first sample description text, the first sample parameter information of the first sample image is generated; based on multiple sample images, the sample description texts of the multiple sample images and the sample parameter information, a sample set is constructed.

[0213] Optionally, the parameter generation model training module is further configured to perform region detection on the first sample image based on the first sample description text to determine the key area position information of the first sample image; perform visual segmentation on the first sample image based on the key area position information to obtain first sample parameter information of the first sample image.

[0214] Optionally, the device also includes: a fourth acquisition module, configured to obtain parameter configuration conditions; when the first sample parameter information does not meet the parameter configuration conditions, adjusting the first sample parameter information to obtain the adjusted first sample parameter information; a parameter generation model training module, further configured to construct a sample set based on multiple sample images, sample description texts of multiple sample images and the adjusted sample parameter information.

[0215] Optionally, the device also includes: an image generation model training module, configured to obtain a sample set, wherein the sample set includes multiple sample image-text pairs, the sample image-text pairs include sample images and sample description texts, and the sample image-text pairs carry sample parameter information; input the multiple sample description texts and the sample parameter information corresponding to the multiple sample description texts into the initial generation model to obtain predicted images corresponding to the multiple sample description texts; adjust the model parameters of the initial generation model according to the predicted images and the sample images to obtain a trained image generation model.

[0216] By applying the solution of the embodiment of the present disclosure, the visual elements of the image are semantically decomposed by using a parameter generation model to obtain image generation parameters, and then accurate image generation is completed based on the image generation parameters, so that the target image can clearly express the image description text and image generation parameters, thereby improving the interpretability and controllability of image generation.

[0217] The above is a schematic diagram of an image generation device according to this embodiment. It should be noted that the technical solution of this image generation device and the technical solution of the aforementioned image generation method are based on the same concept. For details not described in detail in the technical solution of the image generation device, please refer to the description of the technical solution of the aforementioned image generation method.

[0218] Corresponding to the above-mentioned parameter generation model training method embodiment, the present disclosure also provides a parameter generation model training device embodiment. Figure 12 shows a schematic diagram of the structure of a parameter generation model training device provided by one embodiment of the present disclosure. As shown in Figure 12, the device is applied to a cloud-side device and includes:

[0219] The second acquisition module 1202 is configured to acquire a sample set, wherein the sample set includes a plurality of sample image-text pairs, each of the sample image-text pairs includes a sample image and a sample description text, and each of the sample image-text pairs carries sample parameter information;

[0220] The third input module 1204 is configured to input a plurality of sample image-text pairs and prediction prompt information into the pre-trained language model to obtain image prediction parameters corresponding to the plurality of sample image-text pairs;

[0221] The adjustment module 1206 is configured to adjust the model parameters of the pre-trained language model according to the image prediction parameters and the sample parameter information to obtain a parameter generation model that has completed training.

[0222] By applying the solution of the embodiment of the present disclosure, the model parameters of the pre-trained language model are adjusted according to the image prediction parameters and sample parameter information to obtain a parameter generation model that has completed training. By continuously adjusting the model parameters of the pre-trained language model, the parameter generation model finally obtained can be made more accurate.

[0223] The above is a schematic diagram of a parameter generation model training device according to this embodiment. It should be noted that the technical solution of the parameter generation model training device and the technical solution of the parameter generation model training method described above are based on the same concept. For details not described in detail in the technical solution of the parameter generation model training device, please refer to the description of the technical solution of the parameter generation model training method described above.

[0224] Corresponding to the above-mentioned automatic question-answering method embodiment, the present disclosure also provides an automatic question-answering device embodiment. FIG13 shows a schematic structural diagram of an automatic question-answering device provided by one embodiment of the present disclosure. As shown in FIG13 , the device includes:

[0225] The first receiving module 1302 is configured to receive an image question and answer request, wherein the image question and answer request carries an image description text;

[0226] A fourth input module 1304 is configured to input the image description text and the generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and the sample parameter information carried by the multiple sample image-text pairs;

[0227] The fifth input module 1306 is configured to input image generation parameters and image description text into the image generation model to obtain a reply image corresponding to the image question and answer request.

[0228] By applying the solution of the embodiment of the present disclosure, the visual elements of the image are semantically decomposed by using a parameter generation model to obtain image generation parameters, and then accurate image generation is completed based on the image generation parameters, so that the target image can clearly express the image description text and image generation parameters, thereby improving the interpretability and controllability of the automatic question-answering process.

[0229] The above is a schematic diagram of an automatic question-answering device according to this embodiment. It should be noted that the technical solution of this automatic question-answering device and the technical solution of the automatic question-answering method described above are based on the same concept. For details not described in detail in the technical solution of the automatic question-answering device, please refer to the description of the technical solution of the automatic question-answering method described above.

[0230] Figure 14 shows a block diagram of a computing device according to an embodiment of the present disclosure. Components of the computing device 1400 include, but are not limited to, a memory 1410 and a processor 1420. The processor 1420 is connected to the memory 1410 via a bus 1430, and a database 1450 is used to store data.

[0231] Computing device 1400 also includes an access device 1440 that enables computing device 1400 to communicate via one or more networks 1460. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 1440 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a World Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0232] In one embodiment of the present disclosure, the aforementioned components of computing device 1400 and other components not shown in FIG14 may also be connected to each other, for example, via a bus. It should be understood that the block diagram of the computing device structure shown in FIG14 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.

[0233] Computing device 1400 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1400 may also be a mobile or stationary server.

[0234] Among them, the processor 1420 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned image generation method or parameter generation model training method or automatic question-answering method.

[0235] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solutions of the aforementioned image generation method, parameter generation model training method, and automatic question-answering method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the aforementioned image generation method, parameter generation model training method, or automatic question-answering method.

[0236] An embodiment of the present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned image generation method or parameter generation model training method or automatic question-answering method.

[0237] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the aforementioned image generation method, parameter generation model training method, and automatic question-answering method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the aforementioned image generation method, parameter generation model training method, or automatic question-answering method.

[0238] An embodiment of the present disclosure further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned image generation method or parameter generation model training method or automatic question-answering method.

[0239] The above is an illustrative embodiment of a computer program. It should be noted that the technical solution of this computer program is based on the same concept as the technical solutions of the aforementioned image generation method, parameter generation model training method, and automatic question-answering method. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the aforementioned image generation method, parameter generation model training method, or automatic question-answering method.

[0240] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0241] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0242] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present disclosure.

[0243] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0244] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. A method for generating an image, comprising: Get the image description text; Input the image description text and the generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and sample parameter information carried by the multiple sample image-text pairs; The image generation parameters and the image description text are input into an image generation model to obtain a target image corresponding to the image description text.

2. According to the method of claim 1, the image generation model comprises a parameter encoding unit and a coding and decoding unit; The step of inputting the image generation parameters and the image description text into an image generation model to obtain a target image corresponding to the image description text includes: Inputting the image generation parameters and the image description text into an image generation model, and encoding the image generation parameters through the parameter encoding unit to obtain parameter encoding features; The encoding and decoding unit generates a target image corresponding to the image description text according to the parameter encoding features and the image description text.

3. According to the method of claim 2, the parameter encoding unit comprises a one-dimensional parameter encoding unit, a two-dimensional parameter encoding unit and a feature aggregation unit; The parameter encoding unit generates parameter encoding for the image to obtain parameter encoding features, including: The one-dimensional parameter encoding unit encodes the one-dimensional parameter in the image generation parameter to obtain a one-dimensional encoding feature; The two-dimensional parameter encoding unit encodes the two-dimensional parameters in the image generation parameters to obtain two-dimensional encoding features; The one-dimensional coding feature and the two-dimensional coding feature are aggregated by the feature aggregation unit to obtain a parameter coding feature.

4. The method according to any one of claims 1 to 3, wherein after inputting the image generation parameters and the image description text into an image generation model to obtain a target image corresponding to the image description text, the method further comprises: Acquire image update parameters for the target image; Determining an image editing area in the target image according to the image update parameters; Masking the image editing area to obtain a mask generation sequence; The mask generation sequence and the target image are input into the image generation model to obtain an updated target image.

5. The method according to claim 4, after inputting the mask generation sequence and the target image into the image generation model to obtain the updated target image, further comprising: The target image and the updated target image are sent to a client, so that the client displays the target image and the updated target image to a user.

6. The method according to claim 5, after sending the target image and the updated target image to the client, further comprising: The image selection information sent by the user through the client is received, and the model parameters of the image generation model are adjusted based on the image selection information.

7. The method according to any one of claims 1 to 6, after inputting the image generation parameters and the image description text into an image generation model to obtain a target image corresponding to the image description text, further comprising: Receive adjustment sample data sent by the user through the client, and adjust the image generation according to the adjustment sample data The adjustment sample data is constructed based on the target image.

8. The method according to any one of claims 1 to 7, before inputting the image description text and the generation prompt information into a parameter generation model to obtain the image generation parameters corresponding to the image description text, further comprising: Acquire a sample set, wherein the sample set includes a plurality of sample image-text pairs, the sample image-text pairs include a sample image and a sample description text, and the sample image-text pairs carry sample parameter information; Inputting the plurality of sample image-text pairs and prediction prompt information into a pre-trained language model to obtain image prediction parameters corresponding to the plurality of sample image-text pairs respectively; According to the image prediction parameters and the sample parameter information, the model parameters of the pre-trained language model are adjusted to obtain a parameter generation model that has completed training.

9. The method according to claim 8, wherein obtaining a sample set comprises: Acquire multiple sample images; For a first sample image, inputting the first sample image and construction prompt information into a pre-trained language model to obtain a first sample description text of the first sample image; generating first sample parameter information of the first sample image according to the first sample image and the first sample description text; A sample set is constructed according to the multiple sample images, the sample description texts of the multiple sample images, and the sample parameter information.

10. The method according to claim 9, wherein generating first sample parameter information of the first sample image according to the first sample image and the first sample description text comprises: Performing region detection on the first sample image according to the first sample description text to determine key region position information of the first sample image; Visually segment the first sample image according to the key area position information to obtain first sample parameter information of the first sample image.

11. The method according to claim 9, after generating the first sample parameter information of the first sample image according to the first sample image and the first sample description text, further comprising: Get parameter configuration conditions; If the first sample parameter information does not meet the parameter configuration condition, adjust the first sample parameter information to obtain adjusted first sample parameter information; The constructing a sample set according to the plurality of sample images, sample description texts of the plurality of sample images and sample parameter information includes: A sample set is constructed according to the multiple sample images, the sample description texts of the multiple sample images, and the adjusted sample parameter information.

12. The method according to any one of claims 1 to 11, before inputting the image generation parameters and the image description text into an image generation model to obtain a target image corresponding to the image description text, further comprising: Acquire a sample set, wherein the sample set includes a plurality of sample image-text pairs, the sample image-text pairs include a sample image and a sample description text, and the sample image-text pairs carry sample parameter information; Inputting a plurality of sample description texts and sample parameter information corresponding to the plurality of sample description texts into an initial generation model to obtain prediction images corresponding to the plurality of sample description texts respectively; According to the predicted image and the sample image, the model parameters of the initial generation model are adjusted to obtain a completed training Image generation model trained.

13. A parameter generation model training method, applied to a cloud-side device, comprising: Acquire a sample set, wherein the sample set includes a plurality of sample image-text pairs, the sample image-text pairs include a sample image and a sample description text, and the sample image-text pairs carry sample parameter information; Inputting the plurality of sample image-text pairs and prediction prompt information into a pre-trained language model to obtain image prediction parameters corresponding to the plurality of sample image-text pairs respectively; According to the image prediction parameters and the sample parameter information, the model parameters of the pre-trained language model are adjusted to obtain a parameter generation model that has completed training.

14. An automatic question-answering method, comprising: receiving an image question and answer request, wherein the image question and answer request carries an image description text; Input the image description text and the generation prompt information into a parameter generation model to obtain image generation parameters corresponding to the image description text, wherein the image generation parameters are used to describe the visual features of the image, and the parameter generation model is trained based on multiple sample image-text pairs and sample parameter information carried by the multiple sample image-text pairs; The image generation parameters and the image description text are input into an image generation model to obtain a reply image corresponding to the image question and answer request.

15. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 12 or claim 13 or claim 14 are implemented.

16. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method described in any one of claims 1 to 12 or claim 13 or claim 14.

17. A computer program, wherein when the computer program is executed by a computer, the method according to any one of claims 1 to 12 or claim 13 or claim 14 is implemented.

Citation Information

Patent Citations

  • Pre-training model training method and device, pre-training model application method and device, electronic equipment and medium

    CN116186545A

  • Training method of image editing model and image editing method and device

    CN116363261A

  • Figure graph model training method and text graph method

    CN116935169A

  • Generating images using sequences of generative neural networks

    US20230377226A1

Cited By

  • Hyperspectral image super-resolution method and device based on iterative reconstruction and spectral refining

    CN121526883A

  • Model medium-term training sample selection method and device based on value perception

    CN121834285A