Image generation method and device, equipment, storage medium and vehicle

By expanding the prompt words entered by the user, and determining the image generation model and sampler with the preset correspondence relationship, the problem of low image generation quality in the prior art is solved, and higher quality and personalized image generation is achieved.

CN120182971APending Publication Date: 2025-06-20BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311755370.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, the prompt words entered by the user in the literary picture task cannot fully express the user's personalized needs, resulting in a low quality of the generated image.

Method used

By obtaining the prompt words entered by the user, text expansion is performed to enrich the text content, and appropriate image generation models and samplers are determined based on pre-set correspondence, and noise reduction parameters are updated to generate target images.

Benefits of technology

It improves the accuracy and richness of image generation, improves image quality, and can express users' needs more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182971A_ABST
    Figure CN120182971A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and device, equipment, a storage medium and a vehicle. The method comprises the steps of obtaining a cue word input by a user, wherein the cue word is used for indicating generation of a target image; according to a preset first corresponding relation, an image generation model and a sampler corresponding to the cue word are determined, the first corresponding relation is the corresponding relation among the cue word, the image generation model and the sampler, and the sampler is used for controlling the noise reduction mode of the image generation model; performing text expansion on the cue word to obtain a final text; loading the sampler into the image generation model to update noise reduction parameters of the image generation model; and converting the final text into a target image by using the image generation model after the noise reduction parameter is updated. According to the embodiment of the invention, the quality of the generated image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to an image generation method, apparatus, device, storage medium, and vehicle. Background Art

[0002] Text-to-image is a computer generation task aimed at converting text descriptions or natural language texts into corresponding images. In this task, the computer model needs to understand the prompt words input by the user and generate images that match the prompt words.

[0003] However, in the related art, all text-to-image tasks are inferred based on the same set of standard algorithmic processes, that is, the user's prompt words are input into a pre-set image generation model, and the image generation model uses the same set of algorithms to convert the prompt words into images that match the prompt words. Since the prompt words input by the user are often relatively simple and cannot fully express all the user's personalized needs, the images obtained by inferring the concise prompt words using the same set of standard algorithmic processes are often monotonous, resulting in the generated images being unable to accurately express the user's needs, that is, the quality of image generation is low. Summary of the Invention

[0004] Embodiments of this application provide an image generation method, apparatus, device, storage medium, and vehicle, which can solve the problem of low quality of images obtained by converting existing prompt words.

[0005] In a first aspect, embodiments of this application provide an image generation method, the method including:

[0006] Obtain the prompt words input by the user, where the prompt words are used to indicate the generation of a target image;

[0007] According to a pre-set first correspondence, determine the image generation model and sampler corresponding to the prompt words, where the first correspondence is the correspondence among the prompt words, the image generation model, and the sampler, and the sampler is used to control the noise reduction method of the image generation model;

[0008] Perform text expansion on the prompt words to obtain the final text;

[0009] Load the sampler into the image generation model to update the noise reduction parameters of the image generation model;

[0010] Use the image generation model with updated noise reduction parameters to convert the final text into the target image.

[0011] In some embodiments, the image generation model includes a basic image generation model and a specific object generation model, and the first correspondence includes a model correspondence and a sampler correspondence. Determining the image generation model and sampler corresponding to the prompt according to the preset first correspondence includes:

[0012] Perform semantic recognition on the prompt to determine the image generation intention corresponding to the prompt;

[0013] Identify the specific generation information included in the image generation intention corresponding to the prompt; the specific generation information includes a specific subject and / or a specific style;

[0014] According to the model correspondence, determine the basic image generation model and the specific object generation model corresponding to the prompt. The model correspondence includes the correspondence between the specific generation information and the specific object generation model, and the correspondence between the image generation intention and the basic image generation model;

[0015] According to the sampler correspondence, determine the sampler corresponding to the prompt, the basic image generation model, and the specific object generation model. The sampler correspondence includes the correspondence among the image generation intention, the specific generation information, the basic image generation model, the specific object generation model, and the sampler.

[0016] In some embodiments, determining the sampler corresponding to the prompt, the basic image generation model, and the specific object generation model according to the sampler correspondence includes:

[0017] Obtain the identification information of the image generation intention, the specific generation information, the basic image generation model, and the specific object generation model;

[0018] Determine the identification information as the initialization parameter of the sampler;

[0019] Match the sampler corresponding to the initialization parameter from the sampler correspondence. The sampler correspondence includes the correspondence between the initialization parameter and the sampler.

[0020] In some embodiments, expanding the text of the prompt to obtain the final text includes:

[0021] Obtain the model keyword corresponding to the image generation model, and add the model keyword to the prompt to obtain the first intermediate text;

[0022] In the case where there is a special phrase representing a special object in the prompt, obtain the object keyword corresponding to the special object, and add the object keyword to the first intermediate text to obtain a second intermediate text;

[0023] Convert the second intermediate text into the final text.

[0024] In some embodiments, the converting the second intermediate text into the final text includes:

[0025] In the case where an informal phrase is detected in the second intermediate text, query the common phrase corresponding to the informal phrase according to a preset second correspondence, where the second correspondence is the correspondence between the informal phrase and the common phrase;

[0026] Replace the corresponding informal phrase in the second intermediate text with the common phrase to obtain the final text.

[0027] In some embodiments, the image generation model includes a basic image generation model and a specific object generation model. After determining the image generation model and sampler corresponding to the prompt according to a preset first correspondence, the method further includes:

[0028] Determine a specific object module in the basic image generation model that matches the function of the specific object generation model;

[0029] Obtain the input data of the specific object module, and input the input data into the specific object generation model;

[0030] Obtain the module output data obtained by the specific object module in response to the input data, and obtain the model output data obtained by the specific object generation model in response to the input data;

[0031] Merge the module output data and the model output data to obtain fusion data;

[0032] Update the fusion data as the model parameters corresponding to the specific object in the basic image generation model, so as to mount the specific object generation model to the basic image generation model.

[0033] In some embodiments, the obtaining the model output data obtained by the specific object generation model in response to the input data includes:

[0034] Obtain the model weights in the specific object generation model;

[0035] Input the model weights into the specific object generation model to update the weight parameters of the specific object generation model;

[0036] Using the specific object generation model with updated weight parameters to generate a model response to the input data to obtain the model output data.

[0037] In some embodiments, the specific object generation model includes at least one specific object generation node, the model weights include at least one node weight, and each specific object generation node corresponds to a node weight. Inputting the model weights into the specific object generation model includes:

[0038] Adding at least one identity node to the specific object generation model, the output of the identity node being equal to the input of the identity node, and each identity node corresponding to a specific object generation node;

[0039] Inputting each node weight in the at least one node weight into the corresponding specific object generation node through the corresponding identity node.

[0040] In some embodiments, the sampler includes a single sampling algorithm and an iteration count. Using the image generation model updated with the noise reduction parameter to convert the final text into the target image includes:

[0041] Obtaining a randomly generated noise image;

[0042] Encoding the final text using a text vectorization encoding method to obtain a basic text embedding vector;

[0043] Embedding the basic text embedding vector into the image generation model and using the image generation model to perform N noise reduction processes on the noise image according to the single sampling algorithm to obtain a basic latent feature image, where N is the iteration count;

[0044] Decoding the basic latent feature image using an image decoding method to obtain the target image.

[0045] In some embodiments, decoding the basic latent feature image using an image decoding method to obtain the target image includes:

[0046] When the image style of the target image is a realistic style, encoding the final text to obtain a refined text embedding vector;

[0047] Embedding the refined text embedding vector into an image refinement model and using the image refinement model embedded with the refined text embedding vector to perform refined noise reduction processing on the basic latent feature image to obtain a refined latent feature image;

[0048] Decode the refined latent feature image to obtain the target image.

[0049] In some embodiments, the using the image generation model to perform N denoising processes on the noise image according to the single-sampling algorithm includes:

[0050] For the i-th denoising process among the N denoising processes, obtain an intermediate noise image and obtain the noise standard deviation of the intermediate noise image, where the intermediate noise image is the latent space image obtained by subjecting the noise image to (i - 1) denoising processes, and i is any positive integer less than or equal to N;

[0051] Update the parameters of the single-sampling algorithm using the noise standard deviation;

[0052] Use the image generation model embedding the basic text embedding vector to perform a single denoising process on the intermediate noise image according to the single-sampling algorithm with updated parameters until N denoising processes are completed.

[0053] In some embodiments, the obtaining the noise standard deviation of the intermediate noise image includes:

[0054] Obtain the hyperparameters of the sampler, as well as the minimum noise standard deviation and the maximum noise standard deviation of the sampler;

[0055] When i is less than N, input the hyperparameters, the minimum noise standard deviation, the maximum noise standard deviation, and i into the standard deviation calculation formula to obtain the noise standard deviation;

[0056] When i is equal to N, determine the noise standard deviation as 0.

[0057] In a second aspect, an embodiment of the present application provides an image generation device, the device includes:

[0058] A first acquisition module, configured to acquire a prompt word input by a user, where the prompt word is used to indicate generating a target image;

[0059] A first determination module, configured to determine an image generation model and a sampler corresponding to the prompt word according to a preset first correspondence relationship, where the first correspondence relationship is the correspondence relationship among the prompt word, the image generation model, and the sampler, and the sampler is used to control the denoising method of the image generation model;

[0060] An expansion module, configured to expand the prompt word to obtain a final text;

[0061] An update module, configured to load the sampler into the image generation model to update the noise reduction parameters of the image generation model;

[0062] A conversion module, configured to convert the final text into the target image by using the image generation model with updated noise reduction parameters.

[0063] In a third aspect, an embodiment of the present application provides a text generation device, including: a processor and a memory storing computer program instructions;

[0064] When the processor executes the computer program instructions, the above image generation method is implemented.

[0065] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above image generation method is implemented.

[0066] In a fifth aspect, an embodiment of the present application provides a vehicle, where the vehicle includes computer program instructions, and when the computer program instructions are executed by a processor, the above image generation method is implemented.

[0067] In the present application, after receiving a prompt word input by a user, the prompt word can be expanded to obtain a final text, so as to enrich the user requirements expressed by the text content, and based on the prompt word, an image generation model and a sampler are selected, and different prompt words are converted into target images by using different samplers and models, and the target image is super-resolved to obtain a final target image. In this way, the text content of the prompt word can be expanded to obtain a final text, and a suitable model and sampler are selected based on the prompt word, and the final text is converted into a target image by using the selected sampler. In the above process, the accuracy and richness of the generated image can be improved by selecting the sampler and the model and expanding the text, so as to improve the quality of the generated image. Description of the Drawings

[0068] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced below. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0069] Figure 1 is a schematic flowchart of an image generation method provided by an embodiment of the present application;

[0070] Figure 2 is a schematic structural diagram of an image generation device provided by an embodiment of the present application;

[0071] Figure 3 It is a schematic diagram of the hardware structure of a text generation device provided by an embodiment of the present application. Specific embodiments

[0072] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the objectives, technical solutions and advantages of the present application more clear, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.

[0073] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the elements.

[0074] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The embodiments will be described in detail below in conjunction with the accompanying drawings.

[0075] Specifically, to solve the problems of the prior art, the embodiments of the present application provide an image generation method, apparatus, device, storage medium and vehicle. First, the image generation method provided by the embodiments of the present application will be introduced below.

[0076] Figure 1 A flowchart of an image generation method provided by an embodiment of the present application is shown. This method can be applied to the in-vehicle computer of a vehicle or a cloud server communicatively connected to the vehicle. The method includes the following steps:

[0077] S110, obtaining a prompt word input by a user, where the prompt word is used to indicate generating a target image.

[0078] In this embodiment, the image generation task can convert the prompt word input by the user into a target image, and the prompt word is used to describe the content of the target image that the user wants to generate. For example, the prompt word can be "generate a puppy" or "generate a big tree".

[0079] In addition, the input prompt words may include positive prompt words and negative prompt words. The positive prompt words mean the content that is expected to be included in the target image, while the negative prompt words mean the content that is not expected to be included in the target image. For example, the positive prompt words may be "a cartoon-style Chinese girl running on the beach, with flying seagulls and a gorgeous rainbow behind her, and the overall picture is poetic and picturesque", while the negative prompt words may be "pornographic, nude, ugly, deformed".

[0080] S120, determining the image generation model and sampler corresponding to the prompt word according to a pre-set first correspondence relationship, wherein the first correspondence relationship is a correspondence relationship between the prompt word, the image generation model and the sampler, and the sampler is used to control the denoising method of the image generation model.

[0081] In this embodiment, the prompt word may be input into the image generation model, and the image generation model loaded with the sampler converts the prompt word into an image.

[0082] In this embodiment, the image generation model may include a basic image generation model and a specific object generation model. The basic image generation model may be a base model in a stable diffusion image generation model, which may convert text content to generate a basic image, and the specific object generation model may be a low-rank adaptation of large language models (LoRA image generation model) in a potential diffusion image generation model, which may be used to generate images of certain specific styles or specific subjects. The specific object generation model cannot be used alone, and the specific object generation model needs to be mounted on the basic image generation model to assist the basic image generation model in realizing image generation.

[0083] In this case, both the basic image generation model and the specific object generation model can be models of the UNET structure. During the training process of the basic image generation model and the specific object generation model, the training set can be multiple clear pictures. The training process of each picture mainly consists of two parts, namely, noise addition and denoising. During the noise addition process, noise can be continuously added to a clear initial picture. The denoising process is to use the basic image generation model and the specific object generation model to infer the picture with added noise, predict the noise and eliminate it until it is restored to the restored image. Then, by comparing the initial image and the restored image, the loss function of the model is calculated and adjusted to complete the training of the model.

[0084] Specifically, the prompt words can be used to guide the image generation model to perform multiple iterative denoising processes on the random noise image to obtain the final target image, and the sampler is used to load into the image generation model during the iterative process of repeated denoising of the image generation model to control the intensity of denoising, denoising method and number of iterations.

[0085] The correspondence between the prompt word, the image generation model and the sampler can be defined in advance as the first correspondence, and then whenever a prompt word input by the user is received, the image generation model and the sampler corresponding to the prompt word can be queried based on the first correspondence.

[0086] For example, the first correspondence may include a correspondence between a prompt word and an image generation model, and a correspondence between the image generation model and a sampler.

[0087] S130, performing text expansion on the prompt word to obtain a final text.

[0088] In this embodiment, the prompt word can be a simple sentence output by the user in the form of voice or text. In order to enrich the content of the image converted by the prompt word, the prompt word can be expanded to obtain a final text with richer content.

[0089] Specifically, phrases or expressions in a pre-set descriptive word library can be added to the prompt words, and the prompt words can be described in detail and with extension; for example, when the positive prompt word is "a round-faced, chubby and cute tabby cat with an anime style is basking in the sun on a lawn full of flowers", it can be expanded into the text to obtain the final text "a round-faced, chubby and cute tabby cat with an anime style is basking in the sun on a lawn full of flowers, honest and cute, with big eyes, charming scenery, and professional photography techniques".

[0090] S140, loading the sampler into the image generation model to update the denoising parameters of the image generation model.

[0091] In this embodiment, the sampler is loaded into the image generation model. The sampler will adjust the prediction method of noise, the intensity of noise reduction, the noise reduction method, and the number of iterations of the image generation model on the input noise image by affecting the noise reduction parameters in the image generation model, thereby affecting the image generated by the image generation model.

[0092] S150, using the image generation model updated with the noise reduction parameters to convert the final text into the target image.

[0093] In this embodiment, after selecting the sampler and the image generation model corresponding to the prompt, and loading the determined sampler into the image generation model, a random noise image can be obtained, and the prompt is converted into a text embedding vector, and the text embedding vector is embedded into the image generation model loaded with the sampler to guide the image generation model to perform multiple iterations of noise reduction on the noise image to obtain the target image.

[0094] In the embodiment of the present application, after receiving the prompt input by the user, the text of the prompt can be expanded to obtain the final text, so as to enrich the user needs expressed by the text content, and based on the prompt, the image generation model and the sampler are selected, and different prompts are converted into target images by using different samplers and models, and the target images are super-resolved to obtain the final target image. In this way, the text content of the prompt can be expanded to obtain the final text, and the appropriate model and sampler are selected based on the prompt, and the final text is converted into the target image by using the selected sampler. In the above process, the accuracy and richness of the generated image can be improved by the selection of the sampler and the model and the expansion of the text, thereby improving the quality of the generated image.

[0095] As an optional embodiment, the image generation model includes a basic image generation model and a specific object generation model, and the first corresponding relationship includes a model corresponding relationship and a sampler corresponding relationship. The above S120 may include:

[0096] Perform semantic recognition on the prompt to determine the image generation intention corresponding to the prompt;

[0097] Identify the specific generation information included in the image generation intention corresponding to the prompt; the specific generation information includes a specific subject and / or a specific style;

[0098] According to the model corresponding relationship, determine the basic image generation model and the specific object generation model corresponding to the prompt. The model corresponding relationship includes the corresponding relationship between the specific generation information and the specific object generation model, and the corresponding relationship between the image generation intention and the basic image generation model;

[0099] Determine the samplers corresponding to the prompt, the basic image generation model, and the specific object generation model according to the sampler correspondence relationship, where the sampler correspondence relationship includes the correspondence relationship among the image generation intention, the specific generation information, the basic image generation model, the specific object generation model, and the sampler.

[0100] In this embodiment, the intention expressed by the prompt can be recognized by semantic understanding of the prompt. Since the user wants to generate a target image based on the prompt, the image generation intention includes the attributes of the target image to be generated and the elements included in the target image. Among them, the specific generation information included in the image generation intention is the information of the specific subject or specific style included in the target image.

[0101] Exemplarily, the specific subject may refer to certain special subjects that the user hopes to generate in the image. For example, the specific subject may be a cat, a car, etc. The specific style indicates that the user hopes that the generated image has a certain special style, such as the oil painting style, the watercolor style, etc.

[0102] The model correspondence relationship is the correspondence relationship between the prompt and the image generation model. After determining the image generation intention and the specific generation information corresponding to the prompt, the basic image generation model corresponding to the image generation intention can be queried based on the model correspondence relationship, and the specific object generation model corresponding to the specific generation information can be queried, so as to determine the image generation model.

[0103] The sampler correspondence relationship is the correspondence relationship among the prompt, the image generation model, and the sampler. The sampler corresponding to the image generation intention, the specific generation information, the basic image generation model, and the specific object generation model can be queried based on the sampler correspondence relationship, so as to determine the sampler.

[0104] Exemplarily, there are multiple different basic image generation models and multiple different specific object generation models. Among them, each basic image generation model is used to implement the text-to-image task of one or more subjects in a vertical domain, and each specific object generation model is used to implement the text-to-image task of one or more image styles. Therefore, the vertical domain to which the subject of the prompt belongs can be determined through the image generation intention, and the basic image generation model corresponding to the vertical domain can be determined based on the model correspondence relationship. Then, the image style of the target image can be determined through the specific generation information, and the specific object generation model corresponding to the image style can be determined based on the model correspondence relationship. The sampler can be jointly determined by the subject type, the image style, and the selected basic image generation model and specific object generation model.

[0105] In this way, based on the user intention reflected by the prompt, an appropriate image generation model and sampler can be accurately selected to improve the accuracy of text conversion.

[0106] As an optional embodiment, determining the sampler corresponding to the prompt, the basic image generation model, and the specific object generation model according to the sampler correspondence includes:

[0107] Obtain the image generation intention, the specific generation information, the basic image generation model, and the identification information of the specific object generation model;

[0108] Determine the identification information as the initialization parameter of the sampler;

[0109] Match the sampler corresponding to the initialization parameter from the sampler correspondence, where the sampler correspondence includes the correspondence between the initialization parameter and the sampler.

[0110] In this embodiment, the image generation intention may include the subject type, and the specific generation information may include the specific style. A subject type label set S P = [p cls1 , p cls2 … p clsn can be defined in advance. The elements in this subject type label set can all be used to identify the subject type in the image generation intention. The subject type label may include animals, plants, people, food, etc. A specific style label set S t = [t cls1 , t cls2 … t clsn can also be defined in advance. The elements in this specific style label set can all be used to identify the image style in the specific generation information, and the image style of the target image can also be any element in this specific style label set. That is to say, the subject type label is the identification information of the image generation intention, and the specific style label is the identification information of the specific generation information.

[0111] In addition, the main body and main body type of the prompt can be obtained by extracting the features of the prompt and then processing the extracted features through common neural network structures such as pooling, convolution, and residual. Both the model correspondence and the sampler correspondence can be mapping relationships. Through a predefined mapping table, query the basic image generation model number and the specific object generation model number that have a mapping relationship with the main body type label and the image style label. Among them, the basic image generation model number is the identification information of the basic image generation model, and the specific object generation model number is the identification information of the specific object generation model. Specifically, multiple different basic image generation models and multiple different specific object generation models can be set. Each basic image generation model has a uniquely corresponding basic image generation model number, and each specific object generation model has a uniquely corresponding specific object generation model number. The mapping relationship includes the mapping relationship between the main body type label and the basic image generation model number, and the mapping relationship between the image style label and the specific object generation model number.

[0112] After the query is completed, the image generation intention, specific generation information, basic image generation model, and the identification information of the specific object generation model can be determined as the initialization parameters of the sampler, and then a new set of initialization parameters including the image generation intention, specific generation information, basic image generation model, and the identification information of the specific object generation model is generated:

[0113] [Main body type label, specific style label, basic image generation model number, specific object generation model number]

[0114] Among them, the four elements in the set of initialization parameters are all input parameters required for sampler initialization, and then based on the sampler correspondence, query the sampler that has a mapping relationship with this set of initialization parameters in the predefined mapping table.

[0115] For example, for the input with the set of initialization parameters input1 = [person, realistic, person model, person LORA], the DPM++ sampler that has a mapping relationship with the set of initialization parameters of this sampler can be queried based on the sampler correspondence. The characteristics of the DPM++ sampler are good denoising quality, balanced efficiency, stable denoising effect, and traceability for the same initialization noise. Due to the realistic characteristics, we take the denoising iteration times timesteps = T1. For the input with the set of initialization parameters input2 = [person, anime, anime model, anime LORA], the Euler sampler that has a mapping relationship with the set of initialization parameters of this sampler can be queried based on the sampler correspondence. The characteristics of the Euler sampler are very fast iteration speed and good denoising effect can be achieved in a short number of iterations, and its denoising iteration times timesteps = T2. Among them, T2 is less than T1.

[0116] In the above manner, the image generation model and sampler matching the text-to-image task can be accurately and conveniently determined through the prompt words.

[0117] As an optional embodiment, the above S130 may include:

[0118] Obtain the model keyword corresponding to the image generation model, add the model keyword to the prompt words to obtain a first intermediate text;

[0119] In the case where there is a special phrase representing a special object in the prompt words, obtain the object keyword corresponding to the special object, and add the object keyword to the first intermediate text to obtain a second intermediate text;

[0120] Convert the second intermediate text into the final text.

[0121] In this embodiment, the image generation model may include a basic image generation model and a specific object generation model. Each basic image generation model is used to implement the text-to-image task of one or more main bodies in a vertical domain, and each specific object generation model is used to implement the text-to-image task of one or more image styles.

[0122] For the image generation model, a model keyword mapping table associated with the image generation model can be predefined. In the model keyword mapping table, each image generation model has a model keyword with a mapping relationship with it. After determining the image generation model, the model keyword having a mapping relationship with the determined image generation model can be retrieved from the model keyword mapping table, and these model keywords are added to the prompt words to obtain a first intermediate text.

[0123] Exemplarily, since each basic image generation model corresponds to at least one vertical domain and each specific object generation model corresponds to one image style, the model keyword having a mapping relationship with the basic image generation model is used to describe the vertical domain corresponding to the basic image generation model, and the model keyword having a mapping relationship with the specific object generation model is used to describe the image style corresponding to the specific object generation model.

[0124] In addition, the special object may be a description object that the basic image generation model cannot understand or does not support, such as a certain special person, logo or address, etc. The specific object generation model can be trained by extracting, so that the specific object generation model can understand these special objects, and an object keyword mapping table associated with the special objects is generated. In the object keyword mapping table, each special object has an object keyword with a mapping relationship.

[0125] In the case where a special phrase representing a special object is detected in the prompt, an object keyword having a mapping relationship with the special object can be queried from the object keyword mapping table, and the object keyword is added to the first intermediate text to obtain a second intermediate text, and then the second intermediate text is further converted into a final text.

[0126] By adding the keyword corresponding to the model and the keyword corresponding to the special object to the prompt in the embodiment of the present application, the text content of the prompt is made more abundant, and it is convenient for the model to better understand the text content input into the model, thereby improving the accuracy of the generated image.

[0127] As an optional embodiment, the converting the second intermediate text into the final text includes:

[0128] In the case where an informal phrase is detected in the second intermediate text, a common phrase corresponding to the informal phrase is queried according to a preset second correspondence, and the second correspondence is a correspondence between the informal phrase and the common phrase;

[0129] The corresponding informal phrase in the second intermediate text is replaced with the common phrase to obtain the final text.

[0130] In this embodiment, the informal phrases include regional slang, common sayings, ancient Chinese poems, proverbs, and Internet terms, etc. These informal phrases are usually commonly used language expression forms in a certain region or community, and usually have certain regionality and limitations. The image generation model usually cannot understand these informal phrases.

[0131] The second correspondence can be a mapping relationship. An informal phrase mapping table can be preset. In the informal phrase mapping table, each informal phrase has a corresponding common phrase with a mapping relationship, and the common phrase is an easy-to-understand and popular explanation of the informal phrase.

[0132] Therefore, whenever an informal phrase is detected in the second intermediate text, a common phrase having a mapping relationship with the informal phrase can be queried from the informal phrase mapping table, and then the corresponding informal phrase in the second intermediate text is replaced with the common phrase to obtain the final text.

[0133] Exemplarily, when the second intermediate text is "A cute tiger tabby cat with an anime style and a round face and chubby body is sunbathing on a lawn full of flowers", where "round face and chubby body" is a Chinese slang and common saying, that is, an informal phrase. Therefore, a common phrase "round face, chubby" having a mapping relationship with "round face and chubby body" can be queried, so that the second intermediate text can be converted into the final text "A cute tiger tabby cat with an anime style, round face, and chubby body is sunbathing on a lawn full of flowers".

[0134] In the embodiments of the present application, by converting the obscure and informal phrases in the text content to be input into the model, it is convenient for the model to better understand the text content input into the model, thereby improving the accuracy of the generated images.

[0135] As an optional embodiment, the image generation model includes a basic image generation model and a specific object generation model. After determining the image generation model and sampler corresponding to the prompt word according to the preset first correspondence relationship, the method further includes:

[0136] Determine a specific object module in the basic image generation model that matches the function of the specific object generation model;

[0137] Obtain the input data of the specific object module, and input the input data into the specific object generation model;

[0138] Obtain the module output data obtained by the specific object module in response to the input data, and obtain the model output data obtained by the specific object generation model in response to the input data;

[0139] Merge the module output data and the model output data to obtain fusion data;

[0140] Update the fusion data to the model parameters corresponding to the specific object in the basic image generation model, so as to mount the specific object generation model to the basic image generation model.

[0141] In this embodiment, the specific object generation model can, with relatively low power consumption, assist the basic image generation model in accurately understanding and generating images of special objects that the basic image generation model cannot understand or is difficult to support, or certain specific image styles. The specific object generation model cannot run independently, so it is necessary to mount the specific object generation model to the basic image generation model to obtain the final image generation model.

[0142] Specifically, after selecting the specific object generation model, a specific object module that matches the function of the specific object generation model can be determined in the basic image generation model, and then the specific object generation model can be connected to the basic image generation model according to the connection method of the specific object module.

[0143] Specifically, the input data originally intended to be input into a specific object module can be input into the specific object module and the specific object generation model simultaneously. After that, in response to the input data, the specific object module can obtain module output data; the specific object generation model can also obtain model output data in response to the input data. The module output data and the model output data can be added together to obtain fusion data, and the fusion data can be used as the model parameters corresponding to the specific object and input into the next layer of the specific object module in the basic image generation model, thereby completing the mounting of the specific object generation model.

[0144] In this embodiment, by using the specific object generation model to adjust the parameters of the basic image generation model, the image generation accuracy of the basic image generation model can be improved with relatively low power consumption.

[0145] As an alternative embodiment, obtaining the model output data obtained by the specific object generation model in response to the input data includes:

[0146] Obtaining the model weights in the specific object generation model;

[0147] Inputting the model weights into the specific object generation model to update the weight parameters of the specific object generation model;

[0148] Using the specific object generation model with updated weight parameters to obtain the model output data in response to the input data.

[0149] In this embodiment, during the model mounting process, the model parameters in the specific object generation model can be loaded by reading the weight file of the specific object generation model or directly using the model loading function provided in the programming language. The model parameters include model weights, and the model weights are the learnable parameters in the model used to adjust the influence of the input features.

[0150] Then, in the case of multiple different specific object generation models, these multiple specific object generation models can be respectively mounted on the basic image generation model to obtain multiple combined models, and then the multiple combined models are all exported in the ONNX (Open Neural Network Exchange) format. ONNX (Open Neural Network Exchange) is an open deep learning model representation standard aimed at improving the model interoperability across platforms and frameworks.

[0151] By traversing and comparing the weights of each node in each combined model, ONNX nodes with different weights in each combined model can be marked; since the weights of each node in the basic image generation model are the same, these nodes with different weights are the model weights of the specific object generation model in the combined model. In this way, the model weights corresponding to each specific object generation model can be obtained.

[0152] Subsequently, the model weights can be input into the specific object generation model corresponding to the model weights after mounting to update the weight parameters of the specific object generation model, thereby adjusting the response mode of the specific object generation model to the input data, and thus obtaining model output data in response to the input data.

[0153] Through the above method, the model weights loaded in real time can be converted into the input of the specific object generation model after mounting, thereby optimizing the performance of the model after mounting.

[0154] As an optional embodiment, the specific object generation model includes at least one specific object generation node, the model weights include at least one node weight, and each specific object generation node corresponds to a node weight. The inputting the model weights into the specific object generation model includes:

[0155] Adding at least one identity node to the specific object generation model, the output of the identity node being equal to the input of the identity node, and each identity node corresponding to a specific object generation node;

[0156] Inputting each node weight in the at least one node weight into the corresponding specific object generation node through the corresponding identity node.

[0157] In this embodiment, in the process of inputting the model weights into the specific object generation model, the model weights can be input by adding identity nodes to the specific object generation model and inputting the node weights of each specific object generation node into the specific object generation node through the identity nodes.

[0158] Among them, the identity node can be an Identity node. The Identity node is a special glue node, and the input of the node is equal to the output. Through the Identity node, a bypass path can be introduced to allow information to pass directly without additional transformation.

[0159] Through the above method, the model weights can be accurately and quickly introduced into the specific object generation model after mounting.

[0160] As an alternative embodiment, the sampler includes a single sampling algorithm and the number of iterations. The image generation model updated using the noise reduction parameter converts the final text into the target image, including:

[0161] Obtain a randomly generated noise image;

[0162] Encode the final text using a text vectorization encoding method to obtain a basic text embedding vector;

[0163] Embed the basic text embedding vector into the image generation model, and use the image generation model to perform N noise reduction processes on the noise image according to the single sampling algorithm to obtain a basic latent feature image, where N is the number of iterations;

[0164] Decode the basic latent feature image using an image decoding method to obtain the target image.

[0165] In this embodiment, a randomly generated noise image can be obtained, and the text content of the final text can be mapped to a high-dimensional vector space to obtain a basic text embedding vector. After obtaining the basic text embedding vector, the basic text embedding vector and the noise image can be jointly input into the image generation model, and the basic text embedding vector is used to guide the model to perform N iterative noise reduction processes on the noise image according to the single sampling algorithm to obtain a basic latent feature image. Then, the basic latent feature image can be decoded by using a variational autoencoder image generation model to obtain a target image that can be recognized by the human eye.

[0166] In this embodiment, the conversion from the final text to the target image can be accurately completed.

[0167] As an alternative embodiment, the decoding the basic latent feature image using an image decoding method to obtain the target image includes:

[0168] When the image style of the target image is a realistic style, encode the final text to obtain a refined text embedding vector;

[0169] Embed the refined text embedding vector into an image refinement model, and use the image refinement model embedded with the refined text embedding vector to perform refined noise reduction processing on the basic latent feature image to obtain a refined latent feature image;

[0170] Decode the refined latent feature image to obtain the target image.

[0171] In this embodiment, when the style of the target image is a realistic style, the target image has high requirements for image quality and details. Therefore, not only the basic semantics of the prompt need to be encoded into a basic text embedding vector, but also the detailed semantics of the prompt need to be encoded into a refined text embedding vector.

[0172] Since the basic latent feature image is only an image in the feature space that conforms to the basic features of the prompt, and the target image in the realistic style has high requirements for details, the basic latent feature image and the refined text embedding vector can be further input into the refinement model. The refined text embedding vector is used to guide the refinement model to perform further multiple iterative refinement and noise reduction processing on the basic latent feature image to obtain a refined latent feature image. The target image that can be recognized by the human eye can be obtained by decoding the refined latent feature image using the self-variational encoder model.

[0173] Among them, the refinement model can be the Refiner model in the Stable Diffusion model, which is used to refine images with insufficient details and enrich the details of the images. The refined latent feature image is an image in the feature space that conforms to the detailed features of the prompt.

[0174] In this embodiment, for realistic style images with high requirements for details, the basic image generation model and the refinement model can be cascaded twice. The basic image generation task is completed by the image generation model, and the image obtained by the basic model is refined by the refinement model to ensure the image generation quality.

[0175] As an alternative embodiment, decoding the basic latent feature image by means of image decoding to obtain the target image includes:

[0176] When the image style of the target image is an artistic style, decoding the basic latent feature image to obtain the target image.

[0177] In this embodiment, when the style of the target image is an artistic style, the target image focuses more on the display of the painting style and has low requirements for image details. Therefore, the basic latent feature image can be directly decoded to obtain the target image.

[0178] Among them, the basic model can be the base model in the Stable Diffusion model, which is used to convert text content into an image with insufficient details.

[0179] In this embodiment, for artistic style images with low detail requirements, the basic image generation task can be completed only through the basic model, thus effectively reducing the power consumption of algorithm inference while meeting the requirements.

[0180] As an alternative embodiment, the obtaining of the randomly generated noise image includes:

[0181] Obtaining a pre-set random number seed;

[0182] Converting the random number seed into a random sequence with a normal distribution form to obtain Gaussian noise;

[0183] Encoding the Gaussian noise to obtain the noise image.

[0184] In this embodiment, a random sequence with a Gaussian distribution can be generated in a deterministic manner by obtaining a pre-set random number seed (seed) to form Gaussian noise. After obtaining the Gaussian noise, the Gaussian noise can be mapped to the latent space through an encoder to obtain the latent Gaussian noise, that is, the noise image.

[0185] In this embodiment, the noise image generated each time the code is run can be ensured to be reproducible through a given random number seed.

[0186] As an alternative embodiment, the using of the image generation model to perform N denoising processes on the noise image according to the single-sampling algorithm includes:

[0187] For the i-th denoising process among the N denoising processes, obtaining an intermediate noise image and obtaining the noise standard deviation of the intermediate noise image, where the intermediate noise image is the latent space image obtained by subjecting the noise image to (i - 1) denoising processes, and i is any positive integer less than or equal to N;

[0188] Updating the parameters of the single-sampling algorithm using the noise standard deviation;

[0189] Using the image generation model embedded with the basic text embedding vector to perform a single denoising process on the intermediate noise image according to the single-sampling algorithm with updated parameters until N denoising processes are completed.

[0190] In this embodiment, before applying the sampler, the construction of the sampler needs to be completed first. During the construction of the sampler, the sampler needs to be initialized first, and then the single-sampling algorithm and the number of iterations of the sampler are determined. Among them, the single-sampling algorithm determines the way and intensity of noise removal in each denoising process.

[0191] Specifically, the noise standard deviation of the target image is a model parameter in the determination process of each single-sampling algorithm. The model parameter is a parameter learned during the execution of the algorithm and will be updated in each round of iteration. Therefore, it is necessary to obtain the noise standard deviation of the target image before each noise reduction process is executed, then update the single-sampling algorithm using the noise standard deviation, and perform single noise reduction processing on the input target image through the updated single-sampling algorithm.

[0192] Specifically, for the i-th noise reduction process among N noise reduction processes, its noise reduction formula is as follows:

[0193] D θ (x; σ) = c skip (σ)x + c out (σ)F θ (c in (σ)x; c noise (σ))

[0194] Among them, x represents the latent space image obtained by subjecting the noise image to (i - 1) noise reduction processes, that is, the target image. If i is 1, then x is the noise image, and D θ is the latent space image after the i-th noise reduction process. σ is used to represent the standard deviation of the noise of the latent space image obtained by (i - 1) noise reduction processes. F θ represents the inference of the noise reduction neural network algorithm. And Cskip(σ), Cout(σ), Cin(σ), and Cnoise(σ) represent the four noise reduction process terms in the noise reduction formula respectively. Specifically, Cskip(σ) represents the skip connection scaling process, Cout(σ) represents the output scaling process, Cin(σ) represents the input scaling process, and Cnoise(σ) represents the noise adjustment process.

[0195] In different samplers, the above four noise reduction process terms are not the same. For example, in the DPM sampler:

[0196] Skip scaling c skip (σ)1

[0197] Output scaling c out (σ)-σ

[0198]

[0199] Noise cond.c noise (σ)(M - 1)σ -1 (σ)

[0200] In the DDIM sampler:

[0201] Skip scaling c skip (σ)1

[0202] Output scaling c out (σ)-σ

[0203]

[0204] Noise cond.c noise (σ)M - 1 - arg min j |u j -σ|

[0205] Wherein, M is the initial value of the random noise, and μj is the noise adjustment parameter for the jth time.

[0206] In this way, when the type of the sampler is determined, the sampler can be constructed according to the noise reduction process and the number of iterations specified by this type of sampler. After the sampler is constructed, for each iteration process of noise reduction, the latent space image obtained from the previous noise reduction process and the noise standard deviation of this latent space image are input into the sampler, and a new round of noise reduction can be completed through the sampler.

[0207] As an optional embodiment, obtaining the noise standard deviation of the intermediate noise image includes:

[0208] Obtaining the hyperparameters of the sampler, as well as the minimum value and the maximum value of the noise standard deviation of the sampler;

[0209] When i is less than N, according to the hyperparameters, the minimum value of the noise standard deviation, the maximum value of the noise standard deviation, and i are input into the standard deviation calculation formula to obtain the noise standard deviation;

[0210] When i is equal to N, the noise standard deviation is determined to be 0.

[0211] In this embodiment, when the sampler is constructed and noise reduction is achieved using the sampler, it is necessary to obtain the noise standard deviation of the target image once in each noise reduction process, and input the noise standard deviation together with the target image into the sampler.

[0212] Specifically, in the process of the ith noise reduction process, the hyperparameters of the sampler, as well as the previously specified minimum value and the maximum value of the noise standard deviation can be obtained first. Among them, the hyperparameters are used to control the step size of the variance transformation.

[0213] When i is less than N, the noise standard deviation in the ith iteration process can be calculated according to the following standard deviation calculation formula:

[0214]

[0215] Among them, i represents the current iteration number, N represents the total number of iterations, and σ min and σ max represent the minimum value and the maximum value of the noise standard deviation respectively, and ρ is a hyperparameter.

[0216] Exemplarily, the hyperparameter can be 7, and σ min and σ max are 0.02 and 100 respectively.

[0217] When i is equal to N, σ N is 0.

[0218] In this way, the noise standard deviation of the sampler can be adjusted according to a predetermined step size, avoiding the change of the noise standard deviation being too small or too large.

[0219] As an alternative embodiment, after the above S150, it may further include:

[0220] , perform super-resolution processing on the target image to obtain a target image.

[0221] In this embodiment, since the operation of the image generation model consumes a large amount of computing power, the resolution of the target image output by the image generation model is usually low. The target image can be super-resolved by a trained super-resolution model to improve the overall resolution of the target image and obtain a clearer target image.

[0222] Exemplarily, the super-resolution model can be an Enhanced Super-Resolution Generative Adversarial Network (ESRGAN), and use ESRGAN to perform 2X super-resolution on the target image. Specifically, ESRGAN can use deep learning technology and generative adversarial network to improve the spatial resolution of the target image. And 2X super-resolution can generate an image with a resolution twice that of the target image in both the horizontal and vertical directions by processing the target image, thereby improving the visual quality and details of the image.

[0223] As an optional embodiment, the super-resolution model can be a Residual in Residual Dense Block (RRDB) module without Batch Normalization (BN). Exemplarily, the RRDB module may include three interconnected Dense Blocks, and each Dense Block consists of five convolutional layers. These convolutional layers may have different filters and feature map depths for learning different levels of representations of the image. In this way, the super-resolution model has better generalization ability.

[0224] Based on the image generation method provided in the above embodiment, correspondingly, the present application also provides a specific implementation manner of the image generation device. Please refer to the following embodiments.

[0225] First, refer to Figure 2 , the image generation device 200 provided in the embodiment of the present application includes the following modules:

[0226] The first acquisition module 201 is configured to acquire a prompt word input by a user, where the prompt word is used to indicate the generation of a target image;

[0227] The first determination module 202 is configured to determine an image generation model and a sampler corresponding to the prompt word according to a preset first correspondence, where the first correspondence is a correspondence among the prompt word, the image generation model, and the sampler, and the sampler is used to control the noise reduction method of the image generation model;

[0228] The expansion module 203 is configured to expand the prompt word to obtain a final text;

[0229] The update module 204 is configured to load the sampler into the image generation model to update the noise reduction parameters of the image generation model;

[0230] The conversion module 205 is configured to convert the final text into the target image by using the image generation model with updated noise reduction parameters.

[0231] After receiving the prompt word input by the user, the device can expand the text of the prompt word to obtain the final text, so as to enrich the user requirements expressed by the text content, and select an image generation model and a sampler based on the prompt word, convert different prompt words into target images using different samplers and models, and perform super-resolution on the target images to obtain the final target images. In this way, by expanding the text content of the prompt word, the final text can be obtained, and a suitable model and sampler can be selected based on the prompt word, and the final text can be converted into a target image using the selected sampler. In the above process, the accuracy and richness of the generated images can be improved by selecting the sampler and model and expanding the text, thereby improving the quality of the generated images.

[0232] As an implementation manner of the present application, the image generation model includes a basic image generation model and a specific object generation model, the first correspondence includes a model correspondence and a sampler correspondence, and the above first determination module 202 may further include:

[0233] A first recognition unit, configured to perform semantic recognition on the prompt word to determine the image generation intention corresponding to the prompt word;

[0234] A second recognition unit, configured to recognize the specific generation information included in the image generation intention corresponding to the prompt word; the specific generation information includes a specific subject and / or a specific style;

[0235] A first determination unit, configured to determine the basic image generation model and the specific object generation model corresponding to the prompt word according to the model correspondence, where the model correspondence includes the correspondence between the specific generation information and the specific object generation model, and the correspondence between the image generation intention and the basic image generation model;

[0236] A second determination unit, configured to determine the sampler corresponding to the prompt word, the basic image generation model, and the specific object generation model according to the sampler correspondence, where the sampler correspondence includes the correspondence among the image generation intention, the specific generation information, the basic image generation model, the specific object generation model, and the sampler.

[0237] As an implementation manner of the present application, the above second determination unit may further include:

[0238] A first acquisition subunit, configured to acquire the identification information of the image generation intention, the specific generation information, the basic image generation model, and the specific object generation model;

[0239] A first determination subunit, configured to determine the identification information as the initialization parameter of the sampler;

[0240] The first matching subunit is configured to match the sampler corresponding to the initialization parameter from the sampler correspondence, where the sampler correspondence includes the correspondence between the initialization parameter and the sampler.

[0241] As an implementation manner of this application, the above expansion module 203 may further include:

[0242] The first adding unit is configured to obtain the model keyword corresponding to the image generation model, add the model keyword to the prompt word to obtain a first intermediate text;

[0243] The second adding unit is configured to, when there is a special phrase representing a special object in the prompt word, obtain the object keyword corresponding to the special object, and add the object keyword to the first intermediate text to obtain a second intermediate text;

[0244] The conversion unit is configured to convert the second intermediate text into the final text.

[0245] As an implementation manner of this application, the above conversion unit may further include:

[0246] The first query subunit is configured to, when it is detected that there is an informal phrase in the second intermediate text, query the popular phrase corresponding to the informal phrase according to a preset second correspondence, where the second correspondence is the correspondence between the informal phrase and the popular phrase;

[0247] The replacement subunit is configured to replace the corresponding informal phrase in the second intermediate text with the popular phrase to obtain the final text.

[0248] As an implementation manner of this application, the above image generation device 200 may further include:

[0249] The second determination module is configured to determine the specific object module in the basic image generation model that matches the function of the specific object generation model;

[0250] The second obtaining module is configured to obtain the input data of the specific object module, and input the input data into the specific object generation model;

[0251] The third obtaining module is configured to obtain the module output data obtained by the specific object module in response to the input data, and obtain the model output data obtained by the specific object generation model in response to the input data;

[0252] The merging module is configured to merge the module output data and the model output data to obtain the fusion data;

[0253] A mounting module, configured to update the fusion data into model parameters corresponding to the specific object in the basic image generation model, so as to mount the specific object generation model into the basic image generation model.

[0254] As an implementation manner of this application, the above-mentioned third acquisition module may further include:

[0255] A first acquisition unit, configured to acquire model weights in the specific object generation model;

[0256] An input unit, configured to input the model weights into the specific object generation model to update weight parameters of the specific object generation model;

[0257] A response unit, configured to use the specific object generation model with updated weight parameters to respond to the input data to obtain the model output data.

[0258] As an implementation manner of this application, the above-mentioned input unit may further include:

[0259] An adding subunit, configured to add at least one identity node in the specific object generation model, an output of the identity node is equal to an input of the identity node, and each identity node corresponds to one specific object generation node;

[0260] An input subunit, configured to input respective node weights in the at least one node weight into the corresponding specific object generation node through the corresponding identity node.

[0261] As an implementation manner of this application, the above-mentioned conversion module 205 may further include:

[0262] A first acquisition unit, configured to acquire a randomly generated noise image;

[0263] A first encoding unit, configured to encode the final text by using a text vectorization encoding method to obtain a basic text embedding vector;

[0264] A first noise reduction unit, configured to embed the basic text embedding vector into the image generation model, and use the image generation model to perform N times of noise reduction processing on the noise image according to the single sampling algorithm to obtain a basic latent feature image, where N is the number of iterations;

[0265] A first decoding unit, configured to decode the basic latent feature image by using an image decoding method to obtain a target image.

[0266] As an implementation manner of this application, the above-mentioned first decoding unit may further include:

[0267] An encoding subunit, configured to encode the final text to obtain a refined text embedding vector when the image style of the target image is a realistic style;

[0268] A refinement subunit, configured to embed the refined text embedding vector into an image refinement model, and perform refinement and noise reduction processing on the basic latent feature image by using the image refinement model embedded with the refined text embedding vector to obtain a refined latent feature image;

[0269] A decoding subunit, configured to decode the refined latent feature image to obtain the target image.

[0270] As an implementation manner of the present application, the above-mentioned first noise reduction unit may also be used for:

[0271] For the i-th noise reduction process in the N times of noise reduction processing, obtain an intermediate noise image and obtain the noise standard deviation of the intermediate noise image, where the intermediate noise image is a latent space image obtained by subjecting the noise image to (i - 1) times of noise reduction processing, and i is any positive integer less than or equal to N;

[0272] Update the parameters of the single-sampling algorithm by using the noise standard deviation;

[0273] Perform single noise reduction processing on the intermediate noise image by using the image generation model embedded with the basic text embedding vector according to the single-sampling algorithm with updated parameters until N times of noise reduction processing are completed.

[0274] As an implementation manner of the present application, the above-mentioned first noise reduction unit may also be used for:

[0275] Obtain the hyperparameters of the sampler, as well as the minimum noise standard deviation and the maximum noise standard deviation of the sampler;

[0276] When i is less than N, input the hyperparameters, the minimum noise standard deviation, the maximum noise standard deviation, and i into a standard deviation calculation formula to obtain the noise standard deviation;

[0277] When i is equal to N, determine the noise standard deviation as 0.

[0278] The image generation device provided by the embodiments of the present invention can implement each step in the above method embodiments. To avoid repetition, it will not be elaborated here.

[0279] Figure 3 The hardware structure diagram of the text generation device provided by the embodiments of the present application is shown.

[0280] The text generation device may include a processor 301 and a memory 302 storing computer program instructions.

[0281] Specifically, the above-mentioned processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits implementing the embodiments of the present application.

[0282] The memory 302 may include a mass storage for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 302 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid state memory.

[0283] The memory may include a read only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of the present disclosure.

[0284] The processor 301 reads and executes the computer program instructions stored in the memory 302 to implement any one of the image generation methods in the above embodiments.

[0285] In one example, the text generation device may further include a communication interface 303 and a bus 310. Among them, as Figure 3 shown, the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 and complete communication with each other.

[0286] The communication interface 303 is mainly used to implement communication between various modules, devices, units, and / or devices in the embodiments of the present application.

[0287] The bus 310 includes hardware, software, or both, and couples the components of the text generation device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 310 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0288] The text generation device may be based on the above embodiments, so as to implement the image generation method and device in combination with the above.

[0289] In addition, in combination with the image generation method in the above embodiments, the embodiments of the present application may provide a computer storage medium for implementation. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the image generation methods in the above embodiments is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be described in detail here. Among them, the above computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc., which is not limited herein.

[0290] In addition, the embodiments of the present application also provide a vehicle, including computer program instructions, which can implement the steps and corresponding contents of the foregoing method embodiments when executed by a processor.

[0291] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0292] The functional blocks shown in the above structural block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0293] It should also be noted that in the exemplary embodiments mentioned in the present application, some methods or systems are described based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, can be different from the order in the embodiments, or several steps can be executed simultaneously.

[0294] As described above with reference to the flowcharts and / or block diagrams of methods, apparatuses, and vehicles according to embodiments of the present disclosure. It should be understood that each block in the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It is also understood that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0295] The above is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. An image generation method, characterized in that, The method includes: Obtaining a prompt word input by a user, where the prompt word is used to indicate the generation of a target image; Determining an image generation model and a sampler corresponding to the prompt word according to a preset first correspondence relationship, where the first correspondence relationship is the correspondence relationship among the prompt word, the image generation model, and the sampler, and the sampler is used to control the noise reduction method of the image generation model; Performing text expansion on the prompt word to obtain a final text; Loading the sampler into the image generation model to update the noise reduction parameters of the image generation model; Using the image generation model with updated noise reduction parameters to convert the final text into the target image.

2. The image generation method according to claim 1, characterized in that, The image generation model includes a basic image generation model and a specific object generation model, and the first correspondence relationship includes a model correspondence relationship and a sampler correspondence relationship. The determining of the image generation model and the sampler corresponding to the prompt word according to the preset first correspondence relationship includes: Performing semantic recognition on the prompt word to determine the image generation intention corresponding to the prompt word; Identifying specific generation information included in the image generation intention corresponding to the prompt word; the specific generation information includes a specific subject and / or a specific style; Determining the basic image generation model and the specific object generation model corresponding to the prompt word according to the model correspondence relationship, where the model correspondence relationship includes the correspondence relationship between the specific generation information and the specific object generation model, and the correspondence relationship between the image generation intention and the basic image generation model; Determining the sampler corresponding to the prompt word, the basic image generation model, and the specific object generation model according to the sampler correspondence relationship, where the sampler correspondence relationship includes the correspondence relationship among the image generation intention, the specific generation information, the basic image generation model, the specific object generation model, and the sampler.

3. The image generation method according to claim 2, characterized in that, The determining of the sampler corresponding to the prompt word, the basic image generation model, and the specific object generation model according to the sampler correspondence relationship includes: Obtaining the identification information of the image generation intention, the specific generation information, the basic image generation model, and the specific object generation model; Determining the identification information as the initialization parameter of the sampler; Matching the sampler corresponding to the initialization parameter from the sampler correspondence relationship, where the sampler correspondence relationship includes the correspondence relationship between the initialization parameter and the sampler.

4. The image generation method according to claim 1, characterized in that, The performing of text expansion on the prompt word to obtain a final text includes: Obtaining the model keyword corresponding to the image generation model, adding the model keyword to the prompt word to obtain a first intermediate text; In the case where there is a special phrase representing a special object in the prompt word, obtaining the object keyword corresponding to the special object and adding the object keyword to the first intermediate text to obtain a second intermediate text; Converting the second intermediate text into the final text.

5. The image generation method according to claim 4, characterized in that, The converting of the second intermediate text into the final text includes: In the case where an informal phrase is detected in the second intermediate text, query the common phrase corresponding to the informal phrase according to a preset second correspondence, where the second correspondence is the correspondence between the informal phrase and the common phrase; Replace the corresponding informal phrase in the second intermediate text with the common phrase to obtain the final text.

6. The image generation method according to claim 1, characterized in that, The image generation model includes a basic image generation model and a specific object generation model. After determining the image generation model and sampler corresponding to the prompt according to a preset first correspondence, the method further includes: Determine a specific object module in the basic image generation model that matches the function of the specific object generation model; Obtain the input data of the specific object module and input the input data into the specific object generation model; Obtain the module output data obtained by the specific object module in response to the input data, and obtain the model output data obtained by the specific object generation model in response to the input data; Merge the module output data and the model output data to obtain fusion data; Update the model parameters corresponding to the specific object in the basic image generation model with the fusion data, so as to mount the specific object generation model to the basic image generation model.

7. The image generation method according to claim 6, characterized in that, The obtaining the model output data obtained by the specific object generation model in response to the input data includes: Obtain the model weights in the specific object generation model; Input the model weights into the specific object generation model to update the weight parameters of the specific object generation model; Use the specific object generation model with updated weight parameters to respond to the input data to obtain the model output data.

8. The image generation method according to claim 7, wherein The specific object generation model includes at least one specific object generation node, the model weights include at least one node weight, and each specific object generation node corresponds to a node weight. The inputting the model weights into the specific object generation model includes: Add at least one identity node in the specific object generation model, the output of the identity node is equal to the input of the identity node, and each identity node corresponds to a specific object generation node; Input each node weight in the at least one node weight into the corresponding specific object generation node through the corresponding identity node.

9. The image generation method according to claim 1, wherein The sampler includes a single sampling algorithm and the number of iterations. The converting the final text into the target image by using the image generation model updated with the noise reduction parameter includes: Obtain a randomly generated noise image; Encode the final text by using a text vectorization encoding method to obtain a basic text embedding vector; Embed the basic text embedding vector into the image generation model, and use the image generation model to perform N times of noise reduction processing on the noise image according to the single sampling algorithm to obtain a basic latent feature image, where N is the number of iterations; Decode the basic latent feature image by using an image decoding method to obtain the target image.

10. The image generation method according to claim 9, wherein Performing decoding on the basic latent feature image in an image decoding manner to obtain a target image includes: When the image style of the target image is a realistic style, encoding the final text to obtain a refined text embedding vector; Embedding the refined text embedding vector into an image refinement model, and using the image refinement model embedded with the refined text embedding vector to perform refinement and noise reduction processing on the basic latent feature image to obtain a refined latent feature image; Performing decoding on the refined latent feature image to obtain the target image.

11. The image generation method according to claim 9, wherein The performing N times of noise reduction processing on the noise image by using the image generation model according to the single-sampling algorithm includes: For the i-th noise reduction processing among the N times of noise reduction processing, obtaining an intermediate noise image and obtaining the noise standard deviation of the intermediate noise image, where the intermediate noise image is a latent space image obtained by subjecting the noise image to (i - 1) times of noise reduction processing, and i is any positive integer less than or equal to N; Updating the parameters of the single-sampling algorithm by using the noise standard deviation; Performing single-time noise reduction processing on the intermediate noise image by using the image generation model embedded with the basic text embedding vector according to the single-sampling algorithm with updated parameters until N times of noise reduction processing are completed.

12. The image generation method according to claim 11, wherein The obtaining the noise standard deviation of the intermediate noise image includes: Obtaining the hyperparameters of the sampler, as well as the minimum noise standard deviation and the maximum noise standard deviation of the sampler; When i is less than N, inputting the hyperparameters, the minimum noise standard deviation, the maximum noise standard deviation, and i into a standard deviation calculation formula to obtain the noise standard deviation; When i is equal to N, determining the noise standard deviation to be 0.

13. An image generation device, wherein The device includes: A first acquisition module, configured to acquire a prompt word input by a user, where the prompt word is used to indicate generating a target image; A first determination module, configured to determine an image generation model and a sampler corresponding to the prompt word according to a preset first correspondence relationship, where the first correspondence relationship is a correspondence relationship among the prompt word, the image generation model, and the sampler, and the sampler is used to control the noise reduction method of the image generation model; An expansion module, configured to perform text expansion on the prompt word to obtain a final text; An update module, configured to load the sampler into the image generation model to update the noise reduction parameters of the image generation model; A conversion module, configured to use the image generation model with updated noise reduction parameters to convert the final text into the target image.

14. A text generation device, wherein The text generation device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image generation method according to any one of claims 1-12 is implemented.

15. A computer storage medium, wherein Computer program instructions are stored on a computer storage medium, and when the computer program instructions are executed by a processor, the image generation method according to any one of claims 1-12 is implemented.

16. A vehicle, wherein The vehicle includes at least one of the following: the image generation device according to claim 13; An image generation device according to claim 14; A computer storage medium according to claim 15.