Image generation method and device, equipment, storage medium and vehicle

By analyzing the drawing instructions input by the user and querying the sampler database, and dynamically selecting the adapted sampler algorithm and parameters, the problem that the image generation model and sampler in the prior art cannot be adapted to different literary drawing tasks, improving the image generation quality.

CN120182973APending Publication Date: 2025-06-20BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311757404.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, fixed image generation models and samplers cannot adapt to the content generation requirements of different literary image tasks, resulting in lower quality of generated images.

Method used

By analyzing the drawing instructions input by the user, the corresponding target image content category is determined, and the adapted sampler algorithms and parameters are queried from the sampler database according to the content category, and the drawing instructions are converted into the target image using these algorithms and parameters.

Benefits of technology

It realizes dynamic selection of adaptive samplers based on the attributes of different literary and artistic tasks, avoiding the limitations of standardized processes and improving the quality of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182973A_ABST
    Figure CN120182973A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and device, equipment, a storage medium and a vehicle. The method comprises the steps of obtaining a drawing instruction input by a user, wherein the drawing instruction is used for indicating generation of a target image; analyzing the drawing instruction, and determining a content category of a target image corresponding to the drawing instruction; sampling device algorithms and parameters corresponding to the content category are determined from a sampling device database according to the content category, and the sampling device database comprises the corresponding relation between the content category and the sampling device algorithms and the parameters corresponding to each sampling device algorithm; and converting the drawing instruction into a target image by using a sampler algorithm and parameters. According to the embodiment of the invention, the quality of the image obtained by converting the drawing instruction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to an image generation method, apparatus, device, storage medium, and vehicle. Background Art

[0002] Text-to-image is a computer generation task aimed at converting text descriptions or natural language texts into corresponding images. In this task, a computer image generation model needs to understand the drawing instructions input by the user and generate an image that matches the drawing instructions. However, in the related art, a fixed image generation model is set up before all text-to-image tasks, and a fixed sampler is loaded in the image generation model. Then, for each text-to-image task that needs to be performed, the same fixed sampler in the image generation model is used to convert the drawing instructions into the target image. Also, since the content generation requirements of different text-to-image tasks are actually different, and each image generation model and sampler can only meet the content generation requirements of one type of image content, then, in the related art, each text-to-image task using the pre-set fixed image generation model and fixed sampler often has the problem that the fixed image generation model and sampler cannot adapt to the content generation requirements of the text-to-image task, resulting in a low quality of the generated image. Summary of the Invention

[0003] Embodiments of this application provide an image generation method, apparatus, device, storage medium, and vehicle, which can solve the problem of low quality of images generated by the existing fixed models and samplers.

[0004] In a first aspect, embodiments of this application provide an image generation method, the method including:

[0005] Obtain a drawing instruction input by a user, where the drawing instruction is used to indicate the generation of a target image;

[0006] Analyze the drawing instruction to determine the content category of the target image corresponding to the drawing instruction;

[0007] Determine the sampler algorithm and parameters corresponding to the content category from a sampler database, where the sampler database includes the correspondence between the content category and the sampler algorithm, and the parameters corresponding to each sampler algorithm;

[0008] Convert the drawing instruction into the target image by using the sampler algorithm and parameters.

[0009] In some embodiments, the determining the sampler algorithm and parameters corresponding to the content category from the sampler database includes:

[0010] Based on a preset first correspondence, determine an image generation model corresponding to the content category, where the first correspondence is a mapping relationship between the content category and the image generation model;

[0011] Determine a sampler algorithm in the sampler database that has a mapping relationship with the content category and the image generation model, and determine the parameters corresponding to the sampler algorithm; the sampler database includes a first mapping relationship, and the first mapping relationship includes the mapping relationship between the content category, the image generation model, and the sampler algorithm.

[0012] In some embodiments, the content category includes a subject category and a style category, and the image generation model includes a basic image generation model and a specific object generation model. Determining the image generation model corresponding to the content category includes:

[0013] Determine the serial numbers of the basic image generation model and the specific object generation model corresponding to the subject category and the style category; the first correspondence includes the mapping relationship between the subject category, the style category, and the serial numbers of the basic image generation model and the specific object generation model. Among them, the serial number of the basic image generation model is used to identify the corresponding basic image generation model, and the serial number of the specific object generation model is used to identify the corresponding specific object generation model;

[0014] The determining the sampler algorithm in the sampler database that has a mapping relationship with the content category and the image generation model includes:

[0015] Query the first mapping relationship table in the sampler database to determine the sampler algorithm corresponding to the subject category, the style category, the serial number of the basic image generation model, and the serial number of the specific object generation model. Among them, the first mapping relationship table includes the first mapping relationship between the subject category, the style category, the serial number of the basic image generation model, the serial number of the specific object generation model, and the sampler algorithm.

[0016] In some embodiments, the querying the first mapping relationship table in the sampler database to determine the sampler algorithm corresponding to the subject category, the style category, the serial number of the basic image generation model, and the serial number of the specific object generation model includes:

[0017] Combine the subject category, the style category, the serial number of the basic image generation model, and the serial number of the specific object generation model as a sampler retrieval condition set;

[0018] According to the sampler retrieval condition set, query the first mapping relationship table to determine the sampler algorithm corresponding to the sampler retrieval condition set.

[0019] In some embodiments, analyzing the drawing instruction to determine the content category of the target image corresponding to the drawing instruction includes:

[0020] When there are multiple drawing instructions, splicing the multiple drawing instructions into a drawing instruction group;

[0021] Analyzing the drawing instruction group to determine a content category set corresponding to the drawing instruction group, where the category elements in the content category set are the content categories of the target images corresponding to the respective drawing instructions in the drawing instruction group;

[0022] Determining the sampler algorithm and parameters corresponding to the content category from the sampler database according to the content category includes:

[0023] Determining a sampler algorithm set from the sampler database according to the content category set, where the sampler elements in the sampler algorithm set and the category elements in the content category set correspond one by one.

[0024] In some embodiments, the sampler algorithm includes a single sampling algorithm and the number of iterations. Converting the drawing instruction into the target image by using the sampler algorithm and parameters includes:

[0025] Obtaining a randomly generated noise image;

[0026] Encoding the drawing instruction to obtain a text embedding vector;

[0027] Embedding the text embedding vector into an image generation model, and using the image generation model to perform N denoising processes on the noise image according to the single sampling algorithm and the parameters to obtain a latent feature image, where N is the number of iterations;

[0028] Performing a decoding process on the latent feature image to obtain the target image.

[0029] In some embodiments, using the image generation model to perform N denoising processes on the noise image according to the single sampling algorithm and the parameters includes:

[0030] For the i-th denoising process among the N denoising processes, obtaining an intermediate image generated by the image generation model, and determining the noise standard deviation of the intermediate image based on the parameters, where the intermediate image is a latent space image obtained by subjecting the noise image to (i - 1) denoising processes, and i is any positive integer less than or equal to N;

[0031] Configuring the noise standard deviation in the single sampling algorithm;

[0032] Use the image generation model to perform single denoising processing on the intermediate image according to the configured single sampling algorithm until N times of denoising processing are completed.

[0033] In some embodiments, determining the noise standard deviation of the intermediate image based on the parameter includes:

[0034] Obtain the minimum value and the maximum value of the noise standard deviation of the sampler algorithm;

[0035] When i is less than N, input the parameter, the minimum value of the noise standard deviation, the maximum value of the noise standard deviation, and i into the noise determination formula to calculate the noise standard deviation;

[0036] When i is equal to N, determine the noise standard deviation as 0.

[0037] In a second aspect, an embodiment of the present application provides an image generation device, which includes:

[0038] An acquisition module, configured to acquire a drawing instruction input by a user, where the drawing instruction is used to indicate generating a target image;

[0039] An analysis module, configured to analyze the drawing instruction to determine the content category of the target image corresponding to the drawing instruction;

[0040] A determination module, configured to determine the sampler algorithm and parameters corresponding to the content category from a sampler database, where the sampler database includes a correspondence between the content category and the sampler algorithm, and parameters corresponding to each sampler algorithm;

[0041] A conversion module, configured to convert the drawing instruction into the target image by using the sampler algorithm and parameters.

[0042] In a third aspect, an embodiment of the present application provides an image generation device, which includes: a processor and a memory storing computer program instructions;

[0043] When the processor executes the computer program instructions, the above image generation method is implemented.

[0044] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above image generation method is implemented.

[0045] In a fifth aspect, an embodiment of the present application provides a vehicle, which includes the above image generation device, image generation device, and computer storage medium.

[0046] In this application, by analyzing the drawing instructions input by the user, the content category of the target image corresponding to the drawing instructions is determined according to the analysis result. Then, the sampler algorithm and parameters corresponding to the content category are queried in the sampler database, and the drawing instructions are converted into the target image by using the sampler algorithm and parameters. In this way, according to the different content categories of the drawing instructions, the appropriate sampler algorithm and parameters can be automatically selected to complete the conversion from the drawing instructions to the target image. Compared with the prior art, for different text-to-image tasks in this application, different samplers with different characteristics are dynamically selected based on the attributes of the drawing instructions, avoiding the standardized process of applying one sampler to all text-to-image tasks. By adapting different samplers to different text-to-image tasks, the quality of image generation is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings required for use in the embodiments of this application will be briefly introduced below. Obviously, the following described drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0048] Figure 1 It is a schematic flowchart of an image generation method provided by an embodiment of this application;

[0049] Figure 2 It is a schematic structural diagram of an image generation device provided by an embodiment of this application;

[0050] Figure 3 It is a schematic hardware structure diagram of an image generation device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the following further describes this application in detail with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit this application. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of this application by showing examples of this application.

[0052] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0053] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The embodiments will be described in detail below in conjunction with the accompanying drawings.

[0054] Specifically, in order to solve the problems of the prior art, the embodiments of the present application provide an image generation method, device, equipment, storage medium and vehicle. The image generation method provided by the embodiments of the present application is first introduced below.

[0055] Figure 1 A flow chart of an image generation method provided by an embodiment of the present application is shown. The method can be applied to a vehicle computer or a cloud server connected to the vehicle in communication, and the method includes the following steps:

[0056] S110, obtaining a drawing instruction input by a user, where the drawing instruction is used to instruct generation of a target image.

[0057] In this embodiment, the image generation task can convert the drawing instruction input by the user into a target image, and the drawing instruction is used to describe the content of the target image that the user wants to generate. The drawing instruction can be a voice instruction, a text instruction, or an image instruction. The voice instruction can be a voice spoken by the user, for example, "generate a puppy", the text instruction can be a text input by the user, for example, "generate a big tree", and the image instruction can be an image input by the user.

[0058] In addition, the input voice command or text command may include a positive command and a negative command. The positive command means the content that is expected to be included in the target image, while the negative command means the content that is not expected to be included in the target image. For example, the positive command may be "a cartoon-style Chinese girl running on the beach, with flying seagulls and a gorgeous rainbow behind her, and the overall picture is poetic and picturesque", while the negative command may be "pornographic, nude, ugly, deformed".

[0059] S120. Analyze the drawing instruction to determine the content category of the target image corresponding to the drawing instruction.

[0060] In this embodiment, in an image generation task, after receiving a drawing instruction, the drawing instruction can be analyzed, and based on the analysis result, the drawing instruction can be converted into corresponding text content, and then the semantic understanding of the text content can be performed, that is, the content category of each item in the target image can be determined.

[0061] In addition, the content of the target image can include objects in the target image, the background in the target image, the image style of the target image, the main color of the target image, and the light and shadow effects of the target image, etc. Contents with similar feature information can be determined as the same content category. Feature information is the attribute feature of the content indicated by the prompt word. Such as object attributes, style attributes, color attributes, etc.

[0062] The content category includes a subject category and a style category. The subject category refers to the type of the subject in the target image. For example, the subject category can be people, scenery, animals, food, transportation tools, etc. When the subject in the drawing instruction is a cat or a dog, the subject category of the drawing instruction is an animal. When the subject in the drawing instruction is coffee or bread, the subject category of the drawing instruction is food. The style category usually refers to the visual appearance and style characteristics of the image. Common style categories can include two major categories: realistic styles and artistic styles.

[0063] Among them, realistic styles can include realistic people, scenery, animals, buildings, etc. Realistic styles emphasize real color expression and strive to restore the real color and light and shadow effects of people. Therefore, they have high requirements for image quality and are very demanding on the details of the image. Artistic styles can include styles such as comics, oil paintings, ink paintings, and line drawings. They focus more on the display of painting styles and have lower requirements for image details than realistic styles.

[0064] S130. Determine the sampler algorithm and parameters corresponding to the content category from the sampler database. The sampler database includes the correspondence between the content category and the sampler algorithm, and the parameters corresponding to each sampler algorithm.

[0065] In this embodiment, the sampler database includes a variety of different sampler algorithms, as well as the correspondence between each sampler algorithm and the content category. The correspondence between each sampler algorithm and the parameters. Among them, the sampler algorithm and parameters are used to control the intensity of noise reduction, the noise reduction method, and the number of iterations. There are various types of sampler algorithms, such as DDIM, Euler, DPM++2M, etc. Different sampler algorithms have different characteristics.

[0066] As an alternative embodiment, determining the sampler algorithm and parameters corresponding to the content category from the sampler database includes:

[0067] Based on a preset first correspondence, determining an image generation model corresponding to the content category, where the first correspondence is a mapping relationship between the content category and the image generation model;

[0068] Determining, in the sampler database, a sampler algorithm that has a mapping relationship with the content category and the image generation model, and determining the parameters corresponding to the sampler algorithm; the sampler database includes a first mapping relationship, and the first mapping relationship includes the mapping relationship between the content category and the image generation model and the sampler algorithm.

[0069] In this embodiment, each image generation model is used to implement one or more image generation tasks in a vertical domain. Therefore, the first correspondence between the content category and the image generation model, the content category, and the first mapping relationship between the image generation model and the sampler algorithm can be set in advance in the sampler database based on the vertical domain corresponding to each image generation model.

[0070] As an alternative embodiment, the content category includes a subject category and a style category, and the image generation model includes a basic image generation model and a specific object generation model. Determining the image generation model corresponding to the content category includes:

[0071] Determining the serial numbers of the basic image generation model and the specific object generation model corresponding to the subject category and the style category; the first correspondence includes the mapping relationship between the subject category and the style category and the serial numbers of the basic image generation model and the specific object generation model, where the serial number of the basic image generation model is used to identify the corresponding basic image generation model, and the serial number of the specific object generation model is used to identify the corresponding specific object generation model;

[0072] Determining, in the sampler database, the sampler algorithm that has a mapping relationship with the content category and the image generation model includes:

[0073] Querying the first mapping relationship table in the sampler database to determine the sampler algorithm corresponding to the subject category, the style category, the serial number of the basic image generation model, and the serial number of the specific object generation model, where the first mapping relationship table includes the first mapping relationship between the subject category, the style category, the serial number of the basic image generation model, the serial number of the specific object generation model, and the sampler algorithm.

[0074] In this embodiment, the image generation model may include a basic image generation model and a specific object generation model. Among them, the basic image generation model may be the basic image generation model (base model) in the latent diffusion image generation model (Stable Diffusion model), which is used to generate basic images, and the specific object generation model may be the low-rank adaptation of large language models (LoRA image generation model) in the latent diffusion image generation model, which is used to generate various stylized images.

[0075] The drawing instruction can be used to guide the image generation model to perform multiple iterations of noise reduction processing on a random noise image to obtain the final target image. The sampler is used to be loaded into the image generation model during the iterative process of repeated noise reduction of the image generation model to control the intensity of noise reduction, the noise reduction method, and the number of iterations.

[0076] Exemplarily, multiple different basic image generation models and multiple different specific object generation models can be set. During the model training process, each basic image generation model is trained with a dataset of the same subject category, and each specific object generation model is trained with the same style category. In this way, a first correspondence relationship between the content category and the image generation model can be defined in the sampler database based on the above relationship.

[0077] Then, the subject category to which the subject of the drawing instruction belongs can be determined, and based on the first correspondence relationship, the basic image generation model corresponding to the subject category can be determined. Based on the first correspondence relationship, the specific object generation model corresponding to the style category can be determined. The sampler is jointly determined by the subject category, the style category, and the selected basic image generation model and specific object generation model.

[0078] Exemplarily, a subject category set S P = [p cls1 , p cls2 … p clsn can be defined in advance. The elements in this subject category set may include animals, plants, people, food, etc. The subject category of the drawing instruction can be any element in this subject category set. A style set S t = [t cls1 , t cls2 … t clsn can also be defined in advance. The style category of the target image can also be any element in this style set.

[0079] In addition, it is also possible to extract features from the drawing instructions, and then process the extracted features through common neural network structures such as pooling, convolution, and residual to obtain the main body of the drawing instructions and the main body category. Then, through a predefined first correspondence, query the basic image generation model number and the specific object generation model number corresponding to the main body category and the style category. Specifically, multiple different basic image generation models and multiple different specific object generation models can be set. Each basic image generation model has a uniquely corresponding basic image generation model number, and each specific object generation model has a uniquely corresponding specific object generation model number. The first correspondence includes the mapping relationship between the main body category and the basic image generation model, and the mapping relationship between the style category and the specific object generation model.

[0080] As an optional embodiment, the querying the first mapping table in the sampler database to determine the sampler algorithm corresponding to the main body category, the style category, the basic image generation model number, and the specific object generation model number includes:

[0081] Combining the main body category, the style category, the basic image generation model number, and the specific object generation model number to form a sampler retrieval condition set;

[0082] According to the sampler retrieval condition set, query the first mapping table to determine the sampler algorithm corresponding to the sampler retrieval condition set.

[0083] In this embodiment, after determining the main body category, the style category, the basic image generation model number, and the specific object generation model number, a new sampler retrieval condition set including the main body category, the style category, the basic image generation model number, and the specific object generation model number can be generated:

[0084] [Main body category, style category, basic image generation model number, specific object generation model number]

[0085] Among them, the four elements in the sampler retrieval condition set are all input parameters required for sampler initialization, and then query the sampler that has a first mapping relationship with this sampler retrieval condition set in the predefined first mapping table.

[0086] For example, for an input where the sampler retrieval condition set is input1 = [person, realistic, person model, person LORA], a DPM++ sampler that has a first mapping relationship with this sampler retrieval condition set can be queried. The characteristics of the DPM++ sampler are good denoising quality, balanced efficiency, stable denoising effect, and the ability to restore and trace for the same initial noise. Due to the realistic characteristic, we take the number of denoising iterations timesteps = T1. For an input where the sampler retrieval condition set is input2 = [person, anime, anime model, anime LORA], an Euler sampler that has a first mapping relationship with this sampler retrieval condition set can be queried. The characteristic of the Euler sampler is that its iteration speed is very fast, and it can achieve a good denoising effect in a relatively short number of iterations. Its number of denoising iterations timesteps = T2. Among them, T2 is less than T1.

[0087] Through the above method, the image generation model and sampler matching the text-to-image task can be accurately and conveniently determined through the drawing instruction.

[0088] S140, convert the drawing instruction into the target image by using the sampler algorithm and parameters.

[0089] In this embodiment, a sampler can be constructed based on the determined sampler algorithm and parameters, and the determined sampler can be loaded into the image generation model. Then, a random noise image is obtained, and the drawing instruction is converted into a text embedding vector. The text embedding vector is embedded into the image generation model loaded with the sampler to guide the image generation model to perform multiple iterations of denoising on the noise image to obtain a latent feature image, and then the latent feature image is decoded into a visible target image.

[0090] In the embodiment of the present application, by analyzing the drawing instruction input by the user, the content category of the target image corresponding to the drawing instruction is determined according to the analysis result. Then, the sampler algorithm and parameters corresponding to the content category are queried in the sampler database, and the drawing instruction is converted into the target image by using the sampler algorithm and parameters. In this way, according to the different content categories of the drawing instructions, the appropriate sampler algorithm and parameters can be automatically selected to complete the conversion from the drawing instruction to the target image. Compared with the prior art, for different text-to-image tasks, the present application dynamically selects samplers with different characteristics based on the attributes of the drawing instructions, avoiding the standardized process of applying one sampler to all text-to-image tasks. By using different samplers to adapt to different text-to-image tasks, the quality of image generation is improved.

[0091] As an optional embodiment, the analyzing the drawing instruction to determine the content category of the target image corresponding to the drawing instruction includes:

[0092] When there are multiple drawing instructions, splice the multiple drawing instructions into a drawing instruction group;

[0093] Analyze the drawing instruction group to determine the content category set corresponding to the drawing instruction group, where the category elements in the content category set are the content categories of the target images corresponding to the respective drawing instructions in the drawing instruction group;

[0094] Determining the sampler algorithm and parameters corresponding to the content category from the sampler database according to the content category includes:

[0095] Determine a sampler algorithm set from the sampler database according to the content category set, and the sampler elements in the sampler algorithm set correspond one-to-one with the category elements in the content category set.

[0096] In this embodiment, the user can input multiple drawing instructions in the same batch. Then, the multiple drawing instructions can be spliced into a drawing instruction group. For all the drawing instructions in the drawing instruction group, analysis can be performed simultaneously, and then the content category corresponding to each drawing instruction can be determined to obtain multiple content categories. Then, a content category set containing the multiple content categories is generated, and the content category corresponding to each drawing instruction is one of the category elements in the content category.

[0097] For this content category set, the sampler algorithm corresponding to each category element can also be determined from the sampler database to obtain multiple sampler algorithms, and these multiple sampler algorithms are also combined into a sampler algorithm set.

[0098] For example, for the content category set [input1, input2..inputN], where each category element is a content category, the sampler algorithm corresponding to each category element in the content category set can be queried separately, and the multiple sampler algorithms obtained by the query are combined into a sampler algorithm set [sampler1, sampler2...samplerN].

[0099] In this way, for the case of inputting multiple drawing instructions in the same batch, the sampler algorithms corresponding to the respective drawing instructions can be screened and output simultaneously, improving the efficiency of image generation.

[0100] As an alternative embodiment, using the sampler and the image generation model to convert the drawing instruction into the target image includes:

[0101] Load the sampler into the basic image generation model and the specific object generation model;

[0102] Mount the specific object generation model to the basic image generation model to obtain an image generation model;

[0103] Use the image generation model to convert the drawing instruction into the target image.

[0104] In this embodiment, since the sampler is used to control the noise reduction intensity, noise reduction method, and number of iterations during the iterative process of repeatedly denoising the image generation model. Therefore, it needs to be loaded into the basic image generation model and the specific object generation model respectively before the image generation model runs. Then, the specific object generation model can be spliced onto the basic image generation model, and the parameter weights of the specific object generation model and the basic image generation model can be adjusted to obtain an image generation model, and the text-to-image task can be completed through the image generation model. In this way, the basic image generation model and the specific object generation model can jointly complete the conversion of the drawing instruction into the target image.

[0105] Through the above method, the specific object generation model can be mounted on the basic image generation model, thereby completing the actual deployment and inference of the image generation model in the text-to-image task.

[0106] As an alternative embodiment, the sampler includes a single sampling algorithm and the number of iterations. The step of using the image generation model to convert the drawing instruction into the target image includes:

[0107] Obtain a randomly generated noise image;

[0108] Encode the drawing instruction to obtain a text embedding vector;

[0109] Embed the text embedding vector into the image generation model, and use the image generation model to perform N times of noise reduction processing on the noise image according to the single sampling algorithm and the parameters, to obtain a latent feature image, where N is the number of iterations;

[0110] Perform decoding processing on the latent feature image to obtain the target image.

[0111] In this embodiment, a randomly generated noise image can be obtained, and the text content of the drawing instruction can be mapped to a high-dimensional vector space to obtain a text embedding vector. After obtaining the text embedding vector, the text embedding vector and the noise image can be jointly input into the image generation model, and the text embedding vector is used to guide the image generation model to perform N times of iterative noise reduction processing on the noise image according to the single sampling algorithm, to obtain a latent feature image. Then, the latent feature image can be decoded by using the variational autoencoder image generation model to obtain a target image that can be recognized by the human eye.

[0112] In this embodiment, the conversion from the drawing instruction to the target image can be accurately completed.

[0113] As an alternative embodiment, the performing N denoising processes on the noise image by using the image generation model according to the single sampling algorithm and the parameters includes:

[0114] For the i-th denoising process among the N denoising processes, obtaining an intermediate image generated by the image generation model, and determining the noise standard deviation of the intermediate image based on the parameters, where the intermediate image is a latent space image obtained by subjecting the noise image to (i - 1) denoising processes, and i is any positive integer less than or equal to N;

[0115] Configuring the noise standard deviation in the single sampling algorithm;

[0116] Performing a single denoising process on the intermediate image by using the image generation model according to the configured single sampling algorithm until N denoising processes are completed.

[0117] In this embodiment, before applying the sampler, it is necessary to first complete the construction of the sampler. During the construction of the sampler, it is necessary to first initialize the sampler, and then determine the single sampling algorithm and the number of iterations of the sampler. Among them, the single sampling algorithm determines the method and intensity of noise removal in each denoising process.

[0118] Specifically, the noise standard deviation of the intermediate image is a model parameter in the determination process of each single sampling algorithm. The model parameter is a parameter learned during the execution of the algorithm, and the model parameter will be updated during each round of iteration. Therefore, it is necessary to obtain the noise standard deviation of the intermediate image before each denoising process is executed, then update the single sampling algorithm by using the noise standard deviation, and perform a single denoising process on the input intermediate image by using the updated single sampling algorithm.

[0119] Specifically, for the i-th denoising process among the N denoising processes, its denoising formula is as follows:

[0120] D θ (x; σ) = c skip (σ)x + c out (σ)F θ (c in (σ)x; c noise (σ))

[0121] where x represents the latent space image obtained by subjecting the noise image to (i - 1) denoising processes, that is, the intermediate image. If i is 1, then x is the noise image, D θIt is the latent space image after the i-th noise reduction process. σ is used to represent the standard deviation of the noise of the latent space image obtained after (i - 1) times of noise reduction processes. F θ represents the inference of the noise reduction neural network algorithm. And Cskip(σ), Cout(σ), Cin(σ), and Cnoise(σ) represent the four noise reduction process terms in the noise reduction formula respectively. Specifically, Cskip(σ) represents the skip connection scaling process, Cout(σ) represents the output scaling process, Cin(σ) represents the input scaling process, and Cnoise(σ) represents the noise adjustment process.

[0122] In different samplers, the above four noise reduction process terms are not the same. For example, in the DPM sampler:

[0123] Skip scaling c skip (σ) 1

[0124] Output scaling c out (σ) -σ

[0125] Input scaling c in (σ)

[0126] Noise cond.c noise (σ) (M - 1)σ -1 (σ)

[0127] In the DDIM sampler:

[0128] Skip scaling c skip (σ) 1

[0129] Output scaling c out (σ) -σ

[0130] Input scaling c in (σ)

[0131] Noise cond.c noise (σ) M - 1 - arg min j |u j -σ|

[0132] Among them, M is the initial value of the random noise, and μj is the j-th noise adjustment parameter.

[0133] In this way, when the type of the sampler is determined, the sampler can be constructed according to the noise reduction process and the number of iterations specified for this type of sampler. After the sampler is constructed, for each iteration of noise reduction, the latent space image obtained from the previous noise reduction process and the noise standard deviation of this latent space image are input into the sampler, and a new round of noise reduction can be completed through the sampler.

[0134] As an alternative embodiment, determining the noise standard deviation of obtaining the intermediate image based on the parameter includes:

[0135] Obtaining the minimum value of the noise standard deviation and the maximum value of the noise standard deviation of the sampler algorithm;

[0136] When i is less than N, the parameter, the minimum value of the noise standard deviation, the maximum value of the noise standard deviation, and i are input into the noise determination formula to calculate the noise standard deviation;

[0137] When i is equal to N, the noise standard deviation is determined to be 0.

[0138] In this embodiment, when the sampler is constructed and noise reduction is achieved using the sampler, it is necessary to obtain the noise standard deviation of the intermediate image once during each noise reduction process, and input the noise standard deviation together with the intermediate image into the sampler.

[0139] Specifically, during the i-th noise reduction process, the parameters of the sampler, as well as the previously specified minimum value of the noise standard deviation and the maximum value of the noise standard deviation, can be obtained first. Among them, the parameter is used to control the step size of the variance transformation.

[0140] When i is less than N, the noise standard deviation during the i-th iteration process can be calculated according to the following noise determination formula:

[0141]

[0142] where i represents the current iteration number, N represents the total number of iterations, σ min and σ max represent the minimum value of the noise standard deviation and the maximum value of the noise standard deviation respectively, and ρ is the parameter.

[0143] Exemplarily, the parameter can be 7, σ min and σ max are 0.02 and 100 respectively.

[0144] When i is equal to N, σ N is 0.

[0145] In this way, the noise standard deviation of the sampler can be adjusted according to a predetermined step size, avoiding the change of the noise standard deviation being too small or too large.

[0146] Based on the image generation method provided in the above embodiments, correspondingly, the present application also provides a specific implementation manner of an image generation device. Please refer to the following embodiments.

[0147] First, refer to Figure 2 , the image generation device 200 provided in the embodiments of the present application includes the following modules:

[0148] An acquisition module 201, configured to acquire a drawing instruction input by a user, where the drawing instruction is used to indicate the generation of a target image;

[0149] An analysis module 202, configured to analyze the drawing instruction to determine the content category of the target image corresponding to the drawing instruction;

[0150] A determination module 203, configured to determine a sampler algorithm and parameters corresponding to the content category from a sampler database according to the content category, where the sampler database includes a correspondence between the content category and the sampler algorithm, and parameters corresponding to each sampler algorithm;

[0151] A conversion module 204, configured to convert the drawing instruction into the target image by using the sampler algorithm and parameters.

[0152] The device can analyze the drawing instruction input by the user, determine the content category of the target image corresponding to the drawing instruction according to the analysis result, then query the sampler algorithm and parameters corresponding to the content category in the sampler database according to the content category, and convert the drawing instruction into the target image by using the sampler algorithm and parameters. In this way, the appropriate sampler algorithm and parameters can be automatically selected according to the different content categories of the drawing instructions to complete the conversion from the drawing instruction to the target image. Compared with the prior art, for different text-to-image tasks, the present application dynamically selects samplers with different characteristics based on the attributes of the drawing instructions, avoiding the standardized process of applying one sampler to all text-to-image tasks, and adapting different samplers to different text-to-image tasks, thereby improving the quality of image generation.

[0153] As an implementation manner of the present application, the above determination module 203 may further include:

[0154] A first determination unit, configured to determine an image generation model corresponding to the content category based on a preset first correspondence, where the first correspondence is a mapping relationship between the content category and the image generation model;

[0155] A second determination unit, configured to determine a sampler algorithm mapped to the content category and the image generation model in the sampler database, and determine parameters corresponding to the sampler algorithm; the sampler database includes a first mapping relationship, and the first mapping relationship includes the mapping relationship between the content category and the image generation model and the sampler algorithm.

[0156] As an implementation manner of this application, the above first determination unit may further include:

[0157] A first query subunit, configured to determine a basic image generation model serial number and a specific object generation model serial number corresponding to the subject category and the style category; the first corresponding relationship includes the mapping relationship between the subject category and the style category and the basic image generation model serial number and the specific object generation model serial number, where the basic image generation model serial number is used to identify the corresponding basic image generation model, and the specific object generation model serial number is used to identify the corresponding specific object generation model;

[0158] The above second determination unit may further include:

[0159] A second query subunit, configured to query a first mapping relationship table in the sampler database to determine a sampler algorithm corresponding to the subject category, the style category, the basic image generation model serial number, and the specific object generation model serial number, where the first mapping relationship table includes a first mapping relationship between the subject category, the style category, the basic image generation model serial number, the specific object generation model serial number, and the sampler algorithm.

[0160] As an implementation manner of this application, the above second query subunit may further be configured to:

[0161] Combine the subject category, the style category, the basic image generation model serial number, and the specific object generation model serial number to form a sampler retrieval condition set;

[0162] Query the first mapping relationship table according to the sampler retrieval condition set to determine a sampler algorithm corresponding to the sampler retrieval condition set.

[0163] As an implementation manner of this application, the above analysis module 202 may include:

[0164] A splicing unit, configured to splice the multiple drawing instructions into a drawing instruction group when the drawing instructions are multiple;

[0165] An analysis unit for analyzing the drawing instruction group to determine a set of content categories corresponding to the drawing instruction group, where the category elements in the set of content categories are the content categories of the target images corresponding to the respective drawing instructions in the drawing instruction group;

[0166] The above-mentioned determination module 203 may further include:

[0167] A third determination unit for determining a set of sampler algorithms from the sampler database according to the set of content categories, where the sampler elements in the set of sampler algorithms correspond one-to-one with the category elements in the set of content categories.

[0168] As an implementation manner of the present application, the above-mentioned conversion module 204 may include:

[0169] An acquisition unit for acquiring a randomly generated noise image;

[0170] An encoding unit for encoding the drawing instruction to obtain a text embedding vector;

[0171] A noise reduction unit for embedding the text embedding vector into an image generation model and using the image generation model to perform N times of noise reduction processing on the noise image according to the single-sampling algorithm and the parameters to obtain a latent feature image, where N is the number of iterations;

[0172] A decoding unit for performing decoding processing on the latent feature image to obtain the target image.

[0173] As an implementation manner of the present application, the above-mentioned noise reduction unit may also be used for:

[0174] For the i-th noise reduction processing among the N times of noise reduction processing, obtaining an intermediate image generated by the image generation model and determining the noise standard deviation of the intermediate image based on the parameters, where the intermediate image is the latent space image obtained by subjecting the noise image to (i - 1) times of noise reduction processing, and i is any positive integer less than or equal to N;

[0175] Configuring the noise standard deviation in the single-sampling algorithm;

[0176] Using the image generation model to perform single noise reduction processing on the intermediate image according to the configured single-sampling algorithm until N times of noise reduction processing are completed.

[0177] As an implementation manner of the present application, the above-mentioned noise reduction unit may also be used for:

[0178] Obtaining the minimum value and the maximum value of the noise standard deviation of the sampler algorithm;

[0179] When i is less than N, the parameter, the minimum value of the noise standard deviation, the maximum value of the noise standard deviation, and i are input into the noise standard deviation determination formula to calculate the noise standard deviation.

[0180] When i is equal to N, the noise standard deviation is determined to be 0.

[0181] The image generation device provided by the embodiment of the present invention can implement each step in the above method embodiment. To avoid repetition, it will not be elaborated here.

[0182] Figure 3 The hardware structure diagram of the image generation device provided by the embodiment of the present application is shown.

[0183] The image generation device may include a processor 301 and a memory 302 storing computer program instructions.

[0184] Specifically, the above-mentioned processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0185] The memory 302 may include a mass storage for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In a suitable case, the memory 302 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 302 may be internal or external to the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid state memory.

[0186] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of the present disclosure.

[0187] The processor 301 reads and executes the computer program instructions stored in the memory 302 to implement any one of the image generation methods in the above embodiments.

[0188] In one example, the image generation device may further include a communication interface 303 and a bus 310. Among them, as Figure 3 shown, the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 to complete communication with each other.

[0189] The communication interface 303 is mainly used to implement communication between various modules, devices, units, and / or devices in the embodiments of the present application.

[0190] The bus 310 includes hardware, software, or both, and couples the components of the image generation device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 310 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0191] The image generation device may be based on the above embodiments, so as to implement the image generation method and device in combination with the above.

[0192] In addition, in combination with the image generation method in the above embodiments, the embodiments of the present application may provide a computer storage medium to implement. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the image generation methods in the above embodiments is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be described in detail here. Among them, the above computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc., which is not limited herein.

[0193] In addition, the embodiments of the present application further provide a vehicle, including computer program instructions, which can implement the steps and corresponding contents of the foregoing method embodiments when the computer program instructions are executed by a processor.

[0194] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0195] The functional blocks shown in the above structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0196] It should also be noted that in the exemplary embodiments mentioned in the present application, some methods or systems are described based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0197] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses, and vehicles according to embodiments of the present disclosure. It should be understood that each block in the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices to generate a machine such that the instructions executed by the processor of the computer or other programmable data processing devices enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0198] The above are only specific implementation manners of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. An image generation method, characterized in that, The method includes: Obtaining a drawing instruction input by a user, where the drawing instruction is used to indicate generating a target image; Analyzing the drawing instruction to determine the content category of the target image corresponding to the drawing instruction; Determining, from a sampler database, a sampler algorithm and parameters corresponding to the content category according to the content category, where the sampler database includes a correspondence between the content category and the sampler algorithm, and parameters corresponding to each sampler algorithm; Converting the drawing instruction into the target image by using the sampler algorithm and parameters.

2. The image generation method according to claim 1, characterized in that, The determining, from the sampler database, a sampler algorithm and parameters corresponding to the content category according to the content category includes: Based on a preset first correspondence, determining an image generation model corresponding to the content category, where the first correspondence is a mapping relationship between the content category and the image generation model; Determining, in the sampler database, a sampler algorithm having a mapping relationship with the content category and the image generation model, and determining parameters corresponding to the sampler algorithm; the sampler database includes a first mapping relationship, and the first mapping relationship includes a mapping relationship between the content category and the image generation model and the sampler algorithm.

3. The image generation method according to claim 2, characterized in that, The content category includes a main body category and a style category, and the image generation model includes a basic image generation model and a specific object generation model. The determining an image generation model corresponding to the content category includes: Determining a basic image generation model serial number and a specific object generation model serial number corresponding to the main body category and the style category; the first correspondence includes a mapping relationship between the main body category and the style category and the basic image generation model serial number and the specific object generation model serial number, where the basic image generation model serial number is used to identify the corresponding basic image generation model, and the specific object generation model serial number is used to identify the corresponding specific object generation model; The determining, in the sampler database, a sampler algorithm having a mapping relationship with the content category and the image generation model includes: Querying a first mapping relationship table in the sampler database to determine a sampler algorithm corresponding to the main body category, the style category, the basic image generation model serial number, and the specific object generation model serial number, where the first mapping relationship table includes a first mapping relationship between the main body category, the style category, the basic image generation model serial number, and the specific object generation model serial number and the sampler algorithm.

4. The image generation method according to claim 3, characterized in that, The querying a first mapping relationship table in the sampler database to determine a sampler algorithm corresponding to the main body category, the style category, the basic image generation model serial number, and the specific object generation model serial number includes: Combining the main body category, the style category, the basic image generation model serial number, and the specific object generation model serial number to form a sampler retrieval condition set; Querying the first mapping relationship table according to the sampler retrieval condition set to determine a sampler algorithm corresponding to the sampler retrieval condition set.

5. The image generation method according to claim 1, characterized in that, Analyzing the drawing instructions to determine the content category of the target image corresponding to the drawing instructions includes: In the case where there are multiple drawing instructions, splicing the multiple drawing instructions into a drawing instruction group; Analyzing the drawing instruction group to determine a content category set corresponding to the drawing instruction group, where the category elements in the content category set are the content categories of the target images corresponding to the respective drawing instructions in the drawing instruction group; Determining the sampler algorithm and parameters corresponding to the content category from the sampler database according to the content category includes: Determining a sampler algorithm set from the sampler database according to the content category set, and the sampler elements in the sampler algorithm set correspond one-to-one with the category elements in the content category set.

6. The image generation method according to claim 1, characterized in that, The sampler algorithm includes a single sampling algorithm and the number of iterations. Converting the drawing instructions into the target image by using the sampler algorithm and parameters includes: Obtaining a randomly generated noise image; Encoding the drawing instructions to obtain a text embedding vector; Embedding the text embedding vector into an image generation model, and using the image generation model to perform N noise reduction processes on the noise image according to the single sampling algorithm and the parameters to obtain a latent feature image, where N is the number of iterations; Performing a decoding process on the latent feature image to obtain the target image.

7. The image generation method according to claim 6, characterized in that, Using the image generation model to perform N noise reduction processes on the noise image according to the single sampling algorithm and the parameters includes: For the i-th noise reduction process among the N noise reduction processes, obtaining an intermediate image generated by the image generation model, and determining the noise standard deviation of the intermediate image based on the parameters, where the intermediate image is the latent space image obtained by subjecting the noise image to (i - 1) noise reduction processes, and i is any positive integer less than or equal to N; Configuring the noise standard deviation in the single sampling algorithm; Using the image generation model to perform a single noise reduction process on the intermediate image according to the configured single sampling algorithm until N noise reduction processes are completed.

8. The image generation method according to claim 7, characterized in that, Determining the noise standard deviation of obtaining the intermediate image based on the parameters includes: Obtaining the minimum noise standard deviation and the maximum noise standard deviation of the sampler algorithm; In the case where i is less than N, inputting the parameters, the minimum noise standard deviation, the maximum noise standard deviation, and i into a noise determination formula to calculate the noise standard deviation; In the case where i is equal to N, determining the noise standard deviation to be 0.

9. An image generation device, characterized in that, The device includes: An acquisition module, configured to acquire drawing instructions input by a user, where the drawing instructions are used to indicate generating a target image; An analysis module, configured to analyze the drawing instructions to determine the content category of the target image corresponding to the drawing instructions; A determination module, configured to determine the sampler algorithm and parameters corresponding to the content category from a sampler database according to the content category, where the sampler database includes the correspondence between the content category and the sampler algorithm, and the parameters corresponding to each sampler algorithm; A conversion module, configured to convert the drawing instruction into the target image by using the sampler algorithm and parameters.

10. An image generation device, characterized in that, The image generation device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image generation method according to any one of claims 1-8 is implemented.

11. A computer storage medium, characterized in that, Computer program instructions are stored on the computer storage medium, and when the computer program instructions are executed by a processor, the image generation method according to any one of claims 1-8 is implemented.

12. A vehicle, characterized in that, The vehicle includes at least one of the following: The image generation device according to claim 9; The image generation device according to claim 10; The computer storage medium according to claim 11.