Image generation method and device, equipment, storage medium and vehicle

By semantic recognition and encoding of prompt words entered by users, and combining the inference process of text embedding vectors, the target image is generated, which solves the problem of high image generation cost in the prior art, and realizes resource optimization and cost reduction.

CN120182970APending Publication Date: 2025-06-20BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311755363.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, literary and artistic drawing tasks require large power consumption, resulting in excessive cost of image generation.

Method used

By semantic recognition of the prompt words entered by the user, the image style of the target image is determined, and the prompt words are encoded according to the image style to obtain a text embedding vector. Then, the inference flow corresponding to the text embed vector is used to convert it into the target image.

Benefits of technology

It realizes automatic selection of text embedding vectors and algorithm processes according to the image style, and reasonably allocates resources, avoids high power consumption of standardized processes, and reduces the cost of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182970A_ABST
    Figure CN120182970A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and device, equipment, a storage medium and a vehicle. The method comprises the steps of obtaining a cue word input by a user, wherein the cue word is used for indicating generation of a target image; semantic recognition is carried out on the cue word, and the image style of the target image is determined according to a semantic recognition result; the prompt words are coded according to coding modes corresponding to the image styles, text embedding vectors corresponding to the image styles are obtained, and different image styles correspond to different coding modes; and converting the text embedding vector according to a reasoning process corresponding to the text embedding vector to obtain a target image. According to the embodiment of the invention, the image generation cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to an image generation method, apparatus, device, storage medium, and vehicle. Background Art

[0002] Text-to-image is a computer generation task aimed at converting text descriptions or natural language texts into corresponding images. In this task, the computer model needs to understand the prompt words input by the user and generate an image that matches the prompt words. However, in the related art, all text-to-image tasks are inferred based on the same set of standard algorithmic processes, that is, first directly generating a target image with a very high level of detail richness using a text-to-image model, and this standard inference process of directly generating a high-richness target image requires extremely high power consumption. That is to say, each text-to-image task in the related art requires relatively high power consumption, resulting in too high a cost for image generation. Summary of the Invention

[0003] Embodiments of this application provide an image generation method, apparatus, device, storage medium, and vehicle, which can solve the problem of too high a cost for existing image generation.

[0004] In a first aspect, embodiments of this application provide an image generation method, the method comprising:

[0005] Obtaining a prompt word input by a user, the prompt word being used to indicate generating a target image;

[0006] Determining an image style of the target image according to a result of semantic recognition by performing semantic recognition on the prompt word;

[0007] Encoding the prompt word according to an encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style, where different image styles correspond to different encoding methods;

[0008] Converting the text embedding vector according to an inference process corresponding to the text embedding vector to obtain the target image.

[0009] In some embodiments, the encoding the prompt word according to an encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style comprises:

[0010] Determining a target fineness degree of the target image according to an attribute of the image style;

[0011] Querying an encoding method corresponding to the target fineness degree according to a pre-set encoding correspondence;

[0012] Encode the prompt according to the encoding method corresponding to the target fineness level to obtain the text embedding vector corresponding to the image style.

[0013] In some embodiments, determining the target fineness level of the target image according to the attributes of the image style includes:

[0014] When the attribute of the image style is the first type of style attribute, determine the target fineness level of the target image as a low fineness level;

[0015] When the attribute of the image style is the second type of style attribute, determine the target fineness level of the target image as a high fineness level.

[0016] In some embodiments, encoding the prompt according to the encoding method corresponding to the target fineness level to obtain the text embedding vector corresponding to the image style includes:

[0017] When the target fineness level of the target image is a high fineness level, encode the prompt to obtain a basic text embedding vector and a refined text embedding vector, where the basic text embedding vector is used to represent the basic semantics of the prompt, and the refined text embedding vector is used to represent the detailed semantics of the prompt;

[0018] When the target fineness level of the target image is a low fineness level, encode the prompt to obtain a basic text embedding vector.

[0019] In some embodiments, converting the text embedding vector according to the inference process corresponding to the text embedding vector to obtain the target image includes:

[0020] Determine the image generation model corresponding to the text embedding vector according to the model correspondence;

[0021] When there is one image generation model, use the image generation model to convert the text embedding vector to obtain the target image;

[0022] When there are at least two image generation models, cascade the at least two image generation models, and use the cascaded at least two image generation models to convert the text embedding vector to obtain the target image.

[0023] In some embodiments, determining the image generation model corresponding to the text embedding vector includes:

[0024] When the target fineness of the image style is high fineness, determine the base model corresponding to the base text embedding vector and the refinement model corresponding to the refined text embedding vector;

[0025] The cascading of the at least two image generation models and using the cascaded at least two image generation models to transform the text embedding vector to obtain the target image includes:

[0026] When the target fineness of the image style is high fineness, embed the base text embedding vector into the base model and embed the refined text embedding vector into the refinement model;

[0027] Obtain a randomly generated noise image;

[0028] Use the base model embedding the base text embedding vector to perform noise reduction processing on the noise image to obtain a first latent feature image;

[0029] Use the refinement model embedding the refined text embedding vector to perform noise reduction processing on the first latent feature image to obtain a second latent feature image;

[0030] Perform decoding processing on the second latent feature image to obtain the target image.

[0031] In some embodiments, the determining the image style of the target image by performing semantic recognition on the prompt word and according to the result of the semantic recognition includes:

[0032] Perform semantic recognition on the prompt word. When there is a specific phrase in the prompt word, determine the image style of the target image as the image style corresponding to the specific phrase, where the specific phrase is used to indicate the image style of the prompt word;

[0033] When the specific phrase does not exist in the prompt word, determine the image style of the target image as the default style.

[0034] In a second aspect, an image generation device provided by an embodiment of the present application includes:

[0035] An acquisition module, configured to acquire a prompt word input by a user, where the prompt word is used to indicate generating a target image;

[0036] A recognition module, configured to determine the image style of the target image by performing semantic recognition on the prompt word and according to the result of the semantic recognition;

[0037] An encoding module, configured to encode the prompt according to the encoding method corresponding to the image style, to obtain a text embedding vector corresponding to the image style, where different image styles correspond to different encoding methods;

[0038] A conversion module, configured to convert the text embedding vector according to the inference process corresponding to the text embedding vector, to obtain the target image.

[0039] In a third aspect, an embodiment of the present application provides an image generation device, including: a processor and a memory storing computer program instructions;

[0040] When the processor executes the computer program instructions, the above image generation method is implemented.

[0041] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above image generation method is implemented.

[0042] In a fifth aspect, an embodiment of the present application provides a vehicle, where the vehicle includes computer program instructions, and when the computer program instructions are executed by a processor, the above image generation method is implemented.

[0043] In the present application, by performing semantic recognition on the prompt input by the user, determining the image style of the target image corresponding to the prompt according to the result of the semantic recognition, then encoding the prompt according to the image style, correspondingly obtaining a text embedding vector, and converting the text embedding vector into a target image by the inference process corresponding to the text embedding vector. In this way, the text embedding vector can be automatically selected according to different image styles, so as to automatically select an algorithm process adapted to the image style for inference, realizing reasonable allocation of resources. Compared with the prior art, it avoids applying a set of standardized processes with high power consumption for all text-to-image tasks, and reduces the cost of image generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0045] Figure 1 is a flowchart of an image generation method provided by an embodiment of the present application;

[0046] Figure 2 is a flowchart of an image generation method provided by another embodiment of the present application;

[0047] Figure 3 It is a schematic flowchart of an image generation method provided by another embodiment of the present application;

[0048] Figure 4 It is a schematic structural diagram of an image generation device provided by an embodiment of the present application;

[0049] Figure 5 It is a schematic hardware structure diagram of an image generation device provided by an embodiment of the present application. Detailed implementation manners

[0050] The features and exemplary embodiments of various aspects of the present application will be described in detail below. To make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.

[0051] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the elements.

[0052] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The embodiments will be described in detail below with reference to the accompanying drawings.

[0053] Specifically, to solve the problems of the prior art, the embodiments of the present application provide an image generation method, device, device, storage medium and vehicle. First, the image generation method provided by the embodiments of the present application will be introduced below.

[0054] Figure 1 It shows a schematic flowchart of an image generation method provided by an embodiment of the present application. The method includes the following steps:

[0055] S110. Obtain the prompt entered by the user, where the prompt is used to indicate the generation of a target image.

[0056] In this embodiment, the image generation task can convert the prompt entered by the user into a target image, and the prompt is used to describe the content of the target image that the user wants to generate. For example, the prompt can be "Generate a puppy" or "Generate a big tree".

[0057] In addition, the entered prompt can include a positive prompt and a negative prompt. The meaning of the positive prompt is the content that is expected to be included in the target image, while the negative prompt indicates the content that is not expected to be included in the target image.

[0058] S120. Determine the image style of the target image according to the result of semantic recognition by performing semantic recognition on the prompt.

[0059] In this embodiment, in the image generation task, the image style generally refers to the visual appearance and style characteristics of the image. Common image styles can be broadly classified into realistic styles and artistic styles.

[0060] Among them, realistic styles can further include realistic people, landscapes, animals, buildings, etc. Realistic styles emphasize the expression of real colors and strive to restore the real colors and lighting effects of people, so they have high requirements for image quality and are very demanding on the details of the image. Artistic styles can include styles such as comics, oil paintings, ink paintings, line drawings, etc. They focus more on the display of painting styles and have lower requirements for the details of the image than realistic styles.

[0061] It is possible to perform semantic recognition on the prompt, determine the user's intention based on the result of the semantic recognition, then obtain the indication information of the image style from the user's intention, and determine the image style of the target image based on the indication information.

[0062] S130. Encode the prompt according to the encoding method corresponding to the image style to obtain the text embedding vector corresponding to the image style, where different image styles correspond to different encoding methods.

[0063] In this embodiment, the text embedding vector is a representation form that maps the text content of the prompt to a high-dimensional vector space. Through this representation form, the image generation model can better understand the meaning of the text embedding vector. Specifically, by capturing the semantic and syntactic information in the prompt, the prompt can be encoded into at least one text embedding vector.

[0064] Among them, since different image styles have different requirements for the target image, the semantic and syntactic information of the prompt can be selectively encoded into at least one text embedding vector based on the parsed image style. Then, the target images of different image styles correspond to different numbers or different types of text embedding vectors.

[0065] As an optional embodiment, encoding the prompt according to the encoding method corresponding to the image style to obtain the text embedding vector corresponding to the image style includes:

[0066] Determine the target fineness of the target image according to the attributes of the image style;

[0067] Query the encoding method corresponding to the target fineness according to the pre-set encoding correspondence;

[0068] Encode the prompt according to the encoding method corresponding to the target fineness to obtain the text embedding vector corresponding to the image style.

[0069] As an optional embodiment, determining the target fineness of the target image according to the attributes of the image style includes:

[0070] When the attribute of the image style is the first type of style attribute, determine the target fineness of the target image as low fineness;

[0071] When the attribute of the image style is the second type of style attribute, determine the target fineness of the target image as high fineness.

[0072] In this embodiment, the target fineness refers to the user's expectation for the fineness of the target image.

[0073] Exemplarily, when the style of the target image is an artistic style, the target image at this time focuses more on the display of the painting style and has low requirements for the details of the image. Therefore, the target fineness of the target image corresponding to the artistic style can be determined as low fineness, which means that some abstract or simplified inference encoding methods can be used to encode the prompt without paying too much attention to details.

[0074] When the style of the target image is a realistic style, the target image at this time has high requirements for both image quality and image details. Therefore, not only the basic semantics of the prompt need to be reflected in the target image, but also the detailed semantics of the prompt need to be reflected in the target image. Therefore, a more complex and precise encoding method can be used to encode the prompt.

[0075] As an optional embodiment, encoding the prompt according to the encoding method corresponding to the target fineness to obtain the text embedding vector corresponding to the image style includes:

[0076] When the target fineness of the target image is high fineness, encoding the prompt to obtain a basic text embedding vector and a refined text embedding vector, where the basic text embedding vector is used to represent the basic semantics of the prompt, and the refined text embedding vector is used to represent the detailed semantics of the prompt;

[0077] When the target fineness of the target image is low fineness, encoding the prompt to obtain a basic text embedding vector.

[0078] In this embodiment, since when the target fineness is high fineness, the user expects more detailed and richer content in the target image. Therefore, not only can the basic semantic information in the prompt be captured and encoded into the basic text embedding vector, but also the detailed semantic information in the prompt can be captured and encoded into the refined text embedding vector. Then, by guiding the inference of the prompt through the basic text embedding vector and the refined text embedding vector in sequence, the final target image is obtained.

[0079] Since when the target fineness is low fineness, there is not high requirement for the details of the image. Therefore, only the basic semantic information in the prompt can be captured and encoded into the basic text embedding vector, and the target image is obtained by guiding the inference of the prompt through the basic text embedding vector.

[0080] S140, converting the text embedding vector according to the inference process corresponding to the text embedding vector to obtain the target image.

[0081] In this embodiment, the image generation task can be completed by an image generation model or jointly completed by multiple cascaded models. Prompts with different image styles will generate different text embedding vectors, and different text embedding vectors will be embedded into different image generation models to complete the conversion from the prompt to the target image, so as to realize converting the prompt into the target image with different inference processes.

[0082] As an optional embodiment, converting the text embedding vector according to the inference process corresponding to the text embedding vector to obtain the target image includes:

[0083] Determine the image generation model corresponding to the text embedding vector according to the model correspondence;

[0084] When there is one image generation model, use the image generation model to transform the text embedding vector to obtain the target image;

[0085] When there are at least two image generation models, cascade the at least two image generation models, and use the cascaded at least two image generation models to transform the text embedding vector to obtain the target image.

[0086] In this embodiment, the model correspondence is a one-to-one correspondence between the text embedding vector and the image generation model set in advance. Among them, the text embedding vector is used to be embedded in each level of the corresponding image generation model, and at each level, guide the image generation model to denoise the input image to obtain the final target image.

[0087] As an optional embodiment, determining the image generation model corresponding to the text embedding vector includes:

[0088] When the target fineness of the image style is high fineness, determine the base model corresponding to the base text embedding vector and the refinement model corresponding to the refined text embedding vector;

[0089] The cascading of the at least two image generation models and using the cascaded at least two image generation models to transform the text embedding vector to obtain the target image includes:

[0090] When the target fineness of the image style is high fineness, embed the base text embedding vector into the base model and embed the refined text embedding vector into the refinement model;

[0091] Obtain a randomly generated noise image;

[0092] Use the base model embedded with the base text embedding vector to perform denoising processing on the noise image to obtain a first latent feature image;

[0093] Use the refinement model embedded with the refined text embedding vector to perform denoising processing on the first latent feature image to obtain a second latent feature image;

[0094] Perform decoding processing on the second latent feature image to obtain the target image.

[0095] In this embodiment, since when the style of the target image is a realistic style, at this time, the target image has high requirements for image quality and image details, so not only the basic semantics of the prompt need to be encoded into the base text embedding vector, but also the detailed semantics of the prompt need to be encoded into the refined text embedding vector.

[0096] After obtaining the basic text embedding vector and the refined text embedding vector through encoding, similarly, the basic model corresponding to the basic text embedding vector and the refined model corresponding to the refined text embedding vector can be further determined. Then, a randomly generated noise image is also obtained, and the noise image and the basic text embedding vector are input into the basic model. The basic text embedding vector is used to guide the basic model to perform multiple iterations of noise reduction processing on the noise image, resulting in a first latent feature image.

[0097] Since the first latent feature image is only an image in the feature space that conforms to the basic features of the prompt, and the target image of the realistic style has high requirements for details, therefore, the first latent feature image and the refined text embedding vector can be further input into the refined model. The refined text embedding vector is used to guide the refined model to perform further multiple iterations of noise reduction processing on the first latent feature image, resulting in a second latent feature image. The target image that can be recognized by the human eye can be obtained by decoding the second latent feature image using the variational autoencoder model.

[0098] Among them, the refined model can be the Refiner model in the Stable Diffusion model, which is used to refine the image with insufficient details and enrich the details of the image. The second latent feature image is an image in the feature space that conforms to the detail features of the prompt.

[0099] As an alternative embodiment, such as Figure 2As shown in the figure, when the style of the target image is a realistic style, the positive prompt and negative prompt can be used as the input of the multi-modal large model CLIP. The CLIP model can use the text encoder to parse the prompt words and convert them into basic text embedding vectors (Text Embeddings) and refined text embedding vectors (Refiner Text Embeddings). In addition, a Gaussian noise GaussianNoise can be initialized according to the random number seed Seed and converted into a noise image Latents. Then, the basic text embedding vector and the noise image are input into the generative large model UNet. Through multiple iterations of the basic model (basemodel) in the UNet model by the first sampler Sampler1, the first latent feature image ConditionedLatents1 is output. Then, the basic text embedding vector and the noise image are input into the generative large model UNet. Through multiple iterations of the refined model (Refinermodel) in the UNet model by the second sampler Sampler2, the second latent feature image Conditioned Latents2 is output. The decoder in the variational autoencoder VAE is used to decode the second latent feature image to obtain the target image (output image) that can be recognized by the human eye.

[0100] In this embodiment, for realistic style images with high detail requirements, the basic model and the refined model can be cascaded twice. The basic model is used to complete the basic image generation task, and the refined model is used to refine the image obtained by the basic model, so as to ensure the image generation quality.

[0101] As an optional embodiment, determining the image generation model corresponding to the text embedding vector includes:

[0102] In the case where the attribute of the image style is the first type of style attribute, determining the basic model corresponding to the basic text embedding vector;

[0103] When the image generation model is one, using the image generation model to convert the text embedding vector to obtain the target image includes:

[0104] In the case where the attribute of the image style is the first type of style attribute, embedding the basic text embedding vector into the basic model;

[0105] Obtaining a randomly generated noise image;

[0106] Using the base model that embeds the base text embedding vector to perform noise reduction processing on the noise image to obtain a first latent feature image;

[0107] Performing decoding processing on the first latent feature image to obtain the target image.

[0108] In this embodiment, since when the style of the target image is an artistic style, the target image at this time focuses more on the display of the painting style and has low requirements for the details of the image, the basic semantics of the prompt can be encoded as the base text embedding vector only.

[0109] After encoding the base text embedding vector, the base model corresponding to the base text embedding vector can be further determined, and then a randomly generated noise image is obtained. The noise image and the base text embedding vector are input into the base model, and the base text embedding vector is used to guide the base model to perform iterative noise reduction processing on the noise image multiple times to obtain a first latent feature image. The first latent feature image is an image that conforms to the basic features of the prompt in the feature space, and the target image that can be recognized by the human eye can be obtained by using the self-variational encoder model to perform decoding processing on the first latent feature image.

[0110] Among them, the base model can be the base model in the latent diffusion model (Stable Diffusion model), which is used to convert text content into an image with insufficient details.

[0111] As an alternative embodiment, as Figure 3 shown, when the style of the target image is an artistic style, the positive prompt Prompt and the negative prompt Negative Prompt can be used as the input of the multi-modal large model CLIP. The CLIP model can use the text encoder to parse the prompt and convert it into the base text embedding vector (TextEmbeddings). In addition, a Gaussian noise GaussianNoise can be initialized according to the random number seed Seed and converted into a noise image Latents, and the base text embedding vector and the noise image are input into the generation large model UNet. Through multiple iterations of the base model in the UNet model by the sampler Sampler, the first latent feature image Conditioned Latents is output, and then the decoder in the self-variational encoder VAE is used to perform decoding processing on the first latent feature image to obtain the target image (output image) that can be recognized by the human eye.

[0112] In this embodiment, for artistic style images with low detail requirements, the basic image generation task can be completed only through the basic model, thus effectively reducing the power consumption of algorithm inference while meeting the requirements.

[0113] In the embodiment of the present application, by performing semantic recognition on the prompt words input by the user, determining the image style of the target image corresponding to the prompt words according to the result of the semantic recognition, then encoding the prompt words according to the image style, correspondingly obtaining a text embedding vector, and converting the text embedding vector into a target image by the inference process corresponding to the text embedding vector. In this way, the text embedding vector can be automatically selected according to the different image styles, so as to automatically select the algorithm process adapted to the image style for inference, realizing the reasonable allocation of resources. Compared with the prior art, it avoids applying a set of standardized processes with high power consumption to all text-to-image tasks, and reduces the cost of image generation.

[0114] As an optional embodiment, the performing semantic recognition on the prompt words and determining the image style of the target image according to the result of the semantic recognition includes:

[0115] Performing semantic recognition on the prompt words, and when there is a specific phrase in the prompt words, determining the image style of the target image as the image style corresponding to the specific phrase, where the specific phrase is used to indicate the image style of the prompt words;

[0116] When there is no such specific phrase in the prompt words, determining the image style of the target image as the default style.

[0117] In this embodiment, the obtained prompt words can be semantically understood to obtain semantic information representing the user's intention or requirement. Then, based on the semantic information, it is determined whether there is a specific phrase indicating the image style in the prompt words. If there is a specific phrase and the user's requirement is to determine the style of the target image as the image style indicated by the specific phrase, then the image style of the target image can be determined as the image style indicated by the specific phrase; if there is no specific phrase, then the image style of the target image can be default determined as the default style, and the default style can be a realistic style. Among them, the specific phrase can include realistic, anime, oil painting, ink painting, sketch, watercolor, and cartoon, etc.

[0118] Exemplarily, if the prompt is "Generate a picture of a puppy" and this prompt does not include any specific phrases, then the image style of the target image corresponding to this prompt can be determined as a realistic style; if the prompt is "Generate a picture of a puppy in anime style" and this prompt includes the specific phrase "anime", and based on the result of semantic understanding, the user's requirement is to determine the target image as an anime style, then the image style of the target image corresponding to this prompt can be determined as an anime style, and the anime style belongs to the art style.

[0119] Through this embodiment, it is possible to accurately and quickly determine the image style of each target image to be generated based on the prompt.

[0120] As an optional embodiment, after determining the image generation model corresponding to the text embedding vector according to the model correspondence, the method further includes:

[0121] By performing semantic recognition on the prompt, determine the main type of the prompt according to the result of semantic recognition;

[0122] Determine the sampler corresponding to the image style and the main type;

[0123] Load the sampler into the image generation model.

[0124] In this embodiment, the sampler is used to control the intensity, parameters, and number of iterations of noise reduction during the iterative process of the model's repeated noise reduction. Before applying the model for image noise reduction, it is necessary to pre-load the selected sampler into the model.

[0125] Specifically, there are various types of samplers, including DDIM, Euler, DPM++2M, etc., and different samplers have different characteristics. During the selection process of the sampler, semantic recognition can be performed on the prompt to determine the main type of the prompt and the image style of the target image, and then according to the pre-set mapping relation table, query the sampler that has a mapping relation with this image style and main type in the mapping relation table, and load this sampler into all the above models.

[0126] In this embodiment, since different main types and image styles have different requirements for the target image, a suitable sampler can be selected based on the main type and image style to control the target image generation process, so that different target images are adapted to different samplers.

[0127] As an optional embodiment, the obtaining of the noise image includes:

[0128] Obtain a pre-set random number seed;

[0129] Convert the random number seed into Gaussian noise;

[0130] Encode the Gaussian noise to obtain the noise image.

[0131] In this embodiment, Gaussian distributed random numbers can be generated in a deterministic manner by obtaining a pre-set random number seed (seed), and these random numbers constitute Gaussian noise. After obtaining the Gaussian noise, the Gaussian noise can be mapped to the latent space through an encoder to obtain the latent Gaussian noise, that is, the noise image.

[0132] In this embodiment, the noise image generated each time the code is run can be ensured to be reproducible by a given random number seed.

[0133] Based on the image generation method provided in the above embodiment, correspondingly, the present application also provides a specific implementation manner of the image generation device. Please refer to the following embodiments.

[0134] First, refer to Figure 4 , the image generation device 400 provided in the embodiment of the present application includes the following modules:

[0135] An acquisition module 401, configured to acquire a prompt word input by a user, where the prompt word is used to indicate generating a target image;

[0136] An identification module 402, configured to determine the image style of the target image according to the result of semantic recognition by performing semantic recognition on the prompt word;

[0137] An encoding module 403, configured to encode the prompt word according to the encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style, where different image styles correspond to different encoding methods;

[0138] A conversion module 404, configured to convert the text embedding vector according to the inference process corresponding to the text embedding vector to obtain the target image.

[0139] The device can perform semantic recognition on the prompt word input by the user, determine the image style of the target image corresponding to the prompt word according to the result of semantic recognition, then encode the prompt word according to the image style, and correspondingly obtain a text embedding vector, and convert the text embedding vector into a target image by the inference process corresponding to the text embedding vector. In this way, the text embedding vector can be automatically selected according to different image styles, so as to automatically select an algorithm process adapted to the image style for inference, realize reasonable allocation of resources, and compared with the prior art, avoid applying a set of standardized processes with high power consumption to all text-to-image tasks, and reduce the cost of image generation.

[0140] As an implementation manner of the present application, the above encoding module 403 may further include:

[0141] A first determination unit, configured to determine the target fineness of the target image according to the attributes of the image style;

[0142] A first query unit, configured to query the encoding method corresponding to the target fineness according to the pre-set encoding correspondence;

[0143] A first encoding unit, configured to encode the prompt according to the encoding method corresponding to the target fineness to obtain a text embedding vector corresponding to the image style.

[0144] As an implementation manner of the present application, the above first determination unit may further include:

[0145] A first determination subunit, configured to determine the target fineness of the target image as a low fineness when the attribute of the image style is a first type of style attribute;

[0146] A second determination subunit, configured to determine the target fineness of the target image as a high fineness when the attribute of the image style is a second type of style attribute.

[0147] As an implementation manner of the present application, the above first encoding unit may further include:

[0148] A first encoding subunit, configured to encode the prompt to obtain a basic text embedding vector and a refined text embedding vector when the target fineness of the target image is a high fineness, where the basic text embedding vector is used to represent the basic semantics of the prompt, and the refined text embedding vector is used to represent the detailed semantics of the prompt;

[0149] A second encoding subunit, configured to encode the prompt to obtain a basic text embedding vector when the target fineness of the target image is a low fineness.

[0150] As an implementation manner of the present application, the above conversion module 404 may further include:

[0151] A second determination unit, configured to determine the image generation model corresponding to the text embedding vector according to the model correspondence;

[0152] A first conversion unit, configured to use the image generation model to convert the text embedding vector to obtain the target image when the image generation model is one;

[0153] A second conversion unit, configured to cascade the at least two image generation models when the image generation model includes at least two, and convert the text embedding vector by using the cascaded at least two image generation models to obtain the target image.

[0154] As an implementation manner of this application, the above second determination unit may further include:

[0155] A third determination subunit, configured to determine the base model corresponding to the base text embedding vector and the refinement model corresponding to the refined text embedding vector when the target fineness of the image style is a high fineness;

[0156] The above second conversion unit may further include:

[0157] An embedding subunit, configured to embed the base text embedding vector into the base model and embed the refined text embedding vector into the refinement model when the target fineness of the image style is a high fineness;

[0158] A generation subunit, configured to obtain a randomly generated noise image;

[0159] A first noise reduction subunit, configured to perform noise reduction processing on the noise image by using the base model embedding the base text embedding vector to obtain a first latent feature image;

[0160] A second noise reduction subunit, configured to perform noise reduction processing on the first latent feature image by using the refinement model embedding the refined text embedding vector to obtain a second latent feature image;

[0161] A decoding subunit, configured to perform decoding processing on the second latent feature image to obtain the target image.

[0162] As an implementation manner of this application, the above recognition module 402 may further include:

[0163] A first recognition unit, configured to perform semantic recognition on the prompt word, and determine the image style of the target image as the image style corresponding to the specific phrase when the specific phrase exists in the prompt word, where the specific phrase is used to indicate the image style of the prompt word;

[0164] A second recognition unit, configured to determine the image style of the target image as a realistic style when the specific phrase does not exist in the prompt word.

[0165] The image generation device provided by the embodiments of the present invention can implement each step in the above method embodiments. To avoid repetition, it will not be elaborated here.

[0166] Figure 5 The figure shows a schematic diagram of the hardware structure of the image generation device provided by the embodiments of the present application.

[0167] The image generation device may include a processor 501 and a memory 502 storing computer program instructions.

[0168] Specifically, the above-mentioned processor 501 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.

[0169] The memory 502 may include a mass storage for data or instructions. By way of example and not limitation, the memory 502 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 502 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 502 may be internal or external to the integrated gateway disaster recovery device. In a specific embodiment, the memory 502 is a non-volatile solid-state memory.

[0170] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of the present disclosure.

[0171] The processor 501 reads and executes the computer program instructions stored in the memory 502 to implement any one of the image generation methods in the above embodiments.

[0172] In one example, the image generation device may further include a communication interface 503 and a bus 510. Among them, as Figure 5 shown, the processor 501, the memory 502, and the communication interface 503 are connected through the bus 510 to complete the communication with each other.

[0173] The communication interface 503 is mainly used to implement the communication between the modules, devices, units, and / or devices in the embodiments of the present application.

[0174] The bus 510 includes hardware, software, or both, and couples components of the image generation device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, the bus 510 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0175] The image generation device may be based on the above embodiments, so as to implement the image generation method and apparatus in combination with the above.

[0176] In addition, in combination with the image generation method in the above embodiments, an embodiment of the present application may provide a computer storage medium to implement. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the image generation methods in the above embodiments is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be described in detail here. Among them, the above computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc., which is not limited here.

[0177] In addition, an embodiment of the present application also provides a vehicle, including computer program instructions, and when the computer program instructions are executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0178] It should be clear that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0179] The functional blocks shown in the above structural block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0180] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps. That is, the steps can be executed in the order mentioned in the embodiments, can be different from the order in the embodiments, or several steps can be executed simultaneously.

[0181] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses, and vehicles according to embodiments of the present disclosure. It should be understood that each block in the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0182] The above is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. An image generation method, characterized in that, The method includes: Obtaining a prompt word input by a user, where the prompt word is used to indicate the generation of a target image; Determining the image style of the target image according to the result of semantic recognition by performing semantic recognition on the prompt word; Encoding the prompt word according to the encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style, where different image styles correspond to different encoding methods; Converting the text embedding vector according to the inference process corresponding to the text embedding vector to obtain the target image.

2. The image generation method according to claim 1, characterized in that, The encoding the prompt word according to the encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style includes: Determining the target fineness of the target image according to the attributes of the image style; Querying the encoding method corresponding to the target fineness according to the pre-set encoding correspondence; Encoding the prompt word according to the encoding method corresponding to the target fineness to obtain a text embedding vector corresponding to the image style.

3. The image generation method according to claim 2, characterized in that, The determining the target fineness of the target image according to the attributes of the image style includes: When the attribute of the image style is the first type of style attribute, determining the target fineness of the target image as low fineness; When the attribute of the image style is the second type of style attribute, determining the target fineness of the target image as high fineness.

4. The image generation method according to claim 2 or 3, characterized in that, The encoding the prompt word according to the encoding method corresponding to the target fineness to obtain a text embedding vector corresponding to the image style includes: When the target fineness of the target image is high fineness, encoding the prompt word to obtain a basic text embedding vector and a refined text embedding vector, where the basic text embedding vector is used to represent the basic semantics of the prompt word, and the refined text embedding vector is used to represent the detailed semantics of the prompt word; When the target fineness of the target image is low fineness, encoding the prompt word to obtain a basic text embedding vector.

5. The image generation method according to claim 1, characterized in that, The converting the text embedding vector according to the inference process corresponding to the text embedding vector to obtain the target image includes: Determining the image generation model corresponding to the text embedding vector according to the model correspondence; When the image generation model is one, using the image generation model to convert the text embedding vector to obtain the target image; When the image generation model includes at least two, cascading the at least two image generation models and using the cascaded at least two image generation models to convert the text embedding vector to obtain the target image.

6. The image generation method according to claim 5, characterized in that, The determining the image generation model corresponding to the text embedding vector includes: When the target fineness of the image style is high fineness, determining the basic model corresponding to the basic text embedding vector and the refined model corresponding to the refined text embedding vector; Cascading the at least two image generation models and using the cascaded at least two image generation models to transform the text embedding vector to obtain the target image, including: When the target fineness of the image style is high fineness, embedding the basic text embedding vector into the basic model and embedding the refined text embedding vector into the refined model; Obtaining a randomly generated noise image; Using the basic model embedded with the basic text embedding vector to perform noise reduction processing on the noise image to obtain a first latent feature image; Using the refined model embedded with the refined text embedding vector to perform noise reduction processing on the first latent feature image to obtain a second latent feature image; Performing decoding processing on the second latent feature image to obtain the target image.

7. The image generation method according to claim 1, characterized in that, Determining the image style of the target image by performing semantic recognition on the prompt, including: Performing semantic recognition on the prompt, and when there is a specific phrase in the prompt, determining the image style of the target image as the image style corresponding to the specific phrase, where the specific phrase is used to indicate the image style of the prompt; When the specific phrase does not exist in the prompt, determining the image style of the target image as the default style.

8. An image generation device, characterized in that, The device includes: An acquisition module, configured to acquire a prompt input by a user, where the prompt is used to indicate the generation of a target image; A recognition module, configured to determine the image style of the target image by performing semantic recognition on the prompt according to the result of the semantic recognition; An encoding module, configured to encode the prompt according to the encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style, where different image styles correspond to different encoding methods; A conversion module, configured to transform the text embedding vector according to the inference process corresponding to the text embedding vector to obtain the target image.

9. An image generation device, characterized in that, The image generation device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image generation method according to any one of claims 1-7 is implemented.

10. A computer storage medium, characterized in that, Computer program instructions are stored on a computer storage medium, and when the computer program instructions are executed by a processor, the image generation method according to any one of claims 1-7 is implemented.

11. A vehicle, characterized in that, The vehicle includes at least one of the following: The image generation device according to claim 8; The image generation device according to claim 9; The computer storage medium according to claim 10.