Image generation method and apparatus, device, storage medium, and vehicle

By text expansion of the prompt words entered by the user and selecting a suitable image generation model and sampler based on the prompt words, the problems of low image generation quality and high cost in the prior art are solved, and higher quality and lower cost image generation are achieved.

WO2025130927A1PCT designated stage expired Publication Date: 2025-06-26BEIJING CO WHEELS TECH CO LTD

Patent Information

Application Number
PCT/CN2024/140323
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-12-18
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the prior art, the image generation model processes simple prompt words to low image quality, and the fixed model and sampler cannot adapt to different content generation requirements, resulting in high image generation costs.

Method used

By obtaining the prompt words entered by the user, text augmentation is performed to enrich content expression, select appropriate image generation models and samplers based on the prompt words, and convert the prompt words into target images using different samplers and models.

Benefits of technology

This improves the accuracy and richness of image generation, reduces the cost of image generation, and avoids the low quality and high cost problems caused by the fixed use of image generation models and samplers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140323_26062025_PF_FP_ABST
    Figure CN2024140323_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses an image generation method and apparatus, a device, a storage medium, and a vehicle. The method comprises: obtaining a prompt word input by a user, wherein the prompt word is used for instructing to generate a target image; on the basis of a preset first correspondence, determining an image generation model and a sampler corresponding to the prompt word, wherein the first correspondence is a correspondence among the prompt word, the image generation model and the sampler, and the sampler is used for controlling a noise reduction mode of the image generation model; performing text expansion on the prompt word to obtain a final text; loading the sampler into the image generation model to update noise reduction parameters of the image generation model; and using the image generation model with updated noise reduction parameters to convert the final text into the target image. According to the embodiments of the present application, the quality of generated images can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation method, device, equipment, storage medium and vehicle

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent applications filed with the Patent Office of China on December 19, 2023, with application number 202311755370.1, entitled “Image generation method, device, equipment, storage medium and vehicle”, filed with the Patent Office of China on December 19, 2023, with application number 202311755363.1, entitled “Image generation method, device, equipment, storage medium and vehicle”, filed with the Patent Office of China on December 19, 2023, with application number 202311757404.0, entitled “Image generation method, device, equipment, storage medium and vehicle”, and filed with the Patent Office of China on December 19, 2023, with application number 202311756160.4, entitled “Model mounting method, device, equipment, storage medium and vehicle”, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application belongs to the field of artificial intelligence technology, and in particular relates to an image generation method, apparatus, device, storage medium and vehicle. Background Art

[0004] A text-to-image generation task aims to convert text descriptions or natural language text into corresponding images. In this task, the computer model needs to understand the prompt words input by the user and generate an image that matches the prompt words.

[0005] However, in related technologies, all text-to-image tasks rely on the same standard algorithmic process for inference. Specifically, the user's prompt word is input into a pre-set image generation model, which then uses the same algorithm to convert the prompt word into an image that matches the prompt word. Because the prompt word entered by the user is often simple and cannot fully express the user's personalized needs, the images generated by inferring the concise prompt word using the same standard algorithmic process are often monotonous. As a result, the generated image fails to accurately express the user's needs, resulting in low-quality image generation.

[0006] Furthermore, in related technologies, all text-based image tasks rely on the same standard algorithmic process for inference. This process first directly generates a highly detailed target image using the text-based image model. However, this standard inference process, which directly generates a highly detailed target image, consumes significant power. This means that each text-based image task in related technologies requires significant power consumption, resulting in excessively high image generation costs.

[0007] Furthermore, in the related art, a fixed image generation model is set before all cultural image tasks, and a fixed sampler is loaded in the image generation model. Then, for each cultural image task that requires it, the fixed sampler in the same image generation model is applied to convert the drawing instructions into the target image. However, since different cultural image tasks actually have different content generation requirements, and each image generation model and sampler can only meet the content generation requirements of one type of image content, then each cultural image task in the related art uses a pre-set fixed image generation model and a fixed sampler, which often leads to the problem that the fixed image generation model and sampler cannot adapt to the content generation requirements of the cultural image task, resulting in lower quality of the generated image.

[0008] In addition, in related technologies, after different LoRA models are mounted on the main model to obtain a combined model, an application programming interface is usually used to dynamically adjust the weight parameters of the LoRA models in different combined models. However, the adjustment process using the application programming interface often causes serious performance regression, thereby reducing the efficiency of the model operation. Summary of the Invention

[0009] The embodiments of the present application provide an image generation method, apparatus, device, storage medium, and vehicle, which can solve the problem of low quality of images obtained by converting existing prompt words.

[0010] In a first aspect, some embodiments of the present application provide an image generation method, the method comprising:

[0011] Obtaining a prompt word input by the user, the prompt word is used to indicate the generation of a target image;

[0012] Determining an image generation model and a sampler corresponding to the prompt word according to a pre-set first correspondence relationship, wherein the first correspondence relationship is a correspondence relationship between the prompt word, the image generation model, and the sampler, and the sampler is used to control a denoising method of the image generation model;

[0013] Perform text expansion on the prompt words to obtain the final text;

[0014] Load the sampler into the image generation model to update the denoising parameters of the image generation model;

[0015] The final text is converted into the target image using the image generation model after the denoising parameters are updated.

[0016] In a second aspect, some embodiments of the present application provide an image generating device, comprising:

[0017] A first acquisition module is used to acquire a prompt word input by a user, where the prompt word is used to instruct the generation of a target image;

[0018] A first determination module is configured to determine an image generation model and a sampler corresponding to a prompt word based on a preset first correspondence relationship, wherein the first correspondence relationship is a correspondence relationship between the prompt word, the image generation model, and the sampler, and the sampler is configured to control a denoising method of the image generation model;

[0019] The expansion module is used to expand the prompt word to obtain the final text;

[0020] An update module, used to load the sampler into the image generation model to update the denoising parameters of the image generation model;

[0021] The first conversion module is used to convert the final text into a target image using the image generation model after the noise reduction parameters are updated.

[0022] In a third aspect, some embodiments of the present application provide a text generation device, the device comprising: a processor and a memory storing computer program instructions;

[0023] When the processor executes the computer program instructions, the above image generation method is implemented.

[0024] In a fourth aspect, some embodiments of the present application provide a computer storage medium having computer program instructions stored thereon, which implement the above-mentioned image generation method when executed by a processor.

[0025] In a fifth aspect, some embodiments of the present application provide a vehicle, which includes computer program instructions, and when the computer program instructions are executed by a processor, the above image generation method is implemented.

[0026] In some embodiments of the present application, after receiving a prompt word input by the user, the prompt word can be expanded to obtain a final text, thereby enriching the user's needs expressed in the text content, and based on the prompt word, an image generation model and sampler are selected, and different prompt words are converted into target images using different samplers and models, and the target image is super-resolved to obtain the final target image. In this way, the final text can be obtained by expanding the text content of the prompt word, and a suitable model and sampler can be selected based on the prompt word, and the final text can be converted into the target image using the selected sampler. In the above process, the accuracy and richness of the generated image can be improved by selecting the sampler and model and expanding the text, thereby improving the quality of the generated image.

[0027] The embodiments of the present application provide an image generation method, apparatus, device, storage medium, and vehicle, which can solve the problem of high cost of existing image generation.

[0028] In a sixth aspect, some embodiments of the present application provide an image generation method, the method comprising:

[0029] Obtaining a prompt word input by a user, wherein the prompt word is used to instruct generation of a target image;

[0030] performing semantic recognition on the prompt word and determining the image style of the target image according to the result of the semantic recognition;

[0031] Encoding the prompt word according to the encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style, wherein different image styles correspond to different encoding methods;

[0032] The text embedding vector is converted according to the inference process corresponding to the text embedding vector to obtain the target image.

[0033] In a seventh aspect, some embodiments of the present application provide an image generating device, the device comprising:

[0034] a fourth acquisition module, configured to acquire a prompt word input by a user, wherein the prompt word is used to instruct generation of a target image;

[0035] a recognition module, configured to perform semantic recognition on the prompt word and determine the image style of the target image according to the result of the semantic recognition;

[0036] an encoding module, configured to encode the prompt word according to an encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style, wherein different image styles correspond to different encoding methods;

[0037] The second conversion module is used to convert the text embedding vector according to the inference process corresponding to the text embedding vector to obtain the target image.

[0038] In an eighth aspect, some embodiments of the present application provide an image generating device, the device comprising: a processor and a memory storing computer program instructions;

[0039] When the processor executes the computer program instructions, the above image generation method is implemented.

[0040] In a ninth aspect, some embodiments of the present application provide a computer storage medium having computer program instructions stored thereon, which implement the above-mentioned image generation method when the computer program instructions are executed by a processor.

[0041] In a tenth aspect, some embodiments of the present application provide a vehicle, comprising computer program instructions, which implement the above image generation method when executed by a processor.

[0042] In some embodiments of the present application, semantic recognition is performed on a prompt word input by the user. The style of the target image corresponding to the prompt word is determined based on the semantic recognition results. The prompt word is then encoded based on the image style, and a corresponding text embedding vector is obtained. The text embedding vector is then converted into the target image using the inference process corresponding to the text embedding vector. In this way, the text embedding vector can be automatically selected based on the different image styles, and the algorithm process adapted to the image style can be automatically selected for inference, achieving reasonable resource allocation. Compared with the existing technology, this avoids the application of a single set of power-intensive standardized processes for all text-to-image tasks, reducing the cost of image generation.

[0043] The embodiments of the present application provide an image generation method, apparatus, device, storage medium, and vehicle, which can solve the problem of low quality of images generated by existing fixed models and samplers.

[0044] In an eleventh aspect, some embodiments of the present application provide an image generation method, the method comprising:

[0045] Acquire a drawing instruction input by a user, where the drawing instruction is used to instruct generation of a target image;

[0046] Analyzing the drawing instruction to determine the content category of the target image corresponding to the drawing instruction;

[0047] Determining, from a sampler database according to the content category, a sampler algorithm and parameters corresponding to the content category, wherein the sampler database includes a correspondence between content categories and sampler algorithms, and parameters corresponding to each sampler algorithm;

[0048] The drawing instructions are converted into the target image using the sampler algorithm and parameters.

[0049] In a twelfth aspect, some embodiments of the present application provide an image generating device, the device comprising:

[0050] a fifth acquisition module, configured to acquire a drawing instruction input by a user, wherein the drawing instruction is used to instruct generation of a target image;

[0051] an analysis module, configured to analyze the drawing instruction and determine a content category of a target image corresponding to the drawing instruction;

[0052] a third determining module, configured to determine, from a sampler database according to the content category, a sampler algorithm and parameters corresponding to the content category, wherein the sampler database includes a correspondence between content categories and sampler algorithms, and parameters corresponding to each sampler algorithm;

[0053] The third conversion module is configured to convert the drawing instruction into the target image by using the sampler algorithm and parameters.

[0054] In a thirteenth aspect, some embodiments of the present application provide an image generating device, the device comprising: a processor and a memory storing computer program instructions;

[0055] When the processor executes the computer program instructions, the above image generation method is implemented.

[0056] In a fourteenth aspect, some embodiments of the present application provide a computer storage medium having computer program instructions stored thereon, which implement the above-mentioned image generation method when the computer program instructions are executed by a processor.

[0057] In a fifteenth aspect, some embodiments of the present application provide a vehicle, comprising the above-mentioned image generating device, image generating apparatus and computer storage medium.

[0058] In some embodiments of the present application, the drawing instructions input by the user are analyzed, and the content category of the target image corresponding to the drawing instructions is determined based on the analysis results. Then, the sampler algorithm and parameters corresponding to the content category are queried in the sampler database based on the content category, and the drawing instructions are converted into the target image using the sampler algorithm and parameters. In this way, the conversion from the drawing instructions to the target image can be completed by automatically selecting the appropriate sampler algorithm and parameters based on the content category of the drawing instructions. Compared with the prior art, the present application dynamically selects samplers with different characteristics based on the attributes of the drawing instructions for different Wensheng drawing tasks, avoiding the standardized process of applying one sampler to all Wensheng drawing tasks. Different samplers are adapted to different Wensheng drawing tasks, thereby improving the quality of image generation.

[0059] The embodiments of the present application provide a model mounting method, apparatus, device, storage medium and vehicle, which can solve the problem of low efficiency of model operation caused by existing model mounting methods.

[0060] In a sixteenth aspect, some embodiments of the present application provide a model mounting method, the method comprising:

[0061] Obtaining a model mounting instruction, wherein the model mounting instruction is used to instruct to mount the target mounting model to the application main model;

[0062] In response to the model mounting instruction, the target mounting model is mounted to the application main model to obtain an application combination model, and the model weight in the target mounting model is adjusted to a null value;

[0063] Obtain a target weight file corresponding to the target mounting model, input the target weight file into the target mounting model in the application combination model, and update the null value to the target weight parameter of the target weight file.

[0064] In a seventeenth aspect, some embodiments of the present application provide a model mounting device, the device comprising:

[0065] a sixth acquisition module, configured to acquire a model mounting instruction, wherein the model mounting instruction is used to instruct the target mounting model to be mounted to the application main model;

[0066] a second mounting module, configured to mount the target mounting model to the application main model in response to the model mounting instruction to obtain an application combination model, and adjust the model weight in the target mounting model to a null value;

[0067] An input module is used to obtain a target weight file corresponding to the target mounting model, input the target weight file into the target mounting model in the application combination model, and update the null value to the target weight parameter of the target weight file.

[0068] In an eighteenth aspect, some embodiments of the present application provide a model mounting device, the device comprising: a processor and a memory storing computer program instructions;

[0069] The processor implements the above model mounting method when executing computer program instructions.

[0070] In the nineteenth aspect, some embodiments of the present application provide a computer storage medium having computer program instructions stored thereon, which implement the above-mentioned model mounting method when the computer program instructions are executed by a processor.

[0071] In aspect 20, some embodiments of the present application provide a vehicle, comprising computer program instructions, which implement the above-mentioned model mounting method when executed by a processor.

[0072] In some embodiments of the present application, when the mount model is mounted on the main model for application, the model weight in the target mount model can be adjusted to a null value, and the target weight file corresponding to the target mount model is input into the target mount model in the application combination model to update the null value to the target weight parameter of the target weight file. In this way, during the application process of the target mount model, the weight that you want to give to the target mount model can be input into the target mount model in the form of input, so as to realize the real-time update of the weight of the target mount model. Compared with the related art, this method of inputting the model weight into the mount model in the form of a file can improve the processing speed compared to calling the interface for weight update, and will not cause the processing speed performance of the application combination model to fall back after the mount model is mounted, thereby accurately and efficiently realizing the dynamic replacement of the weight parameters of the mount model and improving the operating efficiency of the combination model. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0074] FIG1 is a schematic flow chart of an image generation method provided in some embodiments of the present application;

[0075] FIG2 is a schematic structural diagram of an image generating device provided in some embodiments of the present application;

[0076] FIG3 is a flow chart of another image generation method provided in some embodiments of the present application;

[0077] FIG4 is a flowchart of another image generation method provided in some embodiments of the present application;

[0078] FIG5 is a flowchart of another image generation method provided in some embodiments of the present application;

[0079] FIG6 is a schematic structural diagram of another image generating device provided in some embodiments of the present application;

[0080] FIG7 is a flowchart of another image generation method provided in some embodiments of the present application;

[0081] FIG8 is a schematic structural diagram of another image generating device provided by some embodiments of the present application;

[0082] FIG9 is a flow chart of a model mounting method provided in some embodiments of the present application;

[0083] FIG10 is a schematic structural diagram of a model mounting device provided in some embodiments of the present application;

[0084] FIG11 is a schematic diagram of the hardware structure of an example device provided in some embodiments of the present application. DETAILED DESCRIPTION

[0085] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.

[0086] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0087] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The embodiments will be described in detail below with reference to the accompanying drawings.

[0088] Specifically, in order to solve the problems of the prior art, some embodiments of the present application provide an image generation method, apparatus, device, storage medium, and vehicle. The image generation method provided in the embodiments of the present application is first introduced below.

[0089] Figure 1 shows a flow chart of an image generation method provided by some embodiments of the present application. The method can be applied to a vehicle's onboard computer or a cloud server connected to the vehicle, and includes the following steps:

[0090] S101: Obtain a prompt word input by a user, where the prompt word is used to instruct generation of a target image.

[0091] In some embodiments, the image generation task can convert a prompt word input by the user into a target image, where the prompt word is used to describe the content of the target image that the user wants to generate. For example, the prompt word can be "generate a puppy" or "generate a tree."

[0092] Furthermore, the input prompt words can include positive prompt words and negative prompt words. Positive prompt words indicate desired content in the target image, while negative prompt words indicate undesirable content in the target image. For example, a positive prompt word could be "A cartoon-style Chinese girl running on the beach, with flying seagulls and a gorgeous rainbow behind her, the overall picture is poetic and picturesque," while a negative prompt word could be "pornographic, nude, ugly, deformed."

[0093] S102, determining the image generation model and sampler corresponding to the prompt word according to a pre-set first correspondence relationship, where the first correspondence relationship is the correspondence between the prompt word, the image generation model, and the sampler, and the sampler is used to control the denoising method of the image generation model.

[0094] In some embodiments, the prompt word can be input into the image generation model, and the image generation model loaded with the sampler converts the prompt word into an image.

[0095] In some embodiments, the image generation model may include a basic image generation model and a specific object generation model. The basic image generation model may be a base model in a stable diffusion image generation model, which may convert text content into a basic image, and the specific object generation model may be a low-rank adaptation of large language models (LoRA image generation model) in a stable diffusion image generation model, which may be used to generate images of certain specific styles or specific subjects. The specific object generation model cannot be used alone, and the specific object generation model needs to be mounted on the basic image generation model to assist the basic image generation model in achieving image generation.

[0096] In this case, both the basic image generation model and the specific object generation model can be UNET-structured models. During the training process, the training set can be multiple clear images. The training process for each image mainly consists of two parts: denoising and denoising. The denoising process continuously applies noise to a clear initial image. The denoising process uses the basic image generation model and the specific object generation model to infer the noisy image, predict the noise, and remove it until the restored image is restored. Then, by comparing the initial image and the restored image, the model loss function is calculated and adjusted to complete the model training.

[0097] Specifically, the prompt words can be used to guide the image generation model to perform multiple iterative denoising processes on random noise images to obtain the final target image. The sampler is used to load into the image generation model during the iterative denoising process of the image generation model to control the intensity of denoising, denoising method and number of iterations.

[0098] The correspondence between the prompt word, image generation model and sampler can be defined in advance as the first correspondence. Then, whenever a prompt word input by the user is received, the image generation model and sampler corresponding to the prompt word can be queried based on the first correspondence.

[0099] For example, the first correspondence may include a correspondence between the prompt word and the image generation model, and a correspondence between the image generation model and the sampler.

[0100] S103: Expand the prompt word to obtain the final text.

[0101] In some embodiments, the prompt word can be a simple sentence output by the user in the form of voice or text. In order to enrich the content of the image converted by the prompt word, the prompt word can be expanded to obtain a final text with richer content.

[0102] Specifically, phrases or phrases in a pre-set descriptive word library can be added to the prompt words, and the prompt words can be described in detail and extendedly; for example, when the positive prompt word is "a round-faced, chubby and cute tabby cat with an anime style is basking in the sun on a lawn full of flowers", it can be expanded into the text to obtain the final text "a round-faced, chubby and cute tabby cat with an anime style is basking in the sun on a lawn full of flowers, honest and cute, with big eyes, charming scenery, and professional photography techniques".

[0103] S104: Load the sampler into the image generation model to update the denoising parameters of the image generation model.

[0104] In some embodiments, the sampler is loaded into the image generation model, and the sampler will adjust the image generation model's prediction method of the noise on the input noisy image, the intensity of the noise reduction, the noise reduction method and the number of iterations by affecting the noise reduction parameters in the image generation model, thereby affecting the image generated by the image generation model.

[0105] S105, converting the final text into a target image using the image generation model after the denoising parameters are updated.

[0106] In some embodiments, after selecting the sampler and image generation model corresponding to the prompt word and loading the determined sampler into the image generation model, a random noise image can be obtained, and the prompt word can be converted into a text embedding vector, and the text embedding vector can be embedded into the image generation model loaded with the sampler, guiding the image generation model to perform multiple iterations of denoising on the noise image to obtain the target image.

[0107] In an embodiment of the present application, after receiving a prompt word input by the user, the prompt word can be expanded to obtain a final text, thereby enriching the user's needs expressed in the text content, and selecting an image generation model and sampler based on the prompt word, using different samplers and models to convert different prompt words into target images, and super-resolving the target image to obtain the final target image. In this way, the final text can be obtained by expanding the text content of the prompt word, and a suitable model and sampler can be selected based on the prompt word, and the final text can be converted into a target image using the selected sampler. In the above process, the accuracy and richness of the generated image can be improved by selecting the sampler and model and expanding the text, thereby improving the quality of the generated image.

[0108] As an optional embodiment, the image generation model includes a basic image generation model and a specific object generation model, the first correspondence includes a model correspondence and a sampler correspondence, and the above S102 may include:

[0109] Perform semantic recognition on the prompt word to determine the image generation intention corresponding to the prompt word;

[0110] Identify specific generation information contained in the image generation intent corresponding to the prompt word; the specific generation information includes a specific subject and / or a specific style;

[0111] Determining the basic image generation model and the specific object generation model corresponding to the prompt word according to the model correspondence relationship, wherein the model correspondence relationship includes the correspondence relationship between the specific generation information and the specific object generation model, and the correspondence relationship between the image generation intention and the basic image generation model;

[0112] According to the sampler correspondence, the samplers corresponding to the prompt word, the basic image generation model and the specific object generation model are determined. The sampler correspondence includes the correspondence between the image generation intention, the specific generation information, the basic image generation model, the specific object generation model and the sampler.

[0113] In some embodiments, the intent of the prompt word can be identified by semantically understanding the prompt word. Since the user wants to generate a target image based on the prompt word, the image generation intent includes the attributes of the target image to be generated and the elements included in the target image. Specifically, the image generation intent includes information about a specific subject or style included in the target image.

[0114] For example, a specific subject may refer to a specific subject that the user wishes to generate in the image. For example, the specific subject may be a cat, a car, etc. A specific style may refer to a specific style that the user wishes to generate in the image, such as an oil painting style or a watercolor style.

[0115] The model correspondence is the correspondence between the prompt word and the image generation model. After determining the image generation intention and specific generation information corresponding to the prompt word, the basic image generation model corresponding to the image generation intention can be queried based on the model correspondence, and the specific object generation model corresponding to the specific generation information can be queried to determine the image generation model.

[0116] The sampler correspondence is the correspondence between the prompt word, the image generation model and the sampler. Based on the sampler correspondence, the samplers corresponding to the image generation intention, specific generation information, the basic image generation model, and the specific object generation model can be queried to determine the sampler.

[0117] For example, there are multiple different basic image generation models and multiple different specific object generation models. Among them, each basic image generation model is used to implement the subject's cultural image task of one or more vertical domains, and each specific object generation model is used to implement the cultural image task of one or more image styles. Therefore, the vertical domain to which the subject of the prompt word belongs can be determined by the image generation intention, and the basic image generation model corresponding to the vertical domain can be determined based on the model correspondence. Then, the image style of the target image is determined by the specific generation information, and the specific object generation model corresponding to the image style is determined based on the model correspondence. The sampler can be determined by the subject type, image style, and the selected basic image generation model and specific object generation model.

[0118] In this way, the appropriate image generation model and sampler can be accurately selected based on the user intention reflected by the prompt word, thereby improving the accuracy of text conversion.

[0119] As an optional embodiment, determining the samplers corresponding to the prompt word, the basic image generation model, and the specific object generation model according to the sampler correspondence includes:

[0120] Obtaining identification information of an image generation intent, specific generation information, a basic image generation model, and a specific object generation model;

[0121] Determine the identification information as an initialization parameter of the sampler;

[0122] The sampler corresponding to the initialization parameter is matched from the sampler correspondence, where the sampler correspondence includes the correspondence between the initialization parameter and the sampler.

[0123] In some embodiments, the image generation intention may include a subject type, and the specific generation information may include a specific style. A subject type tag set S may be defined in advance. P =[p cls1 ,p cls2 …p clsn ], the elements in the subject type tag set can be used to identify the subject type in the image generation intention. The subject type tags can include animals, plants, people, food, etc. A specific style tag set S can also be defined in advance t =[t cls1 ,t cls2 …t clsn ], each element in the specific style tag set can be used to identify the image style in the specific generated information. The image style of the target image can also be any element in the specific style tag set. In other words, the subject type tag is the identification information of the image generation intention, and the specific style tag is the identification information of the specific generated information.

[0124] In addition, the subject and subject type of the prompt word can be obtained by extracting features from the prompt word and then processing the extracted features through common neural network structures such as pooling, convolution, and residual. Both the model correspondence and the sampler correspondence can be mapping relationships. Through a pre-defined mapping table, the basic image generation model sequence number and the specific object generation model sequence number that have a mapping relationship with the subject type label and the image style label are queried. Among them, the basic image generation model sequence number is the identification information of the basic image generation model, and the specific object generation model sequence number is the identification information of the specific object generation model. Specifically, multiple different basic image generation models and multiple different specific object generation models can be set. Each basic image generation model has a unique corresponding basic image generation model sequence number, and each specific object generation model has a unique corresponding specific object generation model sequence number. The mapping relationship includes the mapping relationship between the subject type label and the basic image generation model sequence number, and the mapping relationship between the image style label and the specific object generation model sequence number.

[0125] After the query is completed, the image generation intent, specific generation information, basic image generation model, and identification information of the specific object generation model can be determined as the initialization parameters of the sampler, and then a new initialization parameter set including the image generation intent, specific generation information, basic image generation model, and identification information of the specific object generation model is generated:

[0126] [Subject type label, specific style label, base image generation model number, specific object generation model number]

[0127] The four elements in the initialization parameter set are all input parameters required for sampler initialization. Then, based on the sampler correspondence, a sampler having a mapping relationship with the initialization parameter set is searched in a predefined mapping table.

[0128] For example, for an input with an initialization parameter set of input1 = [character, realism, character model, character LORA], a DPM++ sampler with a mapping relationship to the sampler's initialization parameter set can be queried based on the sampler correspondence relationship. The DPM++ sampler features excellent denoising quality, balanced efficiency, stable denoising effects, and reproducibility for the same initialization noise. Due to its realism, we set its denoising iteration number as timesteps = T1. For an input with an initialization parameter set of input2 = [character, anime, anime model, anime LORA], a Euler sampler with a mapping relationship to the sampler's initialization parameter set can be queried based on the sampler correspondence relationship. The Euler sampler features a very fast iteration speed, achieving a good denoising effect in a relatively short number of iterations. Its denoising iteration number is timesteps = T2. Where T2 is less than T1.

[0129] Through the above method, the image generation model and sampler matching the text-image task can be accurately and conveniently determined through the prompt words.

[0130] As an optional embodiment, the above S103 may include:

[0131] Obtain model keywords corresponding to the image generation model, and add the model keywords to the prompt words to obtain a first intermediate text;

[0132] When a special phrase representing a special object exists in the prompt word, obtaining an object keyword corresponding to the special object, and adding the object keyword to the first intermediate text to obtain a second intermediate text;

[0133] Convert the second intermediate text into the final text.

[0134] In some embodiments, the image generation model may include a base image generation model and a specific object generation model. Each base image generation model is used to implement one or more vertical subject text image tasks, and each specific object generation model is used to implement one or more image style text image tasks.

[0135] For the image generation model, a model keyword mapping table associated with the image generation model can be pre-defined. In the model keyword mapping table, each image generation model has a model keyword with a mapping relationship with it. After determining the image generation model, the model keyword with a mapping relationship for the image generation model can be retrieved and determined from the model keyword mapping table, and these model keywords are added to the prompt words to obtain the first intermediate text.

[0136] For example, since each basic image generation model corresponds to at least one vertical domain and each specific object generation model corresponds to an image style, the model keywords that have a mapping relationship with the basic image generation model are used to describe the vertical domain corresponding to the basic image generation model, and the model keywords that have a mapping relationship with the specific object generation model are used to describe the image style corresponding to the specific object generation model.

[0137] Furthermore, special objects can be descriptive objects that the basic image generation model cannot understand or support, such as a specific person, logo, or address. The specific object generation model can be trained to understand these special objects and generate an object keyword mapping table associated with the special objects. In the object keyword mapping table, each special object has an object keyword with a mapping relationship.

[0138] When it is detected that there is a special phrase representing a special object in the prompt word, the object keyword having a mapping relationship with the special object can be queried from the object keyword mapping table, and the object keyword can be added to the first intermediate text to obtain the second intermediate text, and then the second intermediate text can be further converted into the final text.

[0139] The embodiment of the present application adds keywords corresponding to the model and keywords corresponding to special objects to the prompt words, so that the text content of the prompt words is richer and it is easier for the model to better understand the text content of the input model, thereby improving the accuracy of the generated image.

[0140] As an optional embodiment, converting the second intermediate text into the final text includes:

[0141] When an informal phrase is detected in the second intermediate text, searching for a colloquial phrase corresponding to the informal phrase according to a preset second correspondence relationship, where the second correspondence relationship is a correspondence relationship between the informal phrase and the colloquial phrase;

[0142] The corresponding informal phrases in the second intermediate text are replaced with the popular phrases to obtain the final text.

[0143] In some embodiments, informal phrases include regional slang, colloquialisms, classical Chinese poetry, proverbs, and internet jargon. These informal phrases are typically commonly used linguistic expressions within a region or community, and are often regional and limited in nature. Image generation models typically cannot understand these informal phrases.

[0144] The second correspondence may be a mapping relationship. An informal phrase mapping table may be pre-set, in which each informal phrase has a popular phrase with a mapping relationship, and the popular phrase is an easy-to-understand, popular explanation of the informal phrase.

[0145] Therefore, whenever an informal phrase is detected in the second intermediate text, a colloquial phrase having a mapping relationship with the informal phrase can be searched in the informal phrase mapping table, and then the corresponding informal phrase in the second intermediate text is replaced with the colloquial phrase to obtain the final text.

[0146] For example, when the second intermediate text is "A cute tabby cat with a chubby face and an anime-style appearance is basking in the sun on a lawn full of flowers", where "hutouhunao" is a Chinese slang or colloquialism, that is, an informal phrase, you can query the popular phrase "round face, chubby" that has a mapping relationship with "hutouhunao", so that you can convert the second intermediate text into the final text "A cute tabby cat with a round face and chubby face and anime-style appearance is basking in the sun on a lawn full of flowers".

[0147] The embodiment of the present application converts obscure informal phrases in the text content to be input into the model, so that the model can better understand the text content of the input model, thereby improving the accuracy of the generated image.

[0148] As an optional embodiment, the image generation model includes a basic image generation model and a specific object generation model. After determining the image generation model and sampler corresponding to the prompt word according to a preset first correspondence, the method further includes:

[0149] Identifying object-specific modules in the base image generation model that match the functionality of the object-specific generation model;

[0150] Obtain input data of the specific object module and input the input data into the specific object generation model;

[0151] Obtain module output data obtained by the specific object module in response to the input data, and obtain model output data obtained by the specific object generation model in response to the input data;

[0152] Merge module output data and model output data to obtain fused data;

[0153] The fused data is updated to the model parameters corresponding to the specific object in the basic image generation model, so as to mount the specific object generation model to the basic image generation model.

[0154] In some embodiments, the object-specific generation model can, with minimal power consumption, assist the base image generation model in accurately understanding and generating images of special objects or certain image styles that the base image generation model cannot understand or support. The object-specific generation model cannot operate independently, so it must be attached to the base image generation model to produce the final image generation model.

[0155] Specifically, after selecting the specific object generation model, a specific object module that matches the function of the specific object generation model can be determined in the basic image generation model, and then the specific object generation model can be connected to the basic image generation model according to the connection method of the specific object module.

[0156] Specifically, the input data originally to be input into the specific object module can be input into the specific object module and the specific object generation model at the same time. After that, the specific object module can obtain module output data in response to the input data; the specific object generation model can also respond to the input data to obtain model output data. The module output data and the model output data can be added to obtain fused data, and the fused data can be input as the model parameters corresponding to the specific object into the next level of the specific object module in the basic image generation model, thereby completing the mounting of the specific object generation model.

[0157] This embodiment uses a specific object generation model to adjust the parameters of the basic image generation model, thereby improving the image generation accuracy of the basic image generation model with less power consumption.

[0158] As an optional embodiment, obtaining model output data obtained by the specific object generation model in response to input data includes:

[0159] Get the model weights in the generative model for a specific object;

[0160] Inputting the model weights into the specific object generation model to update the weight parameters of the specific object generation model;

[0161] The specific object generation model after updating the weight parameters responds to the input data to obtain model output data.

[0162] In some embodiments, during the model mounting process, the model parameters in the specific object generation model can be loaded by reading the weight file of the specific object generation model or directly using the model loading function provided in the programming language. The model parameters include model weights, which are learnable parameters in the model and are used to adjust the influence of input features.

[0163] Then, when there are multiple different specific object generation models, the multiple specific object generation models can be mounted on the basic image generation model respectively to obtain multiple combined models, and then the multiple combined models are exported to ONNX (Open Neural Network Exchange) format. ONNX (Open Neural Network Exchange) is an open deep learning model representation standard that aims to improve cross-platform and cross-framework model interoperability.

[0164] By traversing the weights of each node in each combined model and comparing them, the ONNX nodes with different weights in each combined model can be marked; since the weights of each node in the basic image generation model are consistent, these nodes with different weights are the model weights of the specific object generation model in the combined model. In this way, the model weights corresponding to each specific object generation model can be obtained.

[0165] Subsequently, the model weight can be input into the specific object generation model corresponding to the mounted model weight to update the weight parameters of the specific object generation model, thereby adjusting the response of the specific object generation model to the input data, and obtaining model output data in response to the input data.

[0166] Through the above method, the model weights loaded in real time can be converted into the input of the specific object generation model after mounting, thereby optimizing the performance of the mounted model.

[0167] As an optional embodiment, the specific object generation model includes at least one specific object generation node, the model weight includes at least one node weight, each specific object generation node corresponds to a node weight, and inputting the model weight into the specific object generation model includes:

[0168] Add at least one identity node to the specific object generation model, where the output of the identity node is equal to the input of the identity node, and each identity node corresponds to a specific object generation node;

[0169] Each node weight in the at least one node weight is input into a corresponding specific object generation node through a corresponding identity node.

[0170] In some embodiments, in the process of inputting model weights into a specific object generation model, the input of model weights can be completed by adding identity nodes in the specific object generation model and inputting the node weights of each specific object generation node into the specific object generation node through the identity nodes.

[0171] Among them, the identity node can be an Identity node. The Identity node is a special glue node. The input of the node is equal to the output. A bypass path can be introduced through the Identity node to allow information to pass directly without additional transformation.

[0172] Through the above method, the model weights can be accurately and quickly introduced into the specific object generation model after mounting.

[0173] As an optional embodiment, the sampler includes a single sampling algorithm and a number of iterations, and uses the image generation model updated with the denoising parameters to convert the final text into a target image, including:

[0174] Get a randomly generated noise image;

[0175] The final text is encoded using a text vectorization encoding method to obtain a basic text embedding vector;

[0176] The basic text embedding vector is embedded in the image generation model, and the image generation model is used to perform N denoising processes on the noisy image according to the single sampling algorithm to obtain the basic latent feature image, where N is the number of iterations;

[0177] The basic latent feature image is decoded by image decoding to obtain the target image.

[0178] In some embodiments, a randomly generated noise image can be obtained, and the text content of the final text can be mapped to a high-dimensional vector space to obtain a basic text embedding vector. After obtaining the basic text embedding vector, the basic text embedding vector and the noise image can be input into the image generation model together, and the basic text embedding vector can be used to guide the model to perform N-iteration denoising processing on the noise image according to a single sampling algorithm to obtain a basic latent feature image. The basic latent feature image can then be decoded by using an autovariation encoder image generation model to obtain a target image that can be recognized by the human eye.

[0179] In this embodiment, the conversion from the final text to the target image can be completed accurately.

[0180] As an optional embodiment, decoding the basic latent feature image in an image decoding manner to obtain a target image includes:

[0181] When the target image style is realistic, the final text is encoded to obtain a refined text embedding vector;

[0182] Embedding the refined text embedding vector into the image refinement model, and using the image refinement model embedded with the refined text embedding vector to perform refinement and denoising on the basic latent feature image to obtain a refined latent feature image;

[0183] The refined latent feature image is decoded to obtain the target image.

[0184] In some embodiments, when the style of the target image is realistic, the target image has high requirements on image quality and image details. Therefore, it is necessary not only to encode the basic semantics of the prompt word into a basic text embedding vector, but also to encode the detailed semantics of the prompt word into a refined text embedding vector.

[0185] Since the basic latent feature image is only an image in the feature space that meets the basic features of the prompt word, and the realistic style target image has high requirements for details, the basic latent feature image and the refined text embedding vector can be further input into the refinement model, and the refined text embedding vector can be used to guide the refinement model to perform further multiple iterative refinement and denoising processing on the basic latent feature image to obtain a refined latent feature image. The refined latent feature image can then be decoded using the autovariation encoder model to obtain a target image that can be recognized by the human eye.

[0186] The refinement model may be a refinement model in a stable diffusion model, which is used to refine an image with insufficient details and enrich the details of the image. The refined latent feature image is an image in the feature space that meets the detail features of the prompt word.

[0187] In this embodiment, for realistic style images with high detail requirements, the basic image generation model and the refined model can be cascaded twice, the basic image generation task can be completed by the image generation model, and the image obtained by the basic model can be refined by the refined model to ensure the image generation quality.

[0188] As an optional embodiment, decoding the basic latent feature image in an image decoding manner to obtain a target image includes:

[0189] When the image style of the target image is an artistic style, the basic latent feature image is decoded to obtain the target image.

[0190] In some embodiments, when the style of the target image is artistic, the target image is more inclined to the display of the painting style and does not require high details of the image. Therefore, the basic latent feature image can be directly decoded to obtain the target image.

[0191] The base model may be a base model in a stable diffusion model, which is used to convert text content into an image with insufficient details.

[0192] In this embodiment, for artistic style images that do not require high details, the basic image generation task can be completed only by using the basic model, thereby effectively reducing the power consumption of algorithm reasoning while meeting the requirements.

[0193] As an optional embodiment, obtaining a randomly generated noise image includes:

[0194] Get the preset random number seed;

[0195] Convert the random number seed into a random sequence with a normal distribution to obtain Gaussian noise;

[0196] Encode the Gaussian noise to obtain a noisy image.

[0197] In some embodiments, a pre-set random number seed can be obtained to deterministically generate a Gaussian distributed random sequence to form Gaussian noise. After obtaining the Gaussian noise, the Gaussian noise can be mapped to a latent space by an encoder to obtain a potential representation of the Gaussian noise, i.e., a noise image.

[0198] In this embodiment, a given random number seed can be used to ensure that the noise image generated each time the code is run is reproducible.

[0199] As an optional embodiment, the image generation model is used to perform N denoising processes on the noisy image according to a single sampling algorithm, including:

[0200] For the i-th noise reduction process among the N noise reduction processes, obtain an intermediate noise image and obtain the noise standard deviation of the intermediate noise image, where the intermediate noise image is a latent space image obtained by the noise image after the (i-1)-th noise reduction process, and i is any positive integer less than or equal to N;

[0201] Update the parameters of the single sampling algorithm using the noise standard deviation;

[0202] The image generation model embedded with the basic text embedding vector is used to perform single denoising on the intermediate noisy image according to the single sampling algorithm after parameter update until N denoising processes are completed.

[0203] In some embodiments, before applying a sampler, it is necessary to first complete the sampler construction. During the sampler construction process, it is necessary to first initialize the sampler and then determine the sampler's single sampling algorithm and number of iterations. The single sampling algorithm determines the method and intensity of noise removal during each noise reduction process.

[0204] Specifically, the target image's noise standard deviation serves as a model parameter during each iteration of the single-sampling algorithm. The model parameters are learned during the algorithm's execution and updated during each iteration. Therefore, it's necessary to obtain the target image's noise standard deviation before each denoising process. This noise standard deviation is then used to update the single-sampling algorithm, and the updated single-sampling algorithm is then used to perform a single denoising process on the input target image.

[0205] Specifically, for the i-th noise reduction process in the N-th noise reduction process, the noise reduction formula is as follows: θ (x;σ)=c skip (σ)x+c out (σ)F θ (c in (σ)x;c noise (σ)

[0206] Where x represents the latent space image obtained by (i-1) times of denoising of the noise image, that is, the target image. If i is 1, then x is the noise image, D θ is the latent space image after the i-th denoising process. σ is used to represent the standard deviation of the noise of the latent space image obtained after the (i-1)th denoising process. θ represents the denoising neural network algorithm inference. Cskip(σ), Cout(σ), Cin(σ), and Cnoise(σ) represent the four denoising process terms in the denoising formula. Specifically, Cskip(σ) represents the skip connection scaling process, Cout(σ) represents the output scaling process, Cin(σ) represents the input scaling process, and Cnoise(σ) represents the noise adjustment process.

[0207] The four noise reduction process items mentioned above are different in different samplers. For example, in the DPM sampler:

[0208] Skip scaling c skip (σ)1

[0209] Output scaling c out (σ)-σ

[0210] Noise cond.c noise (σ)(M-1)σ -1 (σ)

[0211] In DDIM Sampler:

[0212] Skip scaling c skip (σ)1

[0213] Output scaling c out (σ)-σ

[0214] Noise cond.c noise (σ)M-1-arg min j |u j -σ|

[0215] Among them, M is the initial value of random noise, u j is the jth noise adjustment parameter.

[0216] In this way, once the sampler type is determined, the sampler can be constructed according to the denoising process and number of iterations specified by that type of sampler. After the sampler is constructed, for each denoising iteration, the latent space image obtained from the previous denoising process and the noise standard deviation of the latent space image are input into the sampler to complete the next round of denoising.

[0217] As an optional embodiment, obtaining the noise standard deviation of the intermediate noise image includes:

[0218] Get the sampler's hyperparameters, as well as the sampler's minimum and maximum noise standard deviations;

[0219] When i is less than N, the noise standard deviation is obtained by inputting the hyperparameters, the minimum noise standard deviation, the maximum noise standard deviation, and i into the standard deviation calculation formula;

[0220] When i is equal to N, the noise standard deviation is determined to be 0.

[0221] In some embodiments, when the sampler is constructed and used to implement noise reduction, it is necessary to obtain the noise standard deviation of the target image once during each noise reduction process and input the noise standard deviation into the sampler together with the target image.

[0222] Specifically, during the i-th denoising process, the sampler's hyperparameters, as well as the pre-specified minimum and maximum noise standard deviations, can be first obtained. The hyperparameters are used to control the step size of the variance transformation.

[0223] When i is less than N, the noise standard deviation in the i-th iteration can be calculated according to the following standard deviation calculation formula:

[0224] Among them, i represents the number of iterations, N represents the total number of iterations, σ min and σ max They represent the minimum and maximum noise standard deviation respectively, and ρ is a hyperparameter.

[0225] For example, the hyperparameters can be 7, σ min and σ max 0.02 and 100 respectively.

[0226] When i equals N, σ N is 0.

[0227] In this way, the noise standard deviation of the sampler can be adjusted according to a predetermined step size to avoid the noise standard deviation from changing too small or too large.

[0228] As an optional embodiment, after the above S105, the following steps may also be included:

[0229] Perform super-resolution processing on the target image to obtain the target image.

[0230] In some embodiments, since the operation of the image generation model requires a large amount of computing power, the resolution of the target image output by the image generation model is usually low. The target image can be super-resolved using a trained super-resolution model to improve the overall resolution of the target image and obtain a clearer target image.

[0231] For example, the super-resolution model can be an Enhanced Super-Resolution Generative Adversarial Network (ESRGAN), which performs 2X super-resolution on the target image. Specifically, ESRGAN uses deep learning techniques and a generative adversarial network to improve the spatial resolution of the target image. 2X super-resolution processes the target image to generate an image with twice the resolution in both the horizontal and vertical directions, thereby improving the visual quality and detail of the image.

[0232] As an optional embodiment, the super-resolution model can be a residual in residual dense block (RRDB) module without batch normalization (BN). For example, the RRDB module may include three interconnected dense connection blocks (Dense Block), each of which has five convolutional layers. These convolutional layers may have different filters and feature map depths for learning representations of different levels of the image. In this way, the super-resolution model has better generalization.

[0233] As an optional embodiment, the image generation method may further include:

[0234] When the attribute of the image style is the first type of style attribute, determining the target fineness of the target image as a low fineness;

[0235] When the attribute of the image style is the second type of style attribute, the target fineness of the target image is determined to be a high fineness.

[0236] In some embodiments, the target fineness refers to the user's expectation of the fineness of the target image.

[0237] As an optional embodiment, a text vectorization encoding method is used to encode the final text to obtain a basic text embedding vector, including:

[0238] When the target fineness of the target image is low, the final text is encoded to obtain a basic text embedding vector; the basic text embedding vector is used to represent the basic semantics of the prompt word.

[0239] In some embodiments, since the requirements for image details are not high when the target level of fineness is low, only the basic semantic information in the final text can be captured, and the basic semantic information can be encoded into a basic text embedding vector. The basic text embedding vector is used to guide the final text to perform reasoning and obtain the target image.

[0240] As an optional embodiment, encoding the final text to obtain a refined text embedding vector includes:

[0241] When the target fineness of the target image is high, the final text is encoded to obtain a basic text embedding vector and a refined text embedding vector; the refined text embedding vector is used to represent the detailed semantics of the prompt word.

[0242] In some embodiments, when the target level of detail is high, users expect the target image to contain more details and richer content. Therefore, not only can the basic semantic information in the final text be captured and encoded into a base text embedding vector, but also the detailed semantic information in the final text can be captured and encoded into a refined text embedding vector. The final text is then inferred using the base text embedding vector and the refined text embedding vector as guidance, resulting in the final target image.

[0243] As an optional embodiment, the image generation model includes a basic image generation model and a specific object generation model, and determining the image generation model and sampler corresponding to the prompt word includes:

[0244] Analyzing the prompt words to determine the content category of the target image corresponding to the prompt words, wherein the content category includes subject category and style category;

[0245] Determine the basic image generation model corresponding to the subject category; the basic image generation model corresponds to the basic image generation model sequence number;

[0246] Determine the specific object generation model corresponding to the style category; the specific object generation model corresponds to the specific object generation model sequence number;

[0247] Query the first mapping relationship table in the sampler database to determine the samplers corresponding to the subject category, style category, basic image generation model number and specific object generation model number, wherein the first mapping relationship table includes the first mapping relationship between the subject category, style category, basic image generation model number and specific object generation model number and the sampler.

[0248] In some embodiments, the subject category refers to the type of subject in the target image, while the style category generally refers to the visual appearance and stylistic features of the image. Each image generation model is used to implement image generation tasks in one or more vertical domains. Therefore, a first correspondence between content categories and image generation models, as well as a first mapping between content categories, image generation models, and samplers can be pre-set in the sampler database based on the vertical domain corresponding to each image generation model.

[0249] As an optional embodiment, the specific object generation model is mounted into the basic image generation model, including:

[0250] The specific object generation model is mounted on the basic image generation model to obtain an application combination model, and the model weights in the basic image generation model are adjusted to null values;

[0251] Get the basic weight parameters of the basic image generation model;

[0252] Convert the basic weight parameters of the basic image generation model into a basic weight file corresponding to the basic image generation model that conforms to the input format;

[0253] The base weight file is input into the base image generation model in the application combination model, and the null value is updated as the base weight parameter of the base weight file.

[0254] In some embodiments, a correspondence between each basic weight file and the basic image generation model can be set in advance. When a basic image generation model is detected to be mounted on the basic image generation model, the basic weight file corresponding to the basic image generation model can be obtained based on the correspondence and input into the mounted basic image generation model. The basic weight file includes the basic weight parameters obtained in advance.

[0255] Specifically, the base weight file is used as an input to the base image generation model in the application combination model, adjusting the model weights of the base image generation model in the application combination model from null values ​​to the base weight parameters in the base weight file. This allows the weight parameters of the mounted model in the combination model to be dynamically switched during the application of the base image generation model.

[0256] Based on the image generation method provided in the above embodiment, the present application also provides a specific implementation of an image generation device. Please refer to the following embodiment.

[0257] First, referring to FIG. 2 , an image generation device 200 provided in some embodiments of the present application includes the following modules:

[0258] The first acquisition module 201 is used to acquire a prompt word input by a user, where the prompt word is used to instruct the generation of a target image;

[0259] A first determination module 202 is configured to determine an image generation model and a sampler corresponding to a prompt word based on a predetermined first correspondence relationship, wherein the first correspondence relationship is a correspondence relationship between the prompt word, the image generation model, and the sampler, and the sampler is configured to control a denoising method of the image generation model;

[0260] An expansion module 203 is used to expand the prompt word to obtain a final text;

[0261] An updating module 204 is configured to load the sampler into the image generation model to update the noise reduction parameters of the image generation model;

[0262] The first conversion module 205 is configured to convert the final text into a target image using the image generation model after the noise reduction parameters are updated.

[0263] As an implementation of the present application, the image generation model includes a basic image generation model and a specific object generation model, the first correspondence includes a model correspondence and a sampler correspondence, and the first determination module 202 may further include:

[0264] The first recognition unit is used to perform semantic recognition on the prompt word and determine the image generation intention corresponding to the prompt word;

[0265] The second recognition unit is used to recognize specific generation information contained in the image generation intention corresponding to the prompt word; the specific generation information includes a specific subject and / or a specific style;

[0266] A first determining unit is configured to determine a basic image generation model and a specific object generation model corresponding to the prompt word based on a model correspondence relationship, wherein the model correspondence relationship includes a correspondence relationship between specific generation information and the specific object generation model, and a correspondence relationship between the image generation intent and the basic image generation model;

[0267] The second determination unit is used to determine the samplers corresponding to the prompt word, the basic image generation model and the specific object generation model according to the sampler correspondence relationship. The sampler correspondence relationship includes the correspondence between the image generation intention, the specific generation information, the basic image generation model, the specific object generation model and the sampler.

[0268] As an implementation of the present application, the second determining unit may further include:

[0269] A first acquisition subunit is configured to acquire identification information of an image generation intention, specific generation information, a basic image generation model, and a specific object generation model;

[0270] A first determining subunit, configured to determine the identification information as an initialization parameter of the sampler;

[0271] The first matching subunit is configured to match a sampler corresponding to the initialization parameter from a sampler correspondence relationship, where the sampler correspondence relationship includes a correspondence relationship between the initialization parameter and the sampler.

[0272] As an implementation of the present application, the expansion module 203 may further include:

[0273] A first adding unit is configured to obtain a model keyword corresponding to the image generation model, and add the model keyword to the prompt word to obtain a first intermediate text;

[0274] a second adding unit configured to obtain an object keyword corresponding to the special object when a special phrase representing a special object exists in the prompt word, and to add the object keyword to the first intermediate text to obtain a second intermediate text;

[0275] The first conversion unit is configured to convert the second intermediate text into a final text.

[0276] As an implementation of the present application, the first conversion unit may further include:

[0277] a first query subunit configured to, when detecting the presence of an informal phrase in the second intermediate text, query for a colloquial phrase corresponding to the informal phrase based on a preset second correspondence relationship, wherein the second correspondence relationship is a correspondence relationship between the informal phrase and the colloquial phrase;

[0278] The replacement subunit is used to replace the corresponding informal phrases in the second intermediate text with the popular phrases to obtain the final text.

[0279] As an implementation of the present application, the image generating device 200 may further include:

[0280] A second determining module is used to determine a specific object module in the basic image generation model that matches the function of the specific object generation model;

[0281] A second acquisition module is used to acquire input data of the specific object module and input the input data into the specific object generation model;

[0282] a third acquisition module, configured to acquire module output data obtained by the specific object module in response to input data, and to acquire model output data obtained by the specific object generation model in response to input data;

[0283] A merging module is used to merge module output data and model output data to obtain fused data;

[0284] The first mounting module is used to update the fusion data into model parameters corresponding to the specific object in the basic image generation model, so as to mount the specific object generation model into the basic image generation model.

[0285] As an implementation of the present application, the third acquisition module may further include:

[0286] A first acquisition unit is used to acquire a model weight in a specific object generation model;

[0287] An input unit, configured to input the model weights into the specific object generation model to update the weight parameters of the specific object generation model;

[0288] The response unit is used to generate a model using the specific object after updating the weight parameters to respond to the input data and obtain model output data.

[0289] As an implementation of the present application, the input unit may further include:

[0290] An adding subunit is used to add at least one identity node in the specific object generation model, the output of the identity node is equal to the input of the identity node, and each identity node corresponds to a specific object generation node;

[0291] The input subunit is used to input each node weight in at least one node weight into the corresponding specific object generation node through the corresponding identity node.

[0292] As an implementation of the present application, the first conversion module 205 may further include:

[0293] A second acquisition unit is used to acquire a randomly generated noise image;

[0294] A first encoding unit is used to encode the final text using a text vectorization encoding method to obtain a basic text embedding vector;

[0295] A first denoising unit is configured to embed the basic text embedding vector into the image generation model, and perform denoising on the noisy image N times according to a single sampling algorithm using the image generation model to obtain a basic latent feature image, where N is the number of iterations;

[0296] The first decoding unit is used to decode the basic latent feature image by adopting an image decoding method to obtain a target image.

[0297] As an implementation of the present application, the first decoding unit may further include:

[0298] The first encoding subunit is configured to encode the final text to obtain a refined text embedding vector when the image style of the target image is a realistic style;

[0299] A refinement subunit is used to embed the refined text embedding vector into the image refinement model, and use the image refinement model embedded with the refined text embedding vector to perform refinement and noise reduction on the basic latent feature image to obtain a refined latent feature image;

[0300] The first decoding subunit is used to decode the refined latent feature image to obtain a target image.

[0301] As an implementation of the present application, the first noise reduction unit may also be used to:

[0302] For the i-th noise reduction process among the N noise reduction processes, obtain an intermediate noise image and obtain the noise standard deviation of the intermediate noise image, where the intermediate noise image is a latent space image obtained by the noise image after the (i-1)-th noise reduction process, and i is any positive integer less than or equal to N;

[0303] Use the noise standard deviation to update the parameters of the single sampling algorithm;

[0304] The image generation model embedded with the basic text embedding vector is used to perform single denoising on the intermediate noisy image according to the single sampling algorithm after parameter update until N denoising processes are completed.

[0305] As an implementation of the present application, the first noise reduction unit may also be used to:

[0306] Get the sampler's hyperparameters, as well as the sampler's minimum and maximum noise standard deviations;

[0307] When i is less than N, the hyperparameters, the minimum noise standard deviation, the maximum noise standard deviation, and i are input into the standard deviation calculation formula to obtain the noise standard deviation;

[0308] When i is equal to N, the noise standard deviation is determined to be 0.

[0309] The image generation device provided in the embodiment of the present application can implement each step in the above-mentioned method embodiment. To avoid repetition, they will not be described here.

[0310] FIG3 shows a flow chart of another image generation method provided in some embodiments of the present application. The method includes the following steps:

[0311] S301: Obtain a prompt word input by a user, where the prompt word is used to instruct generation of a target image.

[0312] In some embodiments, the image generation task can convert a prompt word input by the user into a target image, where the prompt word is used to describe the content of the target image that the user wants to generate. For example, the prompt word can be "generate a puppy" or "generate a tree."

[0313] Furthermore, the input prompt words can include positive prompt words and negative prompt words. Positive prompt words indicate desired content in the target image, while negative prompt words indicate undesirable content in the target image. For example, a positive prompt word could be "A cartoon-style Chinese girl running on the beach, with flying seagulls and a gorgeous rainbow behind her, the overall picture is poetic and picturesque," while a negative prompt word could be "pornographic, nude, ugly, deformed."

[0314] S302 , performing semantic recognition on the prompt word and determining the image style of the target image according to the result of the semantic recognition.

[0315] In some embodiments, in image generation tasks, image style generally refers to the visual appearance and style characteristics of an image. Common image styles can include two categories: realistic style and artistic style.

[0316] Realistic styles can include realistic figures, landscapes, animals, and architecture. They emphasize authentic color expression and strive to restore the true colors and lighting effects of figures. Therefore, they have high requirements for image quality and are very demanding on image detail. Artistic styles can include comics, oil paintings, ink paintings, and line drawings. These styles focus more on the presentation of artistic style and require less detail than realistic styles.

[0317] The semantics of the prompt words can be recognized, and the user intention can be determined based on the result of the semantic recognition. Then, the image style indication information is obtained from the user intention, and the image style of the target image can be determined based on the indication information.

[0318] S303: Encode the prompt word according to the encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style, wherein different image styles correspond to different encoding methods.

[0319] In some embodiments, a text embedding vector is a representation of the text content of a prompt word mapped into a high-dimensional vector space. This representation enables the image generation model to better understand the meaning of the text embedding vector. Specifically, the prompt word can be encoded into at least one text embedding vector by capturing the semantic and grammatical information in the prompt word.

[0320] Since different image styles have different requirements for target images, the semantic and grammatical information of the prompt word can be selectively encoded into at least one text embedding vector based on the analyzed image style. Target images of different image styles then correspond to different numbers or types of text embedding vectors.

[0321] As an optional embodiment, encoding the prompt word according to the encoding method corresponding to the image style to obtain the text embedding vector corresponding to the image style includes:

[0322] Determine the target refinement of the target image based on the attributes of the image style;

[0323] Query the encoding method corresponding to the target precision level according to the pre-set encoding correspondence relationship;

[0324] Encode the prompt words according to the encoding method corresponding to the target level of refinement to obtain the text embedding vector corresponding to the image style.

[0325] As an optional embodiment, determining the target fineness of the target image according to the attributes of the image style includes:

[0326] When the attribute of the image style is the first type of style attribute, determining the target fineness of the target image as a low fineness;

[0327] When the attribute of the image style is the second type of style attribute, the target fineness of the target image is determined to be a high fineness.

[0328] In some embodiments, the target fineness refers to the user's expectation of the fineness of the target image.

[0329] For example, when the style of the target image is artistic, the target image is more inclined to the display of the painting style and does not require high details of the image. Therefore, the target fineness of the target image corresponding to the artistic style can be determined as low fineness, which means that some abstract or simplified reasoning encoding methods can be used to encode the prompt words without paying attention to too many details.

[0330] When the style of the target image is realistic, the target image has very high requirements on image quality and image details. Therefore, not only the basic semantics of the prompt word need to be reflected in the target image, but also the detailed semantics of the prompt word need to be reflected in the target image. Therefore, a more complex and precise encoding method can be used to encode the prompt word.

[0331] As an optional embodiment, the prompt word is encoded according to the encoding method corresponding to the target fineness to obtain a text embedding vector corresponding to the image style, including:

[0332] When the target fineness of the target image is high, the prompt word is encoded to obtain a basic text embedding vector and a refined text embedding vector. The basic text embedding vector is used to represent the basic semantics of the prompt word, and the refined text embedding vector is used to represent the detailed semantics of the prompt word.

[0333] When the target fineness of the target image is low, the prompt word is encoded to obtain a basic text embedding vector.

[0334] In some embodiments, when the target level of detail is high, users expect the target image to contain greater detail and richer content. Therefore, not only can the basic semantic information of the prompt word be captured and encoded into a base text embedding vector, but also the detailed semantic information of the prompt word can be captured and encoded into a refined text embedding vector. The prompt word is then inferred using the base text embedding vector and the refined text embedding vector in sequence to obtain the final target image.

[0335] Since the requirements for image details are not high when the target level of detail is low, we can only capture the basic semantic information in the prompt word and encode the basic semantic information into a basic text embedding vector. The basic text embedding vector is used to guide the prompt word inference to obtain the target image.

[0336] S304: Convert the text embedding vector according to the inference process corresponding to the text embedding vector to obtain a target image.

[0337] In some embodiments, the image generation task can be completed by a single image generation model or by a cascade of multiple models. Different image style prompts generate different text embedding vectors, which are then embedded in different image generation models to complete the conversion from prompts to target images, thereby achieving different inference processes for converting prompts to target images.

[0338] As an optional embodiment, converting the text embedding vector according to the inference process corresponding to the text embedding vector to obtain the target image includes:

[0339] According to the model correspondence, determine the image generation model corresponding to the text embedding vector;

[0340] When there is only one image generation model, the text embedding vector is converted using the image generation model to obtain the target image;

[0341] In the case where the image generation model includes at least two, the at least two image generation models are cascaded, and the text embedding vector is converted using the at least two cascaded image generation models to obtain a target image.

[0342] In some embodiments, the model correspondence is a one-to-one correspondence between a pre-defined text embedding vector and an image generation model. The text embedding vector is used to embed each level in the corresponding image generation model, and at each level, the image generation model is guided to denoise the input image to obtain the final target image.

[0343] As an optional embodiment, determining an image generation model corresponding to a text embedding vector includes:

[0344] When the target refinement level of the image style is high refinement level, determining a base model corresponding to the base text embedding vector and a refined model corresponding to the refined text embedding vector;

[0345] Cascading at least two image generation models, and using the at least two cascaded image generation models to transform the text embedding vector to obtain a target image, including:

[0346] When the target refinement level of the image style is high refinement level, the base text embedding vector is embedded into the base model, and the refined text embedding vector is embedded into the refined model;

[0347] Get a randomly generated noise image;

[0348] Denoising the noisy image using a basic model embedded with a basic text embedding vector to obtain a first latent feature image;

[0349] Performing denoising on the first latent feature image using a refined model embedded with the refined text embedding vector to obtain a second latent feature image;

[0350] The second latent feature image is decoded to obtain a target image.

[0351] In some embodiments, when the style of the target image is realistic, the target image has high requirements on image quality and image details. Therefore, it is necessary not only to encode the basic semantics of the prompt word into a basic text embedding vector, but also to encode the detailed semantics of the prompt word into a refined text embedding vector.

[0352] After encoding to obtain the basic text embedding vector and the refined text embedding vector, similarly, the basic model corresponding to the basic text embedding vector and the refined model corresponding to the refined text embedding vector can be further determined. Then, a randomly generated noise image is obtained, and the noise image and the basic text embedding vector are input into the basic model. The basic text embedding vector is used to guide the basic model to perform multiple iterative denoising processes on the noise image to obtain the first potential feature image.

[0353] Since the first latent feature image is only an image that meets the basic features of the prompt word in the feature space, and the realistic style target image has high requirements for details, the first latent feature image and the refined text embedding vector can be further input into the refinement model, and the refined text embedding vector can be used to guide the refinement model to perform further multiple iterative denoising processes on the first latent feature image to obtain the second latent feature image. The second latent feature image can be decoded by using the autovariation encoder model to obtain a target image that can be recognized by the human eye.

[0354] As an optional embodiment, as shown in Figure 4, when the style of the target image is realistic, the positive prompt word Prompt and the negative prompt word Negative Prompt can be used as inputs of the multimodal large model CLIP. The CLIP model can use a text encoder to parse the prompt word and convert it into basic text embedding vectors (Text Embeddings) and refined text embedding vectors (Refiner Text Embeddings). In addition, a Gaussian noise Gaussian Noise can be initialized according to the random number seed Seed and converted into a noise image Latents, and the basic text embedding vector and the noise image are input into the generative large model UNet, and the base model (base model) in the UNet model is passed through multiple iterations of the first sampler Sampler1 to output the first latent feature image Conditioned Latents1, and then the basic text embedding vector and the noise image are input into the generative large model UNet, and the refinement model (Refiner model) in the UNet model is passed through multiple iterations of the second sampler Sampler2 to output the second latent feature image Conditioned Latents2, and the decoder in the autovariation encoder VAE is used to decode the second latent feature image to obtain a target image (output image) that can be recognized by the human eye.

[0355] In this embodiment, for realistic style images with high detail requirements, the basic model and the refined model can be cascaded twice, the basic image generation task can be completed by the basic model, and the image obtained by the basic model can be refined by the refined model to ensure the image generation quality.

[0356] As an optional embodiment, determining an image generation model corresponding to a text embedding vector includes:

[0357] When the attribute of the image style is the first type of style attribute, determining a basic model corresponding to the basic text embedding vector;

[0358] When there is only one image generation model, the text embedding vector is converted using the image generation model to obtain the target image, including:

[0359] When the attribute of the image style is the first type of style attribute, the basic text embedding vector is embedded into the basic model;

[0360] Get a randomly generated noise image;

[0361] Denoising the noisy image using a basic model embedded with a basic text embedding vector to obtain a first latent feature image;

[0362] The first latent feature image is decoded to obtain a target image.

[0363] In some embodiments, when the style of the target image is artistic, the target image focuses more on the presentation of the painting style and does not require high details of the image. Therefore, only the basic semantics of the prompt word can be encoded into the basic text embedding vector.

[0364] After encoding to obtain the basic text embedding vector, the basic model corresponding to the basic text embedding vector can be further determined, and then a randomly generated noise image can be obtained. The noise image and the basic text embedding vector are input into the basic model, and the basic text embedding vector is used to guide the basic model to perform multiple iterative denoising processes on the noise image to obtain a first latent feature image, wherein the first latent feature image is an image in the feature space that meets the basic features of the prompt word. The first latent feature image can be decoded by using the autovariation encoder model to obtain a target image that can be recognized by the naked eye.

[0365] As an optional embodiment, as shown in Figure 5, when the style of the target image is artistic, the positive prompt word Prompt and the reverse prompt word Negative Prompt can be used as inputs of the multimodal large model CLIP. The CLIP model can use the text encoder (Text Encoder) to parse the prompt word and convert it into a basic text embedding vector (Text Embeddings). In addition, a Gaussian Noise can be initialized according to the random number seed Seed and converted into a noise image Latents, and the basic text embedding vector and the noise image can be input into the generated large model UNet. The base model in the UNet model is passed through multiple iterations of the sampler Sampler to output the first latent feature image Conditioned Latents, and then the decoder in the autovariation encoder VAE is used to decode the first latent feature image to obtain a target image (output image) that can be recognized by the human eye.

[0366] In this embodiment, for artistic style images that do not require high details, the basic image generation task can be completed only by using the basic model, thereby effectively reducing the power consumption of algorithm reasoning while meeting the requirements.

[0367] In the embodiment of the present application, semantic recognition is performed on the prompt word input by the user. The image style of the target image corresponding to the prompt word is determined based on the semantic recognition results. The prompt word is then encoded according to the image style to obtain a corresponding text embedding vector. The text embedding vector is then converted into the target image using the inference process corresponding to the text embedding vector. In this way, the text embedding vector can be automatically selected based on the different image styles, and the algorithm process adapted to the image style can be automatically selected for inference, achieving reasonable resource allocation. Compared with the existing technology, this avoids the application of a single set of power-intensive standardized processes for all text-to-image tasks, reducing the cost of image generation.

[0368] As an optional embodiment, semantic recognition is performed on the prompt word, and the image style of the target image is determined according to the result of the semantic recognition, including:

[0369] Perform semantic recognition on the prompt word. If a specific phrase exists in the prompt word, determine the image style of the target image as the image style corresponding to the specific phrase. The specific phrase is used to indicate the image style of the prompt word.

[0370] When there is no specific phrase in the prompt words, the image style of the target image is determined to be a default style.

[0371] In some embodiments, semantic understanding can be performed on the acquired prompt words to obtain semantic information that represents the user's intention or needs. Based on the semantic information, it is then determined whether a specific phrase indicating the image style exists in the prompt words. If the specific phrase exists, and the user's need is to determine the style of the target image to be the image style indicated by the specific phrase, then the image style of the target image can be determined to be the image style indicated by the specific phrase. If the specific phrase does not exist, then the image style of the target image can be determined to be a default style by default, which can be a realistic style. The specific phrases can include realism, animation, oil painting, ink painting, sketching, watercolor, and cartoon, etc.

[0372] For example, if the prompt word is "generate a picture of a puppy", and the prompt word does not include any specific phrases, then the image style of the target image corresponding to the prompt word can be determined as a realistic style; if the prompt word is "generate a picture of a puppy in anime style", and the prompt word includes the specific phrase "anime", and based on the results of semantic understanding, the user's demand is to determine the target image as anime style, then the image style of the target image corresponding to the prompt word can be determined as anime style, and anime style belongs to the artistic style.

[0373] Through this embodiment, the image style of each target image to be generated can be accurately and quickly determined based on the prompt words.

[0374] As an optional embodiment, after determining the image generation model corresponding to the text embedding vector according to the model correspondence relationship, the method further includes:

[0375] By performing semantic recognition on the prompt word, the subject type of the prompt word is determined according to the result of semantic recognition;

[0376] Determine the sampler corresponding to the image style and subject type;

[0377] Load the sampler into the image generation model.

[0378] In some embodiments, the sampler is used to control the intensity, parameters, and number of iterations of the noise reduction during the iterative process of the model. Before applying the model to reduce the noise of an image, the selected sampler needs to be pre-loaded into the model.

[0379] Specifically, there are various samplers, including DDIM, Euler, and DPM++2M, each with different characteristics. During sampler selection, semantic recognition of the prompt word is performed to determine the subject type of the prompt word and the style of the target image. A pre-set mapping table is then used to search for a sampler that maps the image style and subject type, and the sampler is then loaded into all of the aforementioned models.

[0380] In this embodiment, since different subject types and image styles have different requirements for target images, a suitable sampler can be selected based on the subject type and image style to control the target image generation process so that different target images are adapted to different samplers.

[0381] As an optional embodiment, obtaining a noise image includes:

[0382] Get the preset random number seed;

[0383] Convert the random number seed to Gaussian noise;

[0384] Encode the Gaussian noise to obtain a noisy image.

[0385] In some embodiments, a pre-set random number seed can be obtained to deterministically generate Gaussian-distributed random numbers, which constitute Gaussian noise. After obtaining the Gaussian noise, the Gaussian noise can be mapped to a latent space by an encoder to obtain a potential representation of the Gaussian noise, i.e., a noise image.

[0386] In this embodiment, a given random number seed can be used to ensure that the noise image generated each time the code is run is reproducible.

[0387] Based on the image generation method provided in the above embodiment, the present application also provides a specific implementation of an image generation device. Please refer to the following embodiment.

[0388] First, referring to FIG6 , another image generation device 400 provided in some embodiments of the present application includes the following modules:

[0389] The fourth acquisition module 401 is used to acquire a prompt word input by the user, where the prompt word is used to instruct the generation of a target image;

[0390] Recognition module 402, configured to perform semantic recognition on the prompt word and determine the image style of the target image according to the result of the semantic recognition;

[0391] An encoding module 403 is configured to encode the prompt word according to an encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style, wherein different image styles correspond to different encoding methods;

[0392] The second conversion module 404 is used to convert the text embedding vector according to the inference process corresponding to the text embedding vector to obtain a target image.

[0393] As an implementation of the present application, the encoding module 403 may further include:

[0394] a third determining unit, configured to determine a target fineness of a target image according to attributes of the image style;

[0395] A first query unit is used to query the encoding method corresponding to the target fineness according to a pre-set encoding correspondence relationship;

[0396] The second encoding unit is used to encode the prompt word according to the encoding method corresponding to the target level of refinement to obtain a text embedding vector corresponding to the image style.

[0397] As an implementation of the present application, the third determining unit may further include:

[0398] a second determining subunit, configured to determine the target fineness of the target image as a low fineness when the attribute of the image style is a first type of style attribute;

[0399] The third determining subunit is configured to determine the target fineness of the target image as a high fineness when the attribute of the image style is the second type of style attribute.

[0400] As an implementation of the present application, the second encoding unit may further include:

[0401] A second encoding subunit is configured to encode the prompt word when the target fineness of the target image is high, to obtain a basic text embedding vector and a refined text embedding vector, wherein the basic text embedding vector is used to represent the basic semantics of the prompt word, and the refined text embedding vector is used to represent the detailed semantics of the prompt word;

[0402] The third encoding subunit is configured to encode the prompt word to obtain a basic text embedding vector when the target fineness of the target image is a low fineness.

[0403] As an implementation of the present application, the second conversion module 404 may further include:

[0404] a fourth determining unit, configured to determine an image generation model corresponding to the text embedding vector according to the model correspondence relationship;

[0405] A second conversion unit is configured to convert the text embedding vector using the image generation model to obtain a target image when there is only one image generation model;

[0406] The third conversion unit is used to cascade the at least two image generation models when the image generation models include at least two, and convert the text embedding vector using the at least two cascaded image generation models to obtain a target image.

[0407] As an implementation of the present application, the fourth determining unit may further include:

[0408] a fourth determination subunit, configured to determine, when the target refinement level of the image style is a high refinement level, a base model corresponding to the base text embedding vector and a refined model corresponding to the refined text embedding vector;

[0409] The third conversion unit may further include:

[0410] an embedding subunit for embedding the base text embedding vector into the base model and embedding the refined text embedding vector into the refined model when the target refinement level of the image style is a high refinement level;

[0411] A generating subunit, for obtaining a randomly generated noise image;

[0412] A first denoising subunit is configured to perform denoising on the noisy image using a basic model embedded with a basic text embedding vector to obtain a first latent feature image;

[0413] A second denoising subunit is configured to perform denoising on the first latent feature image using a refined model embedded with the refined text embedding vector to obtain a second latent feature image;

[0414] The second decoding subunit is configured to decode the second latent feature image to obtain a target image.

[0415] As an implementation of the present application, the identification module 402 may further include:

[0416] a third recognition unit, configured to perform semantic recognition on the prompt word, and, if a specific phrase exists in the prompt word, determine the image style of the target image as the image style corresponding to the specific phrase, where the specific phrase is used to indicate the image style of the prompt word;

[0417] The fourth recognition unit is configured to determine the image style of the target image as a realistic style when no specific phrase exists in the prompt words.

[0418] The image generation device provided in the embodiment of the present application can implement each step in the above-mentioned method embodiment. To avoid repetition, they will not be described here.

[0419] FIG7 shows a flow chart of another image generation method provided in some embodiments of the present application. The method can be applied to a vehicle's onboard computer or a cloud server connected to the vehicle, and the method includes the following steps:

[0420] S501: Acquire a drawing instruction input by a user, where the drawing instruction is used to instruct generation of a target image.

[0421] In some embodiments, an image generation task can convert a drawing instruction input by a user into a target image. The drawing instruction is used to describe the content of the target image that the user wants to generate. The drawing instruction can be a voice instruction, a text instruction, or an image instruction. The voice instruction can be a voice message spoken by the user, for example, "Generate a puppy." The text instruction can be a text message input by the user, for example, "Generate a tree." The image instruction can be an image input by the user.

[0422] Furthermore, the input voice or text commands can include both positive and negative instructions. Positive instructions indicate desired content in the target image, while negative instructions indicate undesirable content. For example, a positive instruction might be "an anime-style Chinese girl running on the beach, with flying seagulls and a vibrant rainbow behind her, creating a poetic and picturesque scene," while a negative instruction might be "pornographic, nude, ugly, deformed."

[0423] S502: Analyze the drawing instruction to determine the content category of the target image corresponding to the drawing instruction.

[0424] In some embodiments, in an image generation task, after receiving a drawing instruction, the drawing instruction can be analyzed, and based on the analysis result, the drawing instruction can be converted into corresponding text content, and then the text content can be semantically understood to determine the content category of each content in the target image.

[0425] Furthermore, the content of the target image can include the object in the target image, the background in the target image, the image style of the target image, the primary color of the target image, and the lighting and shadow effects of the target image. Content with similar feature information can be classified as belonging to the same content category. Feature information refers to the attributes of the content generated by the prompt word, such as object attributes, style attributes, and color attributes.

[0426] Content categories include subject categories and style categories. Subject categories refer to the type of subject in the target image. For example, subject categories can include people, scenery, animals, food, vehicles, and so on. If the subject in the drawing instruction is a cat or a dog, the subject category of the drawing instruction is animals. If the subject in the drawing instruction is coffee or bread, the subject category of the drawing instruction is food. Style categories generally refer to the visual appearance and stylistic characteristics of the image. Common style categories include realistic styles and artistic styles.

[0427] Realistic styles can include realistic figures, landscapes, animals, and architecture. They emphasize authentic color expression and strive to restore the true colors and lighting effects of figures. Therefore, they have high requirements for image quality and are very demanding on image detail. Artistic styles can include comics, oil paintings, ink paintings, and line drawings. These styles focus more on the presentation of artistic style and require less detail than realistic styles.

[0428] S503 , determining the sampler algorithm and parameters corresponding to the content category from the sampler database according to the content category. The sampler database includes the correspondence between the content category and the sampler algorithm, and the parameters corresponding to each sampler algorithm.

[0429] In some embodiments, the sampler database includes multiple different sampler algorithms, as well as a mapping between each sampler algorithm and content category. Each sampler algorithm also includes a mapping between parameters. The sampler algorithms and parameters are used to control the intensity, method, and number of iterations of noise reduction. There are many types of sampler algorithms, such as DDIM, Euler, and DPM++2M, and different sampler algorithms have different characteristics.

[0430] As an optional embodiment, determining the sampler algorithm and parameters corresponding to the content category from the sampler database according to the content category includes:

[0431] Determining an image generation model corresponding to the content category based on a preset first correspondence relationship, where the first correspondence relationship is a mapping relationship between the content category and the image generation model;

[0432] A sampler algorithm having a mapping relationship with a content category and an image generation model is determined in a sampler database, and parameters corresponding to the sampler algorithm are determined; the sampler database includes a first mapping relationship, and the first mapping relationship includes a mapping relationship between the content category and the image generation model and the sampler algorithm.

[0433] In some embodiments, each image generation model is used to implement image generation tasks in one or more vertical domains. Therefore, a first correspondence between content categories and image generation models, and a first mapping between content categories, image generation models, and sampler algorithms can be pre-set in the sampler database based on the vertical domain corresponding to each image generation model.

[0434] As an optional embodiment, the content category includes a subject category and a style category, the image generation model includes a basic image generation model and a specific object generation model, and determining the image generation model corresponding to the content category includes:

[0435] Determining the basic image generation model serial number and the specific object generation model serial number corresponding to the subject category and the style category; the first correspondence includes a mapping relationship between the subject category and the style category and the basic image generation model serial number and the specific object generation model serial number, wherein the basic image generation model serial number is used to identify the corresponding basic image generation model, and the specific object generation model serial number is used to identify the corresponding specific object generation model;

[0436] A sampler algorithm that has a mapping relationship between content categories and image generation models is determined in the sampler database, including:

[0437] Query the first mapping relationship table in the sampler database to determine the sampler algorithm corresponding to the subject category, style category, basic image generation model number and specific object generation model number, wherein the first mapping relationship table includes the first mapping relationship between the subject category, style category, basic image generation model number and specific object generation model number and the sampler algorithm.

[0438] Drawing instructions can be used to guide the image generation model to perform multiple iterative denoising processes on random noise images to obtain the final target image. The sampler is used to load into the image generation model during the iterative denoising process of the image generation model to control the intensity of denoising, denoising method and number of iterations.

[0439] For example, multiple different basic image generation models and multiple different specific object generation models can be set. During model training, each basic image generation model is trained using a dataset of the same subject category, and each specific object generation model is trained using a dataset of the same style category. In this way, a first correspondence between content categories and image generation models can be defined in the sampler database based on the aforementioned relationship.

[0440] The subject category of the drawing instruction can then be determined, and the base image generation model corresponding to the subject category can be determined based on the first correspondence. The specific object generation model corresponding to the style category can also be determined based on the first correspondence. The sampler is determined by the subject category, style category, and the selected base image generation model and specific object generation model.

[0441] In addition, the subject and subject category of the drawing instruction can be obtained by extracting features from the drawing instruction and then processing the extracted features through common neural network structures such as pooling, convolution, and residual. Then, through the pre-defined first correspondence, the basic image generation model serial number and the specific object generation model serial number corresponding to the subject category and style category are queried. Specifically, multiple different basic image generation models and multiple different specific object generation models can be set. Each basic image generation model has a unique corresponding basic image generation model serial number, and each specific object generation model has a unique corresponding specific object generation model serial number. The first correspondence includes the mapping relationship between the subject category and the basic image generation model, and the mapping relationship between the style category and the specific object generation model.

[0442] As an optional embodiment, querying the first mapping relationship table in the sampler database to determine the sampler algorithm corresponding to the subject category, style category, basic image generation model number, and specific object generation model number includes:

[0443] The subject category, style category, basic image generation model number, and specific object generation model number are combined as a sampler retrieval condition set;

[0444] According to the sampler retrieval condition set, the first mapping relationship table is queried to determine the sampler algorithm corresponding to the sampler retrieval condition set.

[0445] S504: Convert the drawing instruction into a target image using a sampler algorithm and parameters.

[0446] In some embodiments, a sampler can be constructed based on a determined sampler algorithm and parameters, and the determined sampler can be loaded into an image generation model. Then, a random noise image can be obtained, and the drawing instructions can be converted into a text embedding vector. The text embedding vector is embedded into the image generation model loaded with the sampler, and the image generation model is guided to perform multiple iterations of denoising on the noise image to obtain a latent feature image, which is then decoded into a target image visible to the naked eye.

[0447] In an embodiment of the present application, the drawing instructions input by the user are analyzed, and the content category of the target image corresponding to the drawing instructions is determined based on the analysis results. Then, the sampler algorithm and parameters corresponding to the content category are queried in the sampler database based on the content category, and the drawing instructions are converted into the target image using the sampler algorithm and parameters. In this way, the conversion from the drawing instructions to the target image can be completed by automatically selecting the appropriate sampler algorithm and parameters based on the content category of the drawing instructions. Compared with the prior art, the present application dynamically selects samplers with different characteristics based on the attributes of the drawing instructions for different Wensheng drawing tasks, avoiding the standardized process of applying one sampler to all Wensheng drawing tasks. Different samplers are adapted to different Wensheng drawing tasks, thereby improving the quality of image generation.

[0448] As an optional embodiment, analyzing the drawing instruction to determine the content category of the target image corresponding to the drawing instruction includes:

[0449] In the case of multiple drawing instructions, the multiple drawing instructions are spliced ​​into a drawing instruction group;

[0450] Analyzing the drawing instruction group to determine a content category set corresponding to the drawing instruction group, wherein a category element in the content category set is a content category of a target image corresponding to each drawing instruction in the drawing instruction group;

[0451] Determine the sampler algorithm and parameters corresponding to the content category from the sampler database according to the content category, including:

[0452] A sampler algorithm set is determined from a sampler database according to the content category set, and the sampler elements in the sampler algorithm set correspond one-to-one to the category elements in the content category set.

[0453] In some embodiments, a user may input multiple drawing instructions in the same batch. These multiple drawing instructions may be combined into a drawing instruction group. All drawing instructions in the drawing instruction group may be analyzed simultaneously to determine the content category corresponding to each drawing instruction, thereby obtaining multiple content categories. A content category set containing these multiple content categories is then generated, and the content category corresponding to each drawing instruction is a category element within the content category.

[0454] For the content category set, the sampler algorithm corresponding to each category element can also be determined from the sampler database to obtain multiple sampler algorithms, and the multiple sampler algorithms can also be combined into a sampler algorithm set.

[0455] In this way, when multiple drawing instructions are input in the same batch, the sampler algorithms corresponding to the drawing instructions can be screened and output simultaneously, thereby improving the efficiency of image generation.

[0456] As an optional embodiment, converting the drawing instruction into a target image using a sampler and an image generation model includes:

[0457] Load the samplers into the base image generation model and the object-specific generation model;

[0458] Mounting the specific object generation model to the basic image generation model to obtain the image generation model;

[0459] Use image generation models to convert drawing instructions into target images.

[0460] In some embodiments, since the sampler is used to control the intensity, method, and number of iterations of denoising during the iterative process of repeated denoising in the image generation model, it is necessary to load it into the basic image generation model and the specific object generation model separately before the image generation model is run. The specific object generation model can then be spliced ​​onto the basic image generation model, and the parameter weights of the specific object generation model and the basic image generation model can be adjusted to obtain the image generation model, and the image generation model can be used to complete the text generation task. In this way, the basic image generation model and the specific object generation model can be used together to complete the conversion of the drawing instructions into the target image.

[0461] Through the above method, the specific object generation model can be mounted on the basic image generation model, thereby completing the actual deployment and reasoning of the image generation model in the text-image task.

[0462] As an optional embodiment, the sampler includes a single sampling algorithm and a number of iterations, and uses an image generation model to convert the drawing instruction into a target image, including:

[0463] Get a randomly generated noise image;

[0464] Encode the drawing instructions to obtain the text embedding vector;

[0465] The text embedding vector is embedded into the image generation model, and the image generation model is used to perform N denoising processes on the noisy image according to the single sampling algorithm and parameters to obtain the potential feature image, where N is the number of iterations;

[0466] The latent feature image is decoded to obtain the target image.

[0467] In some embodiments, a randomly generated noise image can be obtained, and the text content of the drawing instruction can be mapped to a high-dimensional vector space to obtain a text embedding vector. After obtaining the text embedding vector, the text embedding vector and the noise image can be input into an image generation model together. The text embedding vector is used to guide the image generation model to perform N iterative denoising processes on the noise image according to a single sampling algorithm to obtain a latent feature image. The latent feature image can then be decoded by utilizing an autovariational encoder image generation model to obtain a target image that can be recognized by the human eye.

[0468] In this embodiment, the conversion from the drawing instruction to the target image can be completed accurately.

[0469] Based on the image generation method provided in the above embodiment, the present application also provides a specific implementation of an image generation device. Please refer to the following embodiment.

[0470] First, referring to FIG8 , another image generating device 600 provided in some embodiments of the present application includes the following modules:

[0471] A fifth acquisition module 601 is configured to acquire a drawing instruction input by a user, where the drawing instruction is used to instruct generation of a target image;

[0472] An analysis module 602 is configured to analyze the drawing instruction and determine the content category of the target image corresponding to the drawing instruction;

[0473] A third determining module 603 is configured to determine, from a sampler database according to the content category, a sampler algorithm and parameters corresponding to the content category, wherein the sampler database includes a correspondence between content categories and sampler algorithms, and parameters corresponding to each sampler algorithm;

[0474] The third conversion module 604 is configured to convert the drawing instruction into a target image using a sampler algorithm and parameters.

[0475] As an implementation of the present application, the third determining module 603 may further include:

[0476] a fifth determining unit, configured to determine an image generation model corresponding to the content category based on a preset first correspondence relationship, where the first correspondence relationship is a mapping relationship between the content category and the image generation model;

[0477] The sixth determination unit is used to determine a sampler algorithm that has a mapping relationship with the content category and the image generation model in the sampler database, and determine the parameters corresponding to the sampler algorithm; the sampler database includes a first mapping relationship, and the first mapping relationship includes a mapping relationship between the content category and the image generation model and the sampler algorithm.

[0478] As an implementation of the present application, the fifth determining unit may further include:

[0479] A second query subunit is configured to determine a base image generation model serial number and a specific object generation model serial number corresponding to a subject category and a style category; the first correspondence includes a mapping relationship between the subject category and the style category and the base image generation model serial number and the specific object generation model serial number, wherein the base image generation model serial number is used to identify the corresponding base image generation model, and the specific object generation model serial number is used to identify the corresponding specific object generation model;

[0480] The sixth determining unit may further include:

[0481] The third query sub-unit is used to query the first mapping relationship table in the sampler database to determine the sampler algorithm corresponding to the subject category, style category, basic image generation model number and specific object generation model number, wherein the first mapping relationship table includes the first mapping relationship between the subject category, style category, basic image generation model number and specific object generation model number and the sampler algorithm.

[0482] As an implementation of the present application, the third query subunit may also be used to:

[0483] The subject category, style category, basic image generation model number, and specific object generation model number are combined as a sampler retrieval condition set;

[0484] According to the sampler retrieval condition set, the first mapping relationship table is queried to determine the sampler algorithm corresponding to the sampler retrieval condition set.

[0485] As an implementation of the present application, the analysis module 602 may include:

[0486] a splicing unit, for splicing the plurality of drawing instructions into a drawing instruction group when there are multiple drawing instructions;

[0487] an analyzing unit, configured to analyze the drawing instruction group and determine a content category set corresponding to the drawing instruction group, wherein a category element in the content category set is a content category of a target image corresponding to each drawing instruction in the drawing instruction group;

[0488] The determination module 203 may further include:

[0489] The seventh determining unit is configured to determine a sampler algorithm set from the sampler database according to the content category set, wherein the sampler elements in the sampler algorithm set correspond one-to-one to the category elements in the content category set.

[0490] As an implementation of the present application, the third conversion module 604 may include:

[0491] a third acquiring unit, configured to acquire a randomly generated noise image;

[0492] A third encoding unit is used to encode the drawing instruction to obtain a text embedding vector;

[0493] A second denoising unit is configured to embed the text embedding vector into the image generation model, and perform denoising on the noisy image N times according to the single sampling algorithm and parameters using the image generation model to obtain a latent feature image, where N is the number of iterations;

[0494] The second decoding unit is used to decode the potential feature image to obtain a target image.

[0495] As an implementation of the present application, the second noise reduction unit may also be used to:

[0496] For the i-th denoising process among the N denoising processes, obtain an intermediate image generated by the image generation model, and determine the noise standard deviation of the intermediate image based on the parameters, where the intermediate image is a latent space image obtained by denoising the noisy image after the (i-1)th denoising process, and i is any positive integer less than or equal to N;

[0497] Configure the noise standard deviation to the single sampling algorithm;

[0498] Use the image generation model to perform a single denoising process on the intermediate image according to the configured single sampling algorithm until N denoising processes are completed.

[0499] As an implementation of the present application, the second noise reduction unit may also be used to:

[0500] Get the minimum and maximum noise standard deviations of the sampler algorithm;

[0501] When i is less than N, the parameter, the minimum noise standard deviation, the maximum noise standard deviation, and i are input into the noise determination formula to calculate the noise standard deviation;

[0502] When i is equal to N, the noise standard deviation is determined to be 0.

[0503] The image generation device provided in the embodiment of the present invention can implement each step in the above method embodiment, and will not be described again here to avoid repetition.

[0504] Figure 9 shows a flow chart of a model mounting method provided in some embodiments of the present application. The method can be applied to a vehicle's onboard computer or a cloud server connected to the vehicle, and the method includes the following steps:

[0505] S701, obtaining a model mounting instruction, where the model mounting instruction is used to instruct to mount a target mounting model to an application main model;

[0506] In some embodiments, the application master model is responsible for the core aspects of model reasoning. The application master model can mount one mounted model at a time. A mounted model can extend the functionality of a particular aspect of the application master model. Different mounted models mounted on the application master model can extend different aspects of the application master model's functionality. A mounted model is an accessory module or component at a certain model level of the master model, typically mounted onto the master model by connection or embedding. A model mount instruction is used to instruct the target mounted model to be mounted onto the application master model.

[0507] Among them, the main model can be a TensorRT model in PyTorch (Baidu deep learning framework) format, and the mounted model can be a LoRA model.

[0508] S702, in response to the model mounting instruction, mounting the target mounting model to the application main model to obtain the application combination model, and adjusting the model weight in the target mounting model to a null value;

[0509] In some embodiments, in response to a model mount instruction, the target mounted model can be mounted onto the application master model, achieving a fusion of the application master model and the target mounted model to form an application composite model. The model weights in the target mounted model are inherent properties of the mounted model and are used to influence the target mounted model's inference algorithm, thereby affecting the target mounted model's response to input data. After the target mounted model is mounted onto the application master model, the model weights in the application composite model need to be adjusted to null values.

[0510] S703 , obtaining a target weight file corresponding to the target mounting model, inputting the target weight file into the target mounting model in the application combination model, and updating the null value to the target weight parameter of the target weight file.

[0511] In some embodiments, the mount model includes a target mount model. A correspondence between each weight file and the mount model can be set in advance. Upon detecting that a target mount model is mounted on the application main model, a target weight file corresponding to the target mount model can be obtained based on the correspondence and input into the mounted target mount model. The weight file includes the weight parameters obtained in advance.

[0512] Specifically, the target weight file is used as an input to the target mount model in the application combination model, adjusting the model weight of the target mount model in the application combination model from a null value to the target weight parameter in the target weight file. This allows the weight parameters of the mount model in the combination model to be dynamically switched during the application of the target mount model.

[0513] In an embodiment of the present application, when the mount model is mounted on the main model for application, the model weight in the target mount model can be adjusted to a null value, and the target weight file corresponding to the target mount model is input into the target mount model in the application combination model to update the null value to the target weight parameter of the target weight file. In this way, during the application process of the target mount model, the weight you want to give to the target mount model can be input into the target mount model in the form of input, so as to realize the real-time update of the weight of the target mount model. Compared with the relevant technology, this method of inputting the model weight into the mount model in the form of a file can improve the processing speed compared to calling the interface for weight update, and will not cause the processing speed performance of the application combination model to regress after the mount model is mounted, thereby accurately and efficiently realizing the dynamic replacement of the weight parameters of the mount model and improving the operating efficiency of the combination model.

[0514] As an optional embodiment, before S703, the following steps may also be included:

[0515] Get the target weight parameters of the target mount model;

[0516] Convert the target weight parameters of the target mount model into the target weight file corresponding to the target mount model that conforms to the input format;

[0517] The above S703 may include:

[0518] Get the target weight file corresponding to the target mount model that conforms to the input format.

[0519] As an optional embodiment, converting the target weight parameters of the target mounting model into a target weight file corresponding to the target mounting model that conforms to the input format includes:

[0520] Convert the model weight format of the target mounted model from matrix format to binary format;

[0521] Write the model weights in binary format to the preset file template to obtain the target weight file that conforms to the input format.

[0522] In some embodiments, the model weights of the target mounted model can be converted from a weight matrix in matrix format to a binary format, and the model weights in binary format are written into a preset file template, which is a file template that conforms to the input format, to obtain a target weight file including the mounted model weights, which conforms to the input format of the model, wherein the input format is a file format allowed as model input.

[0523] In this way, the mounted model weights can be deployed in different application environments and architectures in the form of files, which improves the flexibility of the application of mounted model weights.

[0524] As an optional embodiment, obtaining a target weight parameter of a target mounting model includes:

[0525] When there are at least two mount models of the target category, the at least two mount models are respectively mounted into the training main model to obtain at least two training combination models, wherein the at least two mount models include the target mount model;

[0526] Convert the model formats of at least two trained combined models to a common format;

[0527] Match the attribute features of each node in at least two training combination models in a common format, and determine and obtain the model weight of the mounted model in each training combination model based on the matching results of the attribute features. The nodes include modules and mounted models in the training main model.

[0528] In some embodiments, the training master model is the main model in the training process. Different mounting models can be mounted on the same training master model to obtain different training combination models. In different training combination models, since the main models in the training combination models are consistent, the attribute characteristics of each module on the main model are consistent, while the attribute characteristics of different mounting models in different training combination models are inconsistent. Therefore, the attribute characteristics of each node in at least two training combination models can be matched, and the nodes with different attribute characteristics in any two training combination models are determined as the nodes where the mounting model is located. Among them, the characteristics of each node may include node weight, node level, node bias, etc.

[0529] Specifically, before performing attribute feature matching, the formats of all trained combination models can be exported to a common format. For example, the common format can be ONNX (Open Neural Network Exchange). ONNX (Open Neural Network Exchange) is an open deep learning model representation standard that can achieve attribute feature matching of models across platforms and frameworks.

[0530] As an optional embodiment, the attribute features include node weights, matching the attribute features of each node in at least two training combination models in a common format, and determining and obtaining the model weight of the mounted model in each training combination model based on the matching results of the attribute features include:

[0531] Based on the model structures of the at least two training combination models, matching the node weights of nodes having the same hierarchical structure in the at least two training combination models;

[0532] In the case where there is an abnormal node on the first training combination model, the node weight of the abnormal node is determined as the model weight of the mounted model in the first training combination model, and the model weight of the mounted model in the first training combination model is obtained, wherein the abnormal node is a node with the same hierarchical position and different node weights in the first training combination model and the second training combination model, wherein the first training combination model is any one of the at least two training combination models, and the second training combination model is any one of the at least two training combination models other than the first training combination model.

[0533] In some embodiments, since the model structure of the main model in different training combination models is consistent, and the connection method between the mounted model and the main model in different training combination models is consistent, the model structure of different training combination models is consistent. Since the model weights of different mounted models are different, the weights of each node in each training combination model can be traversed and compared according to the model structure of the training combination model. The ONNX nodes with different weights but the same position in each training combination model can be marked as abnormal nodes in the training combination model. The abnormal node is the mounted model in the training combination model, and then the model weight of the mounted model in ONNX format can be obtained.

[0534] For example, taking the first training combination model and the second training combination model as an example, the first training combination model is any one of the at least two training combination models, and the second training combination model is any one of the at least two training combination models other than the first training combination model. The node weights of each node in each level can be traversed according to the model structure of the first training combination model and the second training combination model, and the nodes in the first training combination model and the second training combination model with the same position but different node weights are determined as the mounting models in the first training combination model. Similarly, the nodes in the second training combination model with the same position but different node weights as the first training combination model are determined as the mounting models in the second training combination model.

[0535] Through the above method, the mounted model can be accurately and quickly queried from the training combination model.

[0536] As an optional embodiment, the target mounting model includes at least one model unit, the weight file includes at least one unit weight, and each model unit corresponds to a unit weight. Inputting the target weight file into the target mounting model in the application combination model includes:

[0537] Add at least one identity unit to the target mount model. The output of the identity unit is equal to the input of the identity unit. Each identity unit corresponds to a model unit.

[0538] Each unit weight in at least one unit weight is input into a corresponding model unit through a corresponding identity unit.

[0539] In some embodiments, in the process of inputting the weight file into the combined model, the input of the model weight can be completed by adding an identity unit in the mounted model and inputting the unit weight of each model unit into the corresponding model unit through the identity unit.

[0540] Among them, the identity unit can be an Identity unit. The Identity unit is a special glue unit. The input of the unit is equal to the output. A bypass path can be introduced through the Identity unit to allow information to pass directly without additional transformation.

[0541] Through the above method, the target weight parameters can be accurately and quickly introduced into the mounting model of the combined model in the form of file input.

[0542] As an optional embodiment, after inputting the weight file into the mounting model in the combined model to update the weight parameters of the mounting model, the method may further include:

[0543] Receive input data of the mounted model;

[0544] The input data is inferred based on the updated weight parameters to obtain the output data of the mounted model.

[0545] In some embodiments, in a combined model, the mounted model can, with minimal power consumption, assist the main model in accurately understanding and generating images for special objects or certain image styles that the main model cannot understand or support. The mounted model cannot operate independently, so it must be attached to the main model to obtain the final image generation model.

[0546] Specifically, after selecting the mount model, a specific object module that matches the function of the mount model can be determined in the main model, and then the mount model can be connected to the main model according to the connection method of the specific object module.

[0547] Specifically, a specific object module in the main model that matches the function of the mounted model can be determined, and the input data originally to be input into the specific object module can be input into the specific object module and the mounted model at the same time. After that, the specific object module can obtain the module output data in response to the input data; the mounted model can also obtain the output data of the mounted model in response to the input data based on the updated weight parameters. The module output data and the output data of the mounted model can be added to obtain fused data, and the fused data can be used as the model output of the level where the mounted model is located in the combined model, and output to the next level, thereby completing the mounting of the mounted model.

[0548] This embodiment uses the mounted model to adjust the parameters of the main model, thereby improving the image generation accuracy of the main model with less power consumption.

[0549] Based on the model mounting method provided in the above embodiment, the present application also provides a specific implementation of a model mounting device. Please refer to the following embodiment.

[0550] First, referring to FIG10 , a model mounting device 800 provided in some embodiments of the present application includes the following modules:

[0551] The sixth acquisition module 801 is used to acquire a model mounting instruction, where the model mounting instruction is used to instruct to mount the target mounting model to the application main model;

[0552] The second mounting module 802 is configured to mount the target mounting model to the application main model in response to the model mounting instruction to obtain the application combination model, and adjust the model weight in the target mounting model to a null value;

[0553] The input module 803 is used to obtain the target weight file corresponding to the target mounting model, input the target weight file into the target mounting model in the application combination model, and update the null value to the target weight parameter of the target weight file.

[0554] As an implementation of the present application, the model mounting device 200 may further include:

[0555] The parameter acquisition module is used to obtain the target weight parameters of the target mounting model;

[0556] A fourth conversion module is used to convert the target weight parameters of the target mounting model into a target weight file corresponding to the target mounting model that conforms to the input format;

[0557] The input module 803 may also be used to:

[0558] Get the target weight file corresponding to the target mount model that conforms to the input format.

[0559] As an implementation of this application, the above-mentioned parameter acquisition module may further include:

[0560] A training unit is configured to, when there are at least two mount models of the target category, mount the at least two mount models respectively into the training main model to obtain at least two training combination models, wherein the at least two mount models include the target mount model;

[0561] a fourth conversion unit, configured to convert model formats of at least two training combined models into a common format;

[0562] A matching unit is used to match the attribute features of each node in at least two training combination models in a common format, and determine and obtain the model weight of the mounted model in each training combination model based on the matching results of the attribute features. The nodes include modules and mounted models in the training main model.

[0563] As an implementation of the present application, the matching unit may further include:

[0564] a matching subunit, configured to match node weights of nodes having the same hierarchical structure in at least two training combination models based on the model structures of the at least two training combination models;

[0565] The fifth determination subunit is used to determine the node weight of the abnormal node as the model weight of the mounted model in the first training combination model when there is an abnormal node on the first training combination model, and obtain the model weight of the mounted model in the first training combination model, wherein the abnormal node is a node with the same position in the hierarchical structure in the first training combination model and the second training combination model and different node weights, wherein the first training combination model is any one of the at least two training combination models, and the second training combination model is any one of the at least two training combination models other than the first training combination model.

[0566] As an implementation of the present application, the fourth conversion module may also be used to:

[0567] Convert the model weight format of the target mounted model from matrix format to binary format;

[0568] Write the model weights in binary format to the preset file template to obtain the target weight file that conforms to the input format.

[0569] As an implementation of the present application, the input module 803 may also be used to:

[0570] Add at least one identity unit to the target mount model. The output of the identity unit is equal to the input of the identity unit. Each identity unit corresponds to a model unit.

[0571] Each unit weight in at least one unit weight is input into a corresponding model unit through a corresponding identity unit.

[0572] The model mounting device provided in the embodiment of the present invention can implement each step in the above-mentioned method embodiment, and will not be described again here to avoid repetition.

[0573] FIG11 shows a schematic diagram of the hardware structure of an example device provided in some embodiments of the present application.

[0574] The text generating device, or the image generating device, or the model mounting device may include a processor 901 and a memory 902 storing computer program instructions.

[0575] Specifically, the processor 901 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0576] The memory 902 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 902 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 902 may include removable or non-removable (or fixed) media. Where appropriate, the memory 902 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 902 is a non-volatile solid-state memory.

[0577] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0578] The processor 901 implements any one of the image generation methods in the above embodiments by reading and executing computer program instructions stored in the memory 902 .

[0579] In one example, the text generation device, or the image generation device, or the model mounting device may further include a communication interface 903 and a bus 910. As shown in FIG11 , the processor 901, the memory 902, and the communication interface 903 are connected via the bus 910 and communicate with each other.

[0580] The communication interface 903 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0581] Bus 910 comprises hardware, software or both, and the parts of text generation device are coupled together.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more above these combination.In suitable situation, bus 310 can comprise one or more buses.Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.

[0582] The text generation device, or the image generation device, or the model mounting device can be based on the above embodiments, thereby realizing the combination of the above methods and apparatuses.

[0583] In addition, in combination with the method in the above embodiment, the embodiment of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by the processor, any one of the image generation methods in the above embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, they will not be described here. Among them, the above-mentioned computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc., which is not limited here.

[0584] In addition, an embodiment of the present application also provides a vehicle, including computer program instructions, which, when executed by a processor, can implement the steps and corresponding contents of the aforementioned method embodiment.

[0585] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0586] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. Programs or code segments can be stored in machine-readable media, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable media" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0587] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0588] Aspects of the present disclosure are described above with reference to the flowcharts and / or block diagrams of the methods, devices and vehicles according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of the boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more boxes in the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of the boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0589] The above is only a specific implementation method of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited to this. Any technician familiar with this technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of this application.

Claims

1. A method for generating an image, wherein: The method comprises: Obtaining a prompt word input by a user, wherein the prompt word is used to instruct to generate a target image; Determine the image generation model and the sampler corresponding to the prompt word according to a preset first correspondence relationship, wherein the first correspondence relationship is a correspondence relationship between the prompt word, the image generation model and the sampler, and the sampler is used to control the denoising method of the image generation model; Performing text expansion on the prompt word to obtain a final text; Loading the sampler into the image generation model to update the denoising parameters of the image generation model; The final text is converted into the target image using the image generation model after the noise reduction parameters are updated.

2. The image generation method according to claim 1, wherein: The image generation model includes a basic image generation model and a specific object generation model, the first corresponding relationship includes a model corresponding relationship and a sampler corresponding relationship, and the image generation model and the sampler corresponding to the prompt word are determined according to the preset first corresponding relationship, including: Performing semantic recognition on the prompt word to determine the image generation intention corresponding to the prompt word; Identify specific generation information contained in the image generation intention corresponding to the prompt word; the specific generation information includes a specific subject and / or a specific style; Determine the basic image generation model and the specific object generation model corresponding to the prompt word according to the model correspondence relationship, wherein the model correspondence relationship includes the correspondence relationship between the specific generation information and the specific object generation model, and the correspondence relationship between the image generation intention and the basic image generation model; According to the sampler correspondence, the samplers corresponding to the prompt word, the basic image generation model and the specific object generation model are determined, and the sampler correspondence includes the correspondence between the image generation intention, the specific generation information, the basic image generation model, the specific object generation model and the sampler.

3. The image generation method according to claim 2, wherein: The step of determining the samplers corresponding to the prompt word, the basic image generation model, and the specific object generation model according to the sampler correspondence relationship includes: Acquire identification information of the image generation intention, the specific generation information, the basic image generation model, and the specific object generation model; Determine the identification information as an initialization parameter of the sampler; The sampler corresponding to the initialization parameter is matched from the sampler correspondence, and the sampler correspondence includes a correspondence between the initialization parameter and the sampler.

4. The image generation method according to any one of claims 1 to 3, wherein: The step of performing text expansion on the prompt word to obtain a final text includes: Obtaining a model keyword corresponding to the image generation model, and adding the model keyword to the prompt word to obtain a first intermediate text; When there is a special phrase representing a special object in the prompt word, obtaining an object keyword corresponding to the special object, and adding the object keyword to the first intermediate text to obtain a second intermediate text; The second intermediate text is converted into the final text.

5. The image generation method according to claim 4, wherein: The converting the second intermediate text into the final text comprises: When it is detected that an informal phrase exists in the second intermediate text, searching for a colloquial phrase corresponding to the informal phrase according to a preset second correspondence relationship, where the second correspondence relationship is a correspondence relationship between the informal phrase and the colloquial phrase; The colloquial phrases are used to replace corresponding informal phrases in the second intermediate text to obtain the final text.

6. The image generation method according to any one of claims 1 to 5, wherein: The image generation model includes a basic image generation model and a specific object generation model. After determining the image generation model and the sampler corresponding to the prompt word according to the preset first corresponding relationship, the method further includes: Determine a specific object module in the basic image generation model that matches the function of the specific object generation model; Acquire input data of the specific object module, and input the input data into the specific object generation model; Acquire module output data obtained by the specific object module in response to the input data, and acquire model output data obtained by the specific object generation model in response to the input data; Merging the module output data and the model output data to obtain fused data; The fusion data is updated to model parameters corresponding to the specific object in the basic image generation model, so as to mount the specific object generation model into the basic image generation model.

7. The image generation method according to any one of claims 1 to 6, wherein: The acquiring model output data obtained by the specific object generation model in response to the input data includes: Obtaining a model weight in the specific object generation model; Inputting the model weight into the specific object generation model to update the weight parameters of the specific object generation model; The specific object generation model after updating the weight parameters is used to respond to the input data to obtain the model output data.

8. The image generation method according to claim 7, wherein: The specific object generation model includes at least one specific object generation node, the model weight includes at least one node weight, each specific object generation node corresponds to a node weight, and inputting the model weight into the specific object generation model includes: Adding at least one identity node in the specific object generation model, the output of the identity node is equal to the input of the identity node, and each identity node corresponds to one specific object generation node; Each node weight in the at least one node weight is input into the corresponding specific object generation node through the corresponding identity node.

9. The image generation method according to any one of claims 1 to 8, wherein: The sampler includes a single sampling algorithm and a number of iterations, and the image generation model updated with the denoising parameter is used to convert the final text into the target image, including: Get a randomly generated noise image; Encoding the final text using a text vectorization encoding method to obtain a basic text embedding vector; Embedding the basic text embedding vector into the image generation model, and using the image generation model to perform N denoising processes on the noisy image according to the single sampling algorithm to obtain a basic latent feature image, where N is the number of iterations; The basic latent feature image is decoded by image decoding to obtain a target image.

10. The image generation method according to claim 9, wherein: The method of decoding the basic potential feature image by image decoding to obtain the target image includes: When the image style of the target image is a realistic style, encoding the final text to obtain a refined text embedding vector; Embedding the refined text embedding vector into an image refinement model, and performing refinement and noise reduction processing on the basic latent feature image using the image refinement model embedded with the refined text embedding vector to obtain a refined latent feature image; The refined latent feature image is decoded to obtain the target image.

11. The image generation method according to claim 9, wherein: The step of performing N-times noise reduction processing on the noisy image using the image generation model according to the single sampling algorithm includes: For the i-th denoising process in the N denoising processes, an intermediate noise image is obtained, and a noise standard deviation of the intermediate noise image is obtained, wherein the intermediate noise image is a latent space image obtained by performing the (i-1)th denoising process on the noise image, and i is any positive integer less than or equal to N; Updating the parameters of the single sampling algorithm using the noise standard deviation; The image generation model embedded with the basic text embedding vector is used to perform a single denoising process on the intermediate noise image according to the single sampling algorithm after parameter update, until N denoising processes are completed.

12. The image generation method according to claim 11, wherein: The obtaining of the noise standard deviation of the intermediate noise image comprises: Obtaining the hyper parameters of the sampler, as well as the minimum value and maximum value of the noise standard deviation of the sampler; When i is less than N, the hyperparameter, the minimum value of the noise standard deviation, the maximum value of the noise standard deviation and i are input into a standard deviation calculation formula to obtain the noise standard deviation; When i is equal to N, the noise standard deviation is determined to be 0.

13. The image generation method according to any one of claims 1 to 10, wherein: The method further comprises: In a case where the attribute of the image style is a first type of style attribute, determining the target fineness of the target image to be a low fineness; In a case where the attribute of the image style is a second-type style attribute, the target fineness of the target image is determined to be a high fineness.

14. The image generation method according to claim 13, wherein: The adopting a text vectorization encoding method to encode the final text to obtain a basic text embedding vector includes: When the target fineness of the target image is a low fineness, the final text is encoded to obtain the basic text embedding vector; the basic text embedding vector is used to represent the basic semantics of the prompt word.

15. The image generation method according to claim 13, wherein: The encoding of the final text to obtain a refined text embedding vector includes: When the target fineness of the target image is a high fineness, the final text is encoded to obtain the basic text embedding vector and the refined text embedding vector; the refined text embedding vector is used to represent the detailed semantics of the prompt word.

16. The image generation method according to claim 1, wherein: The image generation model includes a basic image generation model and a specific object generation model, and the image generation model and sampler corresponding to the prompt word are determined, including: Analyze the prompt word to determine the content category of the target image corresponding to the prompt word, wherein the content category includes a subject category and a style category; Determine the basic image generation model corresponding to the subject category; the basic image generation model corresponds to the basic image generation model sequence number; Determine the specific object generation model corresponding to the style category; the specific object generation model corresponds to the specific object generation model sequence number; Query the first mapping relationship table in the sampler database to determine the samplers corresponding to the subject category, the style category, the basic image generation model number and the specific object generation model number, wherein the first mapping relationship table includes the first mapping relationship between the subject category, the style category, the basic image generation model number and the specific object generation model number and the sampler.

17. The image generation method according to claim 6, wherein: The step of mounting the specific object generation model into the basic image generation model includes: Mounting the specific object generation model to the basic image generation model to obtain an application combination model, and adjusting the model weight in the basic image generation model to a null value; Obtaining basic weight parameters of the basic image generation model; Converting the basic weight parameters of the basic image generation model into a basic weight file corresponding to the basic image generation model that conforms to an input format; The basic weight file is input into the basic image generation model in the application combination model, and the null value is updated to the basic weight parameter of the basic weight file.

18. A method for generating an image, wherein: The method comprises: Obtaining a prompt word input by a user, wherein the prompt word is used to instruct to generate a target image; By performing semantic recognition on the prompt word, the image style of the target image is determined according to the result of the semantic recognition; Encode the prompt word according to the encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style, wherein different image styles correspond to different encoding methods; The text embedding vector is converted according to the inference process corresponding to the text embedding vector to obtain the target image.

19. The image generation method according to claim 18, wherein: The step of encoding the prompt word according to the encoding method corresponding to the image style to obtain a text embedding vector corresponding to the image style includes: Determining a target refinement level of the target image according to an attribute of the image style; Querying the encoding method corresponding to the target precision level according to the encoding correspondence relationship set in advance; The prompt word is encoded according to an encoding method corresponding to the target fineness to obtain a text embedding vector corresponding to the image style.

20. A method for generating an image, wherein: The method comprises: Acquire a drawing instruction input by a user, where the drawing instruction is used to instruct generation of a target image; Analyzing the drawing instruction to determine the content category of the target image corresponding to the drawing instruction; Determining a sampler algorithm and parameters corresponding to the content category from a sampler database according to the content category, wherein the sampler database includes a correspondence between content categories and sampler algorithms, and parameters corresponding to each sampler algorithm; The drawing instructions are converted into the target image using the sampler algorithm and parameters.

21. The image generation method according to claim 20, wherein: The step of determining the sampler algorithm and parameters corresponding to the content category from a sampler database according to the content category includes: Determine an image generation model corresponding to the content category based on a preset first corresponding relationship, where the first corresponding relationship is a mapping relationship between the content category and the image generation model; A sampler algorithm having a mapping relationship with the content category and the image generation model is determined in the sampler database, and parameters corresponding to the sampler algorithm are determined; the sampler database includes a first mapping relationship, and the first mapping relationship includes a mapping relationship between the content category and the image generation model and the sampler algorithm.

22. A model mounting method, wherein: The method comprises: Obtain a model mounting instruction, where the model mounting instruction is used to instruct to mount the target mounting model to the application main model; In response to the model mounting instruction, the target mounting model is mounted to the application main model to obtain an application combination model, and the model weight in the target mounting model is adjusted to a null value; A target weight file corresponding to the target mounting model is obtained, and the target weight file is input into the target mounting model in the application combination model, and the null value is updated to the target weight parameter of the target weight file.

23. The model mounting method according to claim 22, wherein: Before obtaining the target weight file corresponding to the target mounting model, the method includes: Obtaining a target weight parameter of the target mounting model; Converting the target weight parameter of the target mounting model into a target weight file corresponding to the target mounting model conforming to an input format; The step of obtaining a target weight file corresponding to the target mounting model includes: Obtain a target weight file corresponding to the target mounting model that conforms to the input format.

24. An image generating device, wherein: The device comprises: A first acquisition module, used to acquire a prompt word input by a user, wherein the prompt word is used to instruct to generate a target image; A first determination module, used to determine the image generation model and the sampler corresponding to the prompt word according to a preset first correspondence relationship, wherein the first correspondence relationship is a correspondence relationship between the prompt word, the image generation model and the sampler, and the sampler is used to control the denoising method of the image generation model; An expansion module, used for performing text expansion on the prompt word to obtain a final text; An updating module, used for loading the sampler into the image generation model to update the denoising parameters of the image generation model; The first conversion module is used to convert the final text into the target image by using the image generation model after the noise reduction parameter is updated.

25. A text generation device, wherein: The text generation device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image generating method according to any one of claims 1 to 12 is implemented.

26. A computer storage medium, wherein: The computer storage medium stores computer program instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 18.

27. A vehicle, wherein: The vehicle includes at least one of the following: The device as claimed in claim 24; The apparatus as claimed in claim 25; The computer storage medium of claim 26.

Citation Information

Patent Citations

  • Image generation model training method and device, equipment and storage medium

    CN116721334A

  • Image generation method and device, electronic equipment and storage medium

    CN116797684A

  • Image generation method and device, electronic equipment, storage medium and program product

    CN116958323A

  • Specific scene portrait generation method and device, storage medium and equipment

    CN116977461A

  • Image generation method and device and storage medium

    CN117197268A

Cited By

  • Drawing method based on AI model, computer equipment and medium

    CN120524545A

  • Generative steganography method and system based on diffusion model and semantic cue word

    CN120856837A

  • Building effect picture generation method based on multi-condition weighted fusion and block control

    CN120997368A

  • Image generation method and device, computer equipment, storage medium and program product

    CN121010661A

  • Data chart illustration method, system and equipment based on AIGC

    CN121074201A