Image generation method and device

By using language mapping sets to convert Chinese description text to English and coding, the problem of insufficient understanding of Chinese text graphics models is solved, and the accuracy of image generation and user satisfaction is improved.

CN120014111APending Publication Date: 2025-05-16ZHUHAI KINGSOFT OFFICE SOFTWARE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311535910.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The understanding and drawing ability of Chinese literary and artistic graphics models is weaker than that of English models, making it difficult to understand unusual Chinese vocabulary, resulting in low image generation quality.

Method used

By obtaining the description text and the preset language mapping set, the first language text corresponding to the description text is determined, the encoding process is performed, the first encoding feature and the second encoding feature are obtained, and the image generation process is performed to generate the target image.

Benefits of technology

It improves the Chinese literary and artistic graphic model's understanding of description text, enhances the matching degree between the target image and the description text, and improves the accuracy of image generation and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014111A_ABST
    Figure CN120014111A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device, and relates to the technical field of text processing and image generation, and the image generation method comprises the steps: determining a first language text corresponding to a description text according to the obtained description text and a preset language mapping set, and the description text comprises a second language text; respectively encoding the description text and the first language text to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text; and performing image generation processing on the first coding feature and the second coding feature to obtain a target image corresponding to the description text. The text understanding capability can be improved, the matching degree between the target image and the description text is further improved, the accuracy of the target image is improved, and the user satisfaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of text processing technology and image generation technology in multimodal technology, and in particular to an image generation method and device. Background Art

[0002] With the development of stable diffusion models, users can easily generate images of various styles through a text prompt. However, the current mainstream stable diffusion models are based on English as input. Since the amount of Chinese image and text data is far less than that of English image and text data, the understanding and drawing capabilities of Chinese text-generated image models are weaker than those of English models, and it is difficult to understand some uncommon words. Therefore, an effective solution is urgently needed to solve the above problems. Summary of the invention

[0003] In view of the problems existing in the prior art, an embodiment of the present invention provides an image generating method and device.

[0004] The present invention provides an image generation method, comprising:

[0005] Determining, according to the acquired description text and a preset language mapping set, a first language text corresponding to the description text, wherein the description text includes a second language text;

[0006] Encoding the description text and the first language text respectively to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text;

[0007] The first coding feature and the second coding feature are processed for image generation to obtain a target image corresponding to the description text.

[0008] According to an image generation method provided by the present invention, performing image generation processing on the first coding feature and the second coding feature to obtain a target image corresponding to the description text includes:

[0009] Concatenate the first coding feature and the second coding feature to obtain a third coding feature;

[0010] Inputting the third coding feature into the target denoising model for denoising to obtain a fourth coding feature;

[0011] The fourth coding feature is decoded to generate a target image corresponding to the description text.

[0012] According to an image generation method provided by the present invention, before inputting the third coding feature into the target denoising model for denoising to obtain the fourth coding feature, the method further includes:

[0013] A target denoising model is determined from a plurality of initial denoising models according to the description text and / or the first encoding feature.

[0014] According to an image generation method provided by the present invention, determining a target denoising model from a plurality of initial denoising models according to the description text and / or the first coding feature comprises:

[0015] Perform keyword detection on the description text to obtain a detection result; and determine a target denoising model from multiple initial denoising models based on the detection result; or,

[0016] Determine the classification score of each of the initial denoising models according to the first coding feature; determine the target denoising model from the multiple initial denoising models according to each of the classification scores; or,

[0017] Perform keyword detection on the description text to obtain a detection result; if the detection result is that the keyword is detected, determine a target denoising model from multiple initial denoising models based on the keyword; if the detection result is that the keyword is not detected, determine the classification score of each of the initial denoising models based on the first coding feature, and determine the target denoising model from the multiple initial denoising models based on each of the classification scores.

[0018] According to an image generation method provided by the present invention, determining the classification score of each of the initial denoising models according to the first coding feature includes:

[0019] Obtaining a style identifier of each of the initial denoising models;

[0020] According to each of the style identifiers and the first coding feature, a classification score of each of the initial denoising models is determined.

[0021] According to an image generation method provided by the present invention, determining the classification score of each of the initial denoising models according to each of the style identifiers and the first coding feature includes:

[0022] Calculating the matching degree between each of the style identifiers and the first coding feature;

[0023] Adding the matching degrees to obtain a total matching degree;

[0024] For each of the initial denoising models, a ratio of the matching degree corresponding to the initial denoising model to the total matching degree is determined as the classification score of the initial denoising model.

[0025] According to an image generation method provided by the present invention, determining a target denoising model from a plurality of initial denoising models according to each of the classification scores comprises:

[0026] determining a highest category score and a second highest category score among the category scores;

[0027] When the difference between the highest classification score and the second highest classification score is greater than a set threshold, determining the initial denoising model corresponding to the highest classification score as the target denoising model;

[0028] When the difference between the highest classification score and the second highest classification score is less than or equal to a set threshold, the universal denoising model in each of the initial denoising models is determined as the target denoising model.

[0029] According to an image generation method provided by the present invention, the decoding process of the fourth coding feature to generate a target image corresponding to the description text includes:

[0030] Calculating feature distribution according to the fourth coding feature;

[0031] Performing sampling processing on the feature distribution to obtain sampling features;

[0032] The sampled features are decoded to obtain a target image corresponding to the description text.

[0033] According to an image generation method provided by the present invention, the language mapping set includes a first language vocabulary and a second language vocabulary corresponding to at least one initial vocabulary;

[0034] The determining, based on the acquired description text and a preset language mapping set, a first language text corresponding to the description text includes:

[0035] Matching each second language vocabulary in the description text with the second language vocabulary corresponding to each of the initial vocabulary to determine the target vocabulary in each of the initial vocabulary;

[0036] The first language words corresponding to the target words are concatenated to obtain the first language text corresponding to the description text.

[0037] According to an image generation method provided by the present invention, encoding the description text and the first language text respectively to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text includes:

[0038] Performing vector conversion processing on each character unit in the description text to obtain a first coding vector corresponding to the description text; performing position coding processing on the first coding vector to obtain a second coding vector; determining a first hidden layer vector corresponding to the second coding vector based on a first attention mechanism; performing residual connection and normalization processing on the second coding vector and the first hidden layer vector to obtain a second hidden layer vector; determining a third hidden layer vector corresponding to the second hidden layer vector based on a first activation function; performing residual connection and normalization processing on the second hidden layer vector and the third hidden layer vector to obtain a first coding feature corresponding to the description text;

[0039] Perform vector conversion processing on each character unit in the first language text to obtain a third encoding vector corresponding to the first language text; perform position encoding processing on the third encoding vector to obtain a fourth encoding vector; based on a second attention mechanism, determine a fourth hidden layer vector corresponding to the fourth encoding vector; perform residual connection and normalization processing on the fourth encoding vector and the fourth hidden layer vector to obtain a fifth hidden layer vector; based on a second activation function, determine a sixth hidden layer vector corresponding to the fifth hidden layer vector; perform residual connection and normalization processing on the fifth hidden layer vector and the sixth hidden layer vector to obtain a second encoding feature corresponding to the first language text.

[0040] The present invention also provides an image generating device, comprising:

[0041] A first determination module is configured to determine a first language text corresponding to the description text according to the acquired description text and a preset language mapping set, wherein the description text includes a second language text;

[0042] an encoding module configured to encode the description text and the first language text respectively to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text;

[0043] The image generation module is configured to perform image generation processing on the first coding feature and the second coding feature to obtain a target image corresponding to the description text.

[0044] According to an image generating device provided by the present invention, the image generating module comprises:

[0045] a concatenation module, configured to concatenate the first coding feature and the second coding feature to obtain a third coding feature;

[0046] a denoising module, configured to input the third coding feature into a target denoising model for denoising to obtain a fourth coding feature;

[0047] The decoding module is configured to decode the fourth coding feature to generate a target image corresponding to the description text.

[0048] An image generating device according to the present invention further includes:

[0049] The second determination module is configured to determine a target denoising model from a plurality of initial denoising models according to the description text and / or the first encoding feature.

[0050] According to an image generating device provided by the present invention, the second determining module is further configured to:

[0051] Perform keyword detection on the description text to obtain a detection result; and determine a target denoising model from multiple initial denoising models based on the detection result; or,

[0052] Determine the classification score of each of the initial denoising models according to the first coding feature; determine the target denoising model from the multiple initial denoising models according to each of the classification scores; or,

[0053] Perform keyword detection on the description text to obtain a detection result; if the detection result is that the keyword is detected, determine a target denoising model from multiple initial denoising models based on the keyword; if the detection result is that the keyword is not detected, determine the classification score of each of the initial denoising models based on the first coding feature, and determine the target denoising model from the multiple initial denoising models based on each of the classification scores.

[0054] According to an image generating device provided by the present invention, the second determining module is further configured to:

[0055] Obtaining a style identifier of each of the initial denoising models;

[0056] According to each of the style identifiers and the first coding feature, a classification score of each of the initial denoising models is determined.

[0057] According to an image generating device provided by the present invention, the second determining module is further configured to:

[0058] Calculating the matching degree between each of the style identifiers and the first coding feature;

[0059] Adding the matching degrees to obtain a total matching degree;

[0060] For each of the initial denoising models, a ratio of the matching degree corresponding to the initial denoising model to the total matching degree is determined as the classification score of the initial denoising model.

[0061] According to an image generating device provided by the present invention, the second determining module is further configured to:

[0062] determining a highest classification score and a second highest classification score among the classification scores;

[0063] When the difference between the highest classification score and the second highest classification score is greater than a set threshold, determining the initial denoising model corresponding to the highest classification score as the target denoising model;

[0064] When the difference between the highest classification score and the second highest classification score is less than or equal to a set threshold, the universal denoising model in each of the initial denoising models is determined as the target denoising model.

[0065] According to an image generating device provided by the present invention, the decoding module is further configured as follows:

[0066] Calculating feature distribution according to the fourth coding feature;

[0067] Performing sampling processing on the feature distribution to obtain sampling features;

[0068] The sampled features are decoded to obtain a target image corresponding to the description text.

[0069] According to an image generation device provided by the present invention, the language mapping set includes a first language vocabulary and a second language vocabulary corresponding to at least one initial vocabulary;

[0070] The first determining module is further configured to:

[0071] Matching each second language vocabulary in the description text with the second language vocabulary corresponding to each of the initial vocabulary to determine the target vocabulary in each of the initial vocabulary;

[0072] The first language words corresponding to the target words are concatenated to obtain the first language text corresponding to the description text.

[0073] According to an image generating device provided by the present invention, the encoding module is further configured as follows:

[0074] Performing vector conversion processing on each character unit in the description text to obtain a first coding vector corresponding to the description text; performing position coding processing on the first coding vector to obtain a second coding vector; determining a first hidden layer vector corresponding to the second coding vector based on a first attention mechanism; performing residual connection and normalization processing on the second coding vector and the first hidden layer vector to obtain a second hidden layer vector; determining a third hidden layer vector corresponding to the second hidden layer vector based on a first activation function; performing residual connection and normalization processing on the second hidden layer vector and the third hidden layer vector to obtain a first coding feature corresponding to the description text;

[0075] Perform vector conversion processing on each character unit in the first language text to obtain a third encoding vector corresponding to the first language text; perform position encoding processing on the third encoding vector to obtain a fourth encoding vector; based on a second attention mechanism, determine a fourth hidden layer vector corresponding to the fourth encoding vector; perform residual connection and normalization processing on the fourth encoding vector and the fourth hidden layer vector to obtain a fifth hidden layer vector; based on a second activation function, determine a sixth hidden layer vector corresponding to the fifth hidden layer vector; perform residual connection and normalization processing on the fifth hidden layer vector and the sixth hidden layer vector to obtain a second encoding feature corresponding to the first language text.

[0076] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned image generation methods is implemented.

[0077] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the image generating method described in any one of the above is implemented.

[0078] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the image generating method described above is implemented.

[0079] The image generation method and device provided by the present invention determine the first language text corresponding to the description text according to the acquired description text and the preset language mapping set, wherein the description text contains the second language text; respectively encode the description text and the first language text to obtain the first encoding feature corresponding to the description text and the second encoding feature corresponding to the first language text; perform image generation processing on the first encoding feature and the second encoding feature to obtain the target image corresponding to the description text. By determining the second language text corresponding to the description text, some unrecognizable second language words are translated into the first language, and then the second encoding feature of the first language is obtained as a supplement to the description text, that is, as a supplement to the first encoding feature, thereby improving the ability to understand the description text, and then improving the matching degree between the target image and the description text, that is, improving the accuracy of the target image, which is conducive to improving user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0081] Figure 1 It is one of the flowcharts of the image generation method provided by the present invention;

[0082] Figure 2 This is the second flow chart of the image generation method provided by the present invention;

[0083] Figure 3 It is a schematic diagram of a flow chart of determining a target denoising model provided by the present invention;

[0084] Figure 4 It is a structural schematic diagram of the image generating device provided by the present invention;

[0085] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0086] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0087] In order to facilitate a clearer understanding of the embodiments of the present invention, some relevant background knowledge is first introduced as follows.

[0088] The stable diffusion model is a generative AI (generative artificial intelligence) model that can generate unique realistic images based on text and image prompts, namely the text-based image model.

[0089] With the development of text-generated graph models, users can easily generate images of various styles through a text prompt.

[0090] However, the current mainstream text-based graph models are all based on English as input. Since the amount of Chinese text-based graph data is far less than that of English text-based graph data, the understanding and drawing capabilities of Chinese text-based graph models are weaker than those of English models, making it difficult to understand some uncommon words.

[0091] In addition, when generating images, a denoising model is generally required to make the generated image more realistic. However, the images generated by the denoising model often have obvious advantages in specific scenes or image styles, and it is difficult to cover all scenes or image styles with one denoising model.

[0092] To this end, the present invention provides an image generation method and device, which determines the first language text corresponding to the description text according to the acquired description text and the preset language mapping set, and the description text contains the second language text; respectively encodes the description text and the first language text to obtain the first encoding feature corresponding to the description text and the second encoding feature corresponding to the first language text; performs image generation processing on the first encoding feature and the second encoding feature to obtain the target image corresponding to the description text. By determining the second language text corresponding to the description text, some unrecognized second language words are translated into the first language, and then the second encoding feature of the first language is obtained as a supplement to the description text, that is, as a supplement to the first encoding feature, so as to improve the ability to understand the description text, and then improve the matching degree between the target image and the description text, that is, improve the accuracy of the target image, which is conducive to improving user satisfaction.

[0093] Combine the following Figure 1-Figure 5 The image generation method and device of the present invention are described.

[0094] Figure 1 is one of the flow charts of the image generation method provided by the present invention, see Figure 1 As shown, it includes steps 101 to 103, wherein:

[0095] Step 101: Determine, based on the acquired description text and a preset language mapping set, a first language text corresponding to the description text, wherein the description text includes a second language text.

[0096] First of all, it should be noted that the execution subject of the present invention can be any electronic device that generates an image, for example, it can be any one of a smart phone, a smart watch, a desktop computer, a laptop computer, etc.

[0097] Specifically, the description text refers to the prompt text used to generate the image, which can be a sentence, several words or a paragraph, etc. The first language can be any language, such as Chinese, English, French, etc.; the second language refers to any language different from the first language; for example, the first language is English and the second language is Chinese. The description text can only contain the first language text, such as "the sky, white clouds and wild geese"; the description text can also be a mixed language text, for example, containing the second language text and other language text, the other language text refers to the language text different from the second language, which can be the first language text or other. For example, the description text contains Chinese text and English text, such as "SKY, white clouds and wild geese". The language mapping set refers to a set of language description sets of at least one vocabulary, and the language description set includes at least the first language vocabulary and the second language vocabulary.

[0098] In actual applications, before determining the first language text corresponding to the description text according to the acquired description text and the preset language mapping set, it is necessary to acquire the description text and acquire the language mapping set.

[0099] There are many methods for obtaining the description text. For example, a user uploads the description text through the upload page provided by the image generation platform, and accordingly, the execution subject obtains the description text. For another example, the execution subject receives an image generation instruction or a description text acquisition instruction, and accordingly, the execution subject obtains the description text from the storage area pointed to by the image generation instruction or the description text acquisition instruction. The present invention is not limited to this.

[0100] There are also many methods for obtaining language mapping sets. For example, some commonly used words in literary images, such as at least one of style words and specified words, are collected, and the first language words and second language words of each word are associated to form a language mapping set; wherein, the specified words refer to some words that cannot be recognized by the model for generating pictures based on the second language, that is, words whose semantics cannot be recognized, such as cyberpunk, graffiti, mosaic, etc. For another example, a user uploads a language mapping set through the upload page provided by the image generation platform, and accordingly, the execution subject obtains the language mapping set. For another example, the execution subject receives a language mapping set acquisition instruction, and accordingly, the execution subject obtains the language mapping set from the storage area pointed to by the language mapping set instruction. The present invention is not limited to this.

[0101] On the basis of obtaining the description text and the language mapping set, the description text is further matched with the language mapping set, that is, the second language vocabulary that exists in both the description text and the language mapping set is converted into first language vocabulary, and the first language vocabulary is spliced ​​to obtain the first language text corresponding to the description text.

[0102] Step 102: Encode the description text and the first language text respectively to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text.

[0103] Specifically, encoding is the process of converting information from one form or format to another, that is, the language translation process, such as converting natural language into machine language. The first encoding feature refers to the encoding feature corresponding to the second language; the second encoding feature refers to the encoding feature corresponding to the first language.

[0104] In practical applications, based on obtaining the description text and the first language text, the description text can be encoded to obtain a first encoding feature corresponding to the description text; and the first language text can be text encoded to obtain a second encoding feature corresponding to the first language text.

[0105] Exemplarily, the same encoder may be used to encode the description text and the first language text respectively, or two encoders may be used to encode the description text and the first language text in parallel.

[0106] It should be noted that the description text and the first language text may be encoded at the same time; the description text may be encoded first and then the first language text; or the first language text may be encoded first and then the description text. The present invention does not impose any limitation on this.

[0107] Step 103: Perform image generation processing on the first coding feature and the second coding feature to obtain a target image corresponding to the description text.

[0108] Specifically, the target image refers to the final generated image, and the format of the target image can be Portable Network Graphics (PNG) format, Joint Photographic Experts Group (JPEG) format, Tag Image File Format (TIFF), Bitmap (BMP) format, Graphics Interchange Format (GIF), or other formats that are currently or in the future. The present invention does not make any limitation to this.

[0109] In practical applications, based on the first coding feature and the second coding feature, image generation processing, such as decoding and denoising, can be performed on the first coding feature and the second coding feature to obtain a target image corresponding to the description text.

[0110] The image generation method provided by the present invention determines the first language text corresponding to the description text according to the acquired description text and the preset language mapping set, wherein the description text contains the second language text; performs encoding processing on the description text and the first language text respectively to obtain the first encoding feature corresponding to the description text and the second encoding feature corresponding to the first language text; performs image generation processing on the first encoding feature and the second encoding feature to obtain the target image corresponding to the description text. By determining the second language text corresponding to the description text, some unrecognized second language words are translated into the first language, and then the second encoding feature of the first language is obtained as a supplement to the description text, that is, as a supplement to the first encoding feature, thereby improving the ability to understand the description text, and then improving the matching degree between the target image and the description text, that is, improving the accuracy of the target image, which is conducive to improving user satisfaction.

[0111] In one or more optional embodiments of the present invention, the image generation process is performed on the first coding feature and the second coding feature to obtain the target image corresponding to the description text, and the specific implementation process may be:

[0112] Decoding the first coding feature and the second coding feature respectively to obtain a first initial image corresponding to the first coding feature and a second initial image corresponding to the second coding feature;

[0113] extracting a first image feature of the first initial image and a second image feature of the second initial image;

[0114] fusing the first image feature and the second image feature to obtain a comprehensive image feature;

[0115] According to the comprehensive image features, a target image corresponding to the description text is obtained.

[0116] In practical applications, the first coding feature representing the second language and the second coding feature representing the first language are decoded respectively to obtain the first initial image representing the second language and the second initial image representing the first language. Further, the first image feature of the first initial image is extracted, and the second image feature of the second initial image is extracted. The first image feature and the second image feature are then spliced ​​and fused, and the target image is generated according to the fused comprehensive image feature. In this way, the generation efficiency and accuracy of the target image can be guaranteed.

[0117] In one or more optional embodiments of the present invention, the image generation process is performed on the first coding feature and the second coding feature to obtain the target image corresponding to the description text, and the specific implementation process may be as follows:

[0118] Concatenate the first coding feature and the second coding feature to obtain a third coding feature;

[0119] Inputting the third coding feature into the target denoising model for denoising to obtain a fourth coding feature;

[0120] The fourth coding feature is decoded to generate a target image corresponding to the description text.

[0121] Specifically, the denoising model refers to a model for removing noise from data, which may be a U-net (U-Net or Unet) model, or other models capable of achieving denoising, and the present invention does not impose any limitation on this.

[0122] In practical applications, the first coding feature and the second coding feature may be concatenated to obtain a third coding feature.

[0123] For example, the shapes of the first coding feature and the second coding feature are both (B, N, C), and the first coding feature and the second coding feature are concatenated in the sequence dimension to obtain a new feature, namely, the third coding feature, and the shape of the third coding feature is (B, 2N, C). Among them, B refers to the number of input texts, preferably 1; N is a fixed length, preferably 77, and C is a feature.

[0124] Furthermore, the third coding feature is input into the target denoising model, and the target denoising model analyzes and denoises the third coding feature to obtain a denoised fourth coding feature.

[0125] Afterwards, the fourth encoded feature is input into the encoder for encoding, so as to obtain the target image corresponding to the description text. For example, the fourth encoded feature is input into a variational auto encoder (VAE), and the VAE decodes the fourth encoded feature into the target image.

[0126] In this way, by concatenating, denoising and decoding the first coding feature and the second coding feature, the understanding of the first language and the second language is integrated, so that the generated target image is more consistent with the description text, thereby improving the accuracy of the target image.

[0127] In one or more optional embodiments of the present invention, since images generated by different denoising models often have obvious advantages in different scenes or image styles, it is necessary to select a target denoising model that matches the image to be generated by the description text for denoising, that is, before inputting the third coding feature into the target denoising model for denoising and obtaining the fourth coding feature, it also includes determining the target denoising model from multiple initial denoising models based on the description text and / or the first coding feature. That is:

[0128] According to the description text, determining a target denoising model from a plurality of initial denoising models; or,

[0129] Determine a target denoising model from a plurality of initial denoising models according to the first coding feature; or,

[0130] A target denoising model is determined from a plurality of initial denoising models according to the description text and the first encoding feature.

[0131] Specifically, the initial denoising model refers to a preset model for removing noise. The model structures of different initial denoising models can be the same, for example, the model structures of all initial denoising models are Unet structures; the model structures of different initial denoising models can also be different.

[0132] In addition, different initial denoising models correspond to different scenes or image styles. For example, one initial denoising model corresponds to a portrait scene or image style, another initial denoising model corresponds to an animation scene or image style, and another initial denoising model corresponds to an ink painting scene or image style, and so on.

[0133] In practical applications, since the description text is the fundamental source of generating the target image, the description text can reflect the scene or image style of the target image to be generated. Therefore, the target denoising model can be determined from multiple initial denoising models based on the description text. In this way, the accuracy of determining the target denoising model can be guaranteed while improving the determination efficiency.

[0134] Since the first coding feature is obtained by encoding the description text, on the basis that the description text can reflect the scene or image style of the target image to be generated, the first coding feature can also reflect the scene or image style of the target image to be generated, therefore, the target denoising model can be determined from multiple initial denoising models according to the description text. In this way, the determination efficiency can be improved while ensuring the accuracy of determining the target denoising model.

[0135] Since both the description text and the first coding feature can reflect the scene or image style of the target image to be generated, the target denoising model can be determined from multiple initial denoising models by combining the description text and the first coding feature. Moreover, the description text is reflected from the text perspective, and the first coding feature is reflected from the feature perspective. Combining the description text and the first coding feature can further improve the accuracy and reliability of determining the target denoising model.

[0136] It should be noted that, in order to ensure that different initial denoising models correspond to different scenes or image styles, one initial denoising model can be trained using images and texts of one scene or image style, and another initial denoising model can be trained using images and texts of another scene or image style. Similarly, an initial denoising model dedicated to a certain scene or image style can be obtained.

[0137] In one or more optional embodiments of the present invention, the target denoising model is determined from a plurality of initial denoising models according to the description text, and the specific implementation process may be as follows:

[0138] Perform keyword detection on the description text to obtain a detection result;

[0139] And according to the detection result, a target denoising model is determined from a plurality of initial denoising models.

[0140] Specifically, keywords refer to words used to describe the style in a text.

[0141] In practical applications, keyword detection can be performed on the description text. For example, when the description text is a sentence, the description text is segmented to obtain at least one initial word, and then keyword detection is performed on the at least one initial word to determine the keywords therein; when the description text is at least one initial word, keyword detection is performed on the at least one initial word to determine the keywords therein.

[0142] Furthermore, based on the detection result, a target denoising model can be determined from multiple initial denoising models: if the detection result is that a keyword is detected, the keyword is matched with each initial denoising model. If the match is successful, the successfully matched initial denoising model is used as the target denoising model; if the match fails, the universal denoising model in each initial denoising model is used as the target denoising model; if the detection result is that the keyword is not detected, the universal denoising model in each initial denoising model is used as the target denoising model.

[0143] In this way, the accuracy and determination efficiency of the target denoising model can be improved, thereby improving the capability and accuracy of the Wensheng map.

[0144] In one or more optional embodiments of the present invention, the target denoising model is determined from multiple initial denoising models according to the first coding feature, and the specific implementation process may be as follows:

[0145] Determining a classification score of each of the initial denoising models according to the first coding feature;

[0146] According to each of the classification scores, a target denoising model is determined from a plurality of initial denoising models.

[0147] In practical applications, the classification scores of the initial denoising models can be determined based on the first coding feature and the scene or image style corresponding to each initial denoising model. Then, based on the classification scores of the initial denoising models, the target denoising model is determined from multiple initial denoising models. For example, the initial denoising model with the highest classification score among multiple initial denoising models is determined as the target denoising model. For example, when the classification scores of each initial denoising model are not higher than the score threshold, the common denoising model among multiple initial denoising models is determined as the target denoising model.

[0148] In this way, the accuracy and determination efficiency of the target denoising model can be improved, thereby improving the capability and accuracy of the Wensheng map.

[0149] In one or more optional embodiments of the present invention, the target denoising model is determined from a plurality of initial denoising models according to the description text and the first encoding feature, and the specific implementation process may be as follows:

[0150] Perform keyword detection on the description text to obtain a detection result;

[0151] If the detection result is that a keyword is detected, determining a target denoising model from a plurality of initial denoising models according to the keyword;

[0152] If the detection result is that no keyword is detected, the classification score of each of the initial denoising models is determined according to the first coding feature, and the target denoising model is determined from the multiple initial denoising models according to each of the classification scores.

[0153] In practical applications, keyword detection can be performed on the description text. Based on the detection result, the target denoising model can be determined from multiple initial denoising models: if the detection result is that the keyword is detected, the keyword is matched with each initial denoising model; if the match is successful, the initial denoising model that successfully matches is used as the target denoising model; if the match fails, the universal denoising model in each initial denoising model is used as the target denoising model, or the target denoising model is determined from multiple initial denoising models according to the first coding feature; if the detection result is that the keyword is not detected, the target denoising model is determined from multiple initial denoising models according to the first coding feature.

[0154] According to the first coding feature, a target denoising model is determined from multiple initial denoising models. The specific implementation process is: according to the first coding feature and the scene or image style corresponding to each initial denoising model, the classification score of each initial denoising model is determined. Then, based on the classification score of each initial denoising model, the target denoising model is determined from the multiple initial denoising models.

[0155] In this way, the target denoising model is first determined according to the description text. If the target denoising model cannot be determined according to the description text, the target denoising model is then determined according to the first encoding feature, instead of directly determining the denoising model as the target denoising model. This can further improve the accuracy and reliability of determining the target denoising model.

[0156] In one or more optional embodiments of the present invention, the classification score of each of the initial denoising models is determined according to the first coding feature, and the specific implementation process may be as follows:

[0157] Obtaining a style identifier of each of the initial denoising models;

[0158] According to each of the style identifiers and the first coding feature, a classification score of each of the initial denoising models is determined.

[0159] Specifically, the style identifier characterizes the scene or image style corresponding to the initial denoising model.

[0160] In practical applications, each initial denoising model has a corresponding scene or image style, i.e., a style identifier. On this basis, the style identifier of each initial denoising model can be first obtained. Then, for each initial denoising model, the classification score of the initial denoising model is determined according to the style identifier and the first coding feature of the initial denoising model. For example, the style identifier and the first coding feature are input into a classification score prediction model to obtain a classification score.

[0161] It should be noted that each initial denoising model has a corresponding scene or image style, i.e., a style identifier. On this basis, the keywords and the initial denoising models can be matched as follows: the target denoising model can be determined by matching the keywords with the style identifiers of the initial denoising models: that is, the keywords are matched with the style identifiers of the initial denoising models, and the initial denoising model corresponding to the successfully matched style identifier is used as the target denoising model. If all the matches fail, the universal denoising model in the initial denoising models is used as the target denoising model.

[0162] Exemplarily, the semantic similarity between the style identifier of each initial denoising model and the keyword is calculated: the initial denoising model corresponding to the style identifier with the highest semantic similarity is determined as the target denoising model; or, the initial denoising model corresponding to the style identifier with the highest semantic similarity higher than the similarity threshold is determined as the target denoising model. If there is no semantic similarity higher than the similarity threshold, the match fails, and the common denoising model among the initial denoising models is used as the target denoising model.

[0163] In one or more optional embodiments of the present invention, determining the classification score of each of the initial denoising models according to each of the style identifiers and the first encoding features includes:

[0164] Calculating the matching degree between each of the style identifiers and the first coding feature;

[0165] The matching degree is used as the classification score of the initial denoising model.

[0166] Specifically, the size of the classification score represents the degree of adaptation between the initial denoising model and the description text in this text generation process; the higher the classification score, the greater the degree of adaptation, and the lower the classification score, the smaller the degree of adaptation.

[0167] In practical applications, for each initial denoising model, the style identifier of the initial denoising model can be encoded to obtain a style coding feature, and then the feature similarity between the style coding feature and the first coding feature, i.e., the matching degree, is calculated; further, the matching degree is used as the classification score of the initial denoising model. In this way, the efficiency of determining the classification score can be improved.

[0168] In one or more optional embodiments of the present invention, determining the classification score of each of the initial denoising models according to each of the style identifiers and the first encoding features includes:

[0169] Calculating the matching degree between each of the style identifiers and the first coding feature;

[0170] Adding the matching degrees to obtain a total matching degree;

[0171] For each of the initial denoising models, a ratio of the matching degree corresponding to the initial denoising model to the total matching degree is determined as the classification score of the initial denoising model.

[0172] In practical applications, for each initial denoising model, the style identifier of the initial denoising model can be first encoded to obtain a style encoding feature, and then the feature similarity between the style encoding feature and the first encoding feature, i.e., the matching degree, is calculated. Each initial denoising model is traversed to obtain the matching degree between the style identifier of each initial denoising model and the first encoding feature.

[0173] Furthermore, each matching degree is normalized: first, the sum of all matching degrees is calculated to obtain the total matching degree. For each initial denoising model, the matching degree corresponding to the initial denoising model is divided by the total matching degree to obtain the classification score of the initial denoising model. Each initial denoising model is traversed to obtain the classification score of each initial denoising model.

[0174] In this way, by taking the ratio of the matching degree to the total matching degree as the classification score, the difference between the initial denoising models can be more clearly seen, so as to facilitate the selection of the target denoising model.

[0175] In one or more optional embodiments of the present invention, the target denoising model is determined from multiple initial denoising models according to each of the classification scores, and the specific implementation process may be as follows:

[0176] comparing each of the classification scores to a score threshold;

[0177] If there is a target classification score, determining the initial denoising model corresponding to the highest classification score among the classification scores as the target denoising model, wherein the target classification score is higher than the classification score of the score threshold;

[0178] If the target classification score does not exist, a common denoising model among the multiple initial denoising models is determined as the target denoising model.

[0179] In this way, the accuracy of classification can be guaranteed by the score threshold, thereby avoiding the situation where the classification scores of each initial denoising model are relatively average and the differences are small, resulting in the adopted target denoising model being inaccurate, thereby improving the accuracy of the target denoising model and the accuracy of the generated image.

[0180] In one or more optional embodiments of the present invention, the target denoising model is determined from multiple initial denoising models according to each of the classification scores, and the specific implementation process may be as follows:

[0181] determining a highest category score and a second highest category score among the category scores;

[0182] When the difference between the highest classification score and the second highest classification score is greater than a set threshold, determining the initial denoising model corresponding to the highest classification score as the target denoising model;

[0183] When the difference between the highest classification score and the second highest classification score is less than or equal to a set threshold, the universal denoising model in each of the initial denoising models is determined as the target denoising model.

[0184] Specifically, the set threshold refers to a preset difference threshold, preferably 0.3.

[0185] In practical applications, after the classification scores of the initial denoising models are determined, the first highest classification score, ie, the highest classification score, is determined from the classification scores, and the second highest classification score, ie, the second highest classification score, is determined from the classification scores.

[0186] Furthermore, the difference between the highest classification score and the second highest classification score is calculated, and then the difference is compared with the set threshold. If the difference is greater than the set threshold, it means that the classification is relatively accurate. At this time, the initial denoising model corresponding to the highest classification score can be determined as the target denoising model. If the difference is less than or equal to the set threshold, it means that the classification accuracy is not high, so the general denoising model among the initial denoising models is selected as the target denoising model.

[0187] In this way, the accuracy of classification can be further guaranteed, thereby avoiding the situation where the classification scores of each initial denoising model are relatively average and the differences are small, resulting in the adopted target denoising model being inaccurate, thereby improving the accuracy of the target denoising model and the accuracy of the generated image.

[0188] In one or more optional embodiments of the present invention, the target denoising model includes an input layer and a denoising layer;

[0189] The third coding feature is input into the target denoising model for denoising to obtain the fourth coding feature. The specific implementation process may be as follows:

[0190] Inputting the third encoding feature into the input layer so that the input layer randomly samples noise from a Gaussian distribution to obtain random noise;

[0191] Set the current denoising times to one;

[0192] Step A: inputting the third coding feature into the denoising layer, so that the denoising layer performs denoising according to the third coding feature, the current denoising times and the previous denoising result to obtain a current denoising result;

[0193] Step B: adding one to the current denoising times, and determining whether the current denoising times is greater than the set total denoising times;

[0194] If not, continue to execute step A and step B; if so, determine the current denoising result as the fourth coding feature;

[0195] When the current denoising times is one, the last denoising result is used as the random noise.

[0196] In practical applications, the target denoising model is the Unet model for illustration: the target image of the corresponding text is generated after the third encoding feature is denoised by the selected Unet denoising model (target denoising model): first, noise is randomly sampled from the Gaussian distribution, the shape of the tensor is (64, 64, 4), and the total number of denoising steps (the total number of denoising times) is set to M, where M is a positive integer. For each denoising, Unet takes the text encoding, the current number of steps, and the result of the previous denoising as input, predicts the noise to be removed, and calculates the feature distribution after denoising, repeats N times, and finally decodes the feature into an image through VAE.

[0197] In one or more optional embodiments of the present invention, the decoding process of the fourth coding feature to generate a target image corresponding to the description text includes:

[0198] Calculating feature distribution according to the fourth coding feature;

[0199] Performing sampling processing on the feature distribution to obtain sampling features;

[0200] The sampled features are decoded to obtain a target image corresponding to the description text.

[0201] In practical applications, on the basis of obtaining the fourth coding feature, the mean and variance of the fourth coding feature can be calculated, and the feature distribution of the fourth coding feature can be determined according to the calculated variance and mean. Further, the feature distribution is sampled to obtain the sampled features, i.e., the sampling features; the sampled features obtained by sampling are then sent to the decoder, and the sampling features in the latent space are mapped to the image data space, thereby generating the target image. In this way, by calculating the feature distribution, sampling and decoding, the decoding accuracy can be improved, thereby improving the accuracy of the target image.

[0202] In one or more optional embodiments of the present invention, the language mapping set includes a first language vocabulary and a second language vocabulary corresponding to at least one initial vocabulary;

[0203] The determining, based on the acquired description text and a preset language mapping set, a first language text corresponding to the description text includes:

[0204] Matching each second language vocabulary in the description text with the second language vocabulary corresponding to each of the initial vocabulary to determine the target vocabulary in each of the initial vocabulary;

[0205] The first language words corresponding to the target words are concatenated to obtain the first language text corresponding to the description text.

[0206] Specifically, the initial vocabulary is at least one of a style vocabulary and a designated vocabulary, and the first language vocabulary and the second language vocabulary corresponding to each vocabulary are associated to form a language mapping set, wherein the designated vocabulary refers to some vocabulary that cannot be recognized by the model for generating pictures based on the second language, that is, vocabulary whose semantics cannot be recognized, such as cyberpunk, graffiti, mosaic, etc.

[0207] In practical applications, on the basis of obtaining the description text and the language mapping set, the description text is further matched with the language mapping set: when the description text is a sentence, the description text is segmented to obtain at least one second language vocabulary, or when the description text is at least one word, the second language vocabulary in each word is determined; further, the second language vocabulary is matched with the second language vocabulary corresponding to each initial vocabulary in the language mapping set, and the initial vocabulary with the same semantics and / or vocabulary itself as the second language vocabulary in the description text is determined as the target vocabulary. Afterwards, the first language vocabulary corresponding to each target vocabulary in the language mapping set is spliced ​​to obtain the first language text corresponding to the description text.

[0208] In this way, by using the first language text as a supplement to the text information (descriptive text), the text comprehension ability can be improved, thereby improving the drawing ability.

[0209] It should be noted that in the case where the description text is a mixed-language text, such as when the description text contains text in the first language and text in the second language, the first-language vocabulary corresponding to the target vocabulary and the text in the first language included in the description text can be concatenated at this time to obtain the final text in the first language corresponding to the description text.

[0210] Exemplarily, the description text contains Chinese and English, such as "Anime Blue Sky Bus Cyberpunk", and the target vocabulary is Cyberpunk, and its English is "Cyberpunk". Then, "Bus" and "Cyberpunk" are concatenated to obtain the English text "Cyberpunk, Bus" corresponding to the description text.

[0211] It should be noted that when concatenating the first-language vocabulary, it can be concatenated using a comma ",", a separator "-", a space, or other symbols. The present invention does not make any limitation on this.

[0212] In one or more alternative embodiments of the present invention, the encoding the description text and the text in the first language respectively to obtain the first encoding feature corresponding to the description text and the second encoding feature corresponding to the text in the first language includes:

[0213] Inputting the description text into a second-language encoder for encoding to obtain the first encoding feature corresponding to the description text; and inputting the text in the first language into a first-language encoder for encoding to obtain the second encoding feature corresponding to the text in the first language.

[0214] Specifically, the second-language encoder refers to a text encoder that encodes text containing the second language. The first-language encoder refers to a text encoder that encodes text containing the first language.

[0215] In practical applications, in order to improve the encoding efficiency and accuracy, the description text is encoded using the second-language encoder to obtain the first encoding feature, and the text in the first language is encoded using the first-language encoder to obtain the second encoding feature. In this way, by encoding in parallel using two encoders, the efficiency of determining the first encoding feature and the second encoding feature can be improved, and on the basis of understanding the first language, a first-language encoder is added to supplement the understanding ability of the second-language encoder. Thus, the accuracy of the target image is improved.

[0216] In one or more alternative embodiments of the present invention, the inputting the description text into a second-language encoder for encoding to obtain the first encoding feature corresponding to the description text includes:

[0217] Performing vector conversion processing on each character unit in the description text to obtain a first encoding vector corresponding to the description text;

[0218] Performing position encoding processing on the first encoding vector to obtain a second encoding vector;

[0219] Determine, based on the first attention mechanism, a first hidden layer vector corresponding to the second encoding vector;

[0220] Performing residual connection and normalization processing on the second encoding vector and the first hidden layer vector to obtain a second hidden layer vector;

[0221] Based on the first activation function, determining a third hidden layer vector corresponding to the second hidden layer vector;

[0222] The second hidden layer vector and the third hidden layer vector are subjected to residual connection and normalization processing to obtain a first coding feature corresponding to the description text.

[0223] In practical applications, the second language encoder includes a first embedding layer, a first position encoding layer, a first attention mechanism, a first residual and normalization layer, a first feedforward network, and a second residual and normalization layer.

[0224] The description text can be input into the first embedding layer, and the first embedding layer converts each word (text unit) in the description text into a dimension as the first coding feature, where the first coding feature has three dimensions, namely, the number of description texts (here 1), the number of text units contained in the description text, and the embedding dimension of each text unit.

[0225] Next, the first coding feature is input into the first position coding layer, which marks the position of each character unit in the description text to obtain a first coding array, and then the first coding array is superimposed with the first coding feature to obtain a second coding vector.

[0226] Then, the second encoding vector is input into the first attention mechanism, and the first attention mechanism performs linear mapping on the second encoding vector to obtain Query (Q), Key (K) and Value (V). The attention matrix is ​​calculated according to Q and K, and then V is weighted according to the attention matrix to obtain the first hidden layer vector corresponding to the second encoding vector.

[0227] Afterwards, the second encoding vector and the first hidden layer vector are input into the first residual and normalization layer, and the first residual and normalization layer performs residual connection on the second encoding vector and the first hidden layer vector to obtain a first splicing vector; the first residual and normalization layer then performs normalization processing (standardization processing) on ​​the first splicing vector to obtain a second hidden layer vector.

[0228] Furthermore, the second hidden layer vector is input into the first feedforward network, the first feedforward network performs two-layer linear mapping on the second hidden layer vector, and processes it through the first activation function to obtain a third hidden layer vector corresponding to the second hidden layer vector.

[0229] Finally, the second hidden layer vector and the third hidden layer vector are residually connected by the second residual and standardization layer to obtain a second concatenated vector; the second concatenated vector is normalized (standardized) by the second residual and standardization layer to obtain the first encoding feature corresponding to the description text.

[0230] This is helpful to improve the accuracy of the first coding feature.

[0231] In one or more optional embodiments of the present invention, the step of inputting the first language text into a first language encoder for encoding to obtain a second encoding feature corresponding to the first language text includes:

[0232] Performing vector conversion processing on each character unit in the first language text to obtain a third encoding vector corresponding to the first language text;

[0233] Performing position encoding processing on the third encoding vector to obtain a fourth encoding vector;

[0234] Based on the second attention mechanism, determining a fourth hidden layer vector corresponding to the fourth encoding vector;

[0235] Performing residual connection and normalization processing on the fourth encoding vector and the fourth hidden layer vector to obtain a fifth hidden layer vector;

[0236] Based on the second activation function, determining a sixth hidden layer vector corresponding to the fifth hidden layer vector;

[0237] Residual connection and normalization are performed on the fifth hidden layer vector and the sixth hidden layer vector to obtain a second encoding feature corresponding to the first language text.

[0238] In practical applications, the first language encoder includes a second embedding layer, a second position encoding layer, a second attention mechanism, a third residual and normalization layer, a second feedforward network, and a fourth residual and normalization layer.

[0239] The first language text can be input into the second embedding layer, and the second embedding layer converts each character (text unit) in the first language text into a third encoding feature with a dimension, wherein the third encoding feature has three dimensions, namely, the number of first language texts, the number of text units contained in the first language text, and the embedding dimension of each text unit.

[0240] Next, the third coding feature is input into the second position coding layer, and the second position coding layer marks the position of each character unit in the first language text to obtain a second coding array, and then the second coding array is superimposed with the third coding feature to obtain a fourth coding vector.

[0241] Then, the fourth encoding vector is input into the second attention mechanism, and the second attention mechanism performs linear mapping on the fourth encoding vector to obtain Query (Q), Key (K) and Value (V). The attention matrix is ​​calculated according to Q and K, and then V is weighted according to the attention matrix to obtain the fourth hidden layer vector corresponding to the fourth encoding vector.

[0242] Afterwards, the fourth encoding vector and the fourth hidden layer vector are input into the third residual and normalization layer, and the third residual and normalization layer performs residual connection on the fourth encoding vector and the fourth hidden layer vector to obtain a third splicing vector; the third residual and normalization layer then performs normalization processing (standardization processing) on ​​the third splicing vector to obtain a fifth hidden layer vector.

[0243] Furthermore, the fifth hidden layer vector is input into the second feed-forward network, the second feed-forward network performs two-layer linear mapping on the fifth hidden layer vector, and processes it through the second activation function to obtain the sixth hidden layer vector.

[0244] Finally, the fifth hidden layer vector and the sixth hidden layer vector are subjected to the fourth residual and standardization layer, and the fifth hidden layer vector and the sixth hidden layer vector are residually connected by the fourth residual and standardization layer to obtain a fourth concatenated vector; the fourth concatenated vector is then normalized (standardized) by the fourth residual and standardization layer to obtain the second encoding feature corresponding to the first language text.

[0245] This is helpful to improve the accuracy of the second coding feature.

[0246] Combine the following Figure 2 The image generating method provided by the present invention is further described. Figure 2 This is the second flow chart of the image generation method provided by the present invention. The following is an example in which the first language is English and the second language is Chinese.

[0247] First, we collect some style words commonly used in text mapping and construct a Chinese-English mapping vocabulary (language mapping set). In the text mapping process, we first preprocess the input prompt word (Prompt), that is, the description text: copy the prompt word, convert the Chinese words (target words) that coexist in the copied prompt word and the English mapping dictionary (language mapping set) into English, and concatenate them with the existing English prompt words to obtain the English text corresponding to the prompt word, and use the English text as the input of the English text encoder (Text Encoder, first language encoder), that is, the English encoder, and the original prompt word remains unchanged as the input of the Chinese text encoder (second language encoder), that is, the Chinese encoder.

[0248] Figure 2 In the example, "car, cyberpunk" is used as the prompt words, and "cyberpunk" is used as the target word. The English corresponding to "cyberpunk" is "Cyberpunk", so the English text is also "Cyberpunk".

[0249] Then, the Chinese encoder encodes the prompt word to obtain a feature with a shape of (B, N, C), i.e., the first encoding feature, and the English encoder encodes the English text to obtain a feature with a shape of (B, N, C), i.e., the second encoding feature; further, the obtained features are concatenated in the sequence dimension to obtain a new feature (B, 2N, C), i.e., the third encoding feature (PromptEmbedding).

[0250] Next, the third encoding feature is input into the Unet (target denoising model) selected according to the prompt word and the first encoding feature. The third encoding feature is denoised by the selected Unet to generate the target image corresponding to the prompt word: first, noise is randomly sampled from the Gaussian distribution, and the shape of the tensor is (64, 64, 4). The total number of denoising steps (the total number of denoising times) is set to M. For each denoising, Unet uses the text encoding, the current number of steps, and the result of the previous denoising as input, predicts the noise to be removed, and calculates the feature distribution after denoising. This cycle is repeated M times, and finally the feature is decoded into an image through VAE.

[0251] Specifically, determine the target denoising model. Figure 3 The flowchart of determining the target denoising model provided by the present invention is shown, in which the target Unet (target denoising model) selected according to the prompt word and the first coding feature is: by retrieving keywords from the prompt word or performing text classification on the first coding feature, a Unet for a specific field is selected, such as a Unet for portraits, a Unet for animation, etc.

[0252] Specifically, input the prompt word "anime, car, cyberpunk", and calculate the classification scores corresponding to the input and set codes (codes corresponding to the style identifiers of the initial denoising model), respectively, where the classification identifiers include anime style, related to people, etc., and the keywords include keywords of anime style (such as anime, two-dimensional, etc.) and keywords of drawing people (such as male, female, sister, sister, etc.). If the difference between the highest classification score and the second highest score is more than 0.3, the classification is considered to be relatively accurate; at the same time, "anime" is found in the text in the keyword search; finally, the Unet to be selected is determined based on the keyword search results and the classification results. The keyword search results have a higher priority than the classification results. If no keywords are detected and the classification is valid, the classification results are used. If there are no valid results for both, the default Unet is used. That is, if the words related to the anime style are found in the prompt words, the anime model is selected, otherwise the default model (general denoising model) is used; if the highest value (highest classification score) and the second value (second highest classification score) differ by more than 0.3, the model corresponding to the highest score is selected, otherwise the default model (general denoising model) is used.

[0253] Unet refers to the network structure used for denoising in graph-generated text models.

[0254] In an embodiment of the present invention, the English text corresponding to the description text is determined based on the acquired description text and a preset language mapping set, and the description text contains Chinese text; the description text and the English text are respectively encoded to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the English text; the first encoding feature and the second encoding feature are processed for image generation to obtain a target image corresponding to the description text. On the basis of the Chinese image-to-text model, an English text encoder is added to supplement the understanding ability of the Chinese text encoder. In the Unet selector, a Unet in a specific field is selected to perform a denoising step to improve the drawing ability by keyword retrieval or text classification.

[0255] The image generating device provided by the present invention is described below. The image generating device described below and the image generating method described above can be referred to each other.

[0256] Figure 4 is a schematic diagram of the structure of the image generating device provided by the present invention, such as Figure 4 As shown, the image generating device 400 includes: a first determining module 401, an encoding module 402 and an image generating module 403, wherein:

[0257] A first determination module 401 is configured to determine a first language text corresponding to the description text according to the acquired description text and a preset language mapping set, wherein the description text contains a second language text;

[0258] The encoding module 402 is configured to encode the description text and the first language text respectively to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text;

[0259] The image generation module 403 is configured to perform image generation processing on the first coding feature and the second coding feature to obtain a target image corresponding to the description text.

[0260] The image generation device provided by the present invention determines the first language text corresponding to the description text according to the acquired description text and the preset language mapping set, wherein the description text contains the second language text; performs encoding processing on the description text and the first language text respectively to obtain the first encoding feature corresponding to the description text and the second encoding feature corresponding to the first language text; performs image generation processing on the first encoding feature and the second encoding feature to obtain the target image corresponding to the description text. By determining the second language text corresponding to the description text, some unrecognizable second language words are translated into the first language, and then the second encoding feature of the first language is obtained as a supplement to the description text, that is, as a supplement to the first encoding feature, thereby improving the ability to understand the description text, and then improving the matching degree between the target image and the description text, that is, improving the accuracy of the target image, which is conducive to improving user satisfaction.

[0261] According to one or more optional embodiments of the present invention, the image generation module 403 includes:

[0262] a concatenation module, configured to concatenate the first coding feature and the second coding feature to obtain a third coding feature;

[0263] a denoising module, configured to input the third coding feature into a target denoising model for denoising to obtain a fourth coding feature;

[0264] The decoding module is configured to decode the fourth coding feature to generate a target image corresponding to the description text.

[0265] According to one or more optional embodiments of the present invention, the image generating device 400 further includes:

[0266] The second determination module is configured to determine a target denoising model from a plurality of initial denoising models according to the description text and / or the first encoding feature.

[0267] According to one or more optional embodiments of the present invention, the second determining module is further configured to:

[0268] Perform keyword detection on the description text to obtain a detection result; and determine a target denoising model from multiple initial denoising models based on the detection result; or,

[0269] Determine the classification score of each of the initial denoising models according to the first coding feature; determine the target denoising model from the multiple initial denoising models according to each of the classification scores; or,

[0270] Perform keyword detection on the description text to obtain a detection result; if the detection result is that the keyword is detected, determine a target denoising model from multiple initial denoising models based on the keyword; if the detection result is that the keyword is not detected, determine the classification score of each of the initial denoising models based on the first coding feature, and determine the target denoising model from the multiple initial denoising models based on each of the classification scores.

[0271] According to one or more optional embodiments of the present invention, the second determining module is further configured to:

[0272] Obtaining a style identifier of each of the initial denoising models;

[0273] According to each of the style identifiers and the first coding feature, a classification score of each of the initial denoising models is determined.

[0274] According to one or more optional embodiments of the present invention, the second determining module is further configured to:

[0275] Calculating the matching degree between each of the style identifiers and the first coding feature;

[0276] Adding the matching degrees to obtain a total matching degree;

[0277] For each of the initial denoising models, a ratio of the matching degree corresponding to the initial denoising model to the total matching degree is determined as the classification score of the initial denoising model.

[0278] According to one or more optional embodiments of the present invention, the second determining module is further configured to:

[0279] determining a highest classification score and a second highest classification score among the classification scores;

[0280] When the difference between the highest classification score and the second highest classification score is greater than a set threshold, determining the initial denoising model corresponding to the highest classification score as the target denoising model;

[0281] When the difference between the highest classification score and the second highest classification score is less than or equal to a set threshold, the universal denoising model in each of the initial denoising models is determined as the target denoising model.

[0282] According to an image generating device provided by the present invention, the decoding module is further configured as follows:

[0283] Calculating feature distribution according to the fourth coding feature;

[0284] Performing sampling processing on the feature distribution to obtain sampling features;

[0285] The sampled features are decoded to obtain a target image corresponding to the description text.

[0286] According to one or more optional embodiments of the present invention, the language mapping set includes a first language vocabulary and a second language vocabulary corresponding to at least one initial vocabulary;

[0287] The first determining module 401 is further configured to:

[0288] Matching each second language vocabulary in the description text with the second language vocabulary corresponding to each of the initial vocabulary to determine the target vocabulary in each of the initial vocabulary;

[0289] The first language words corresponding to the target words are concatenated to obtain the first language text corresponding to the description text.

[0290] According to one or more optional embodiments of the present invention, the encoding module 402 is further configured to:

[0291] Performing vector conversion processing on each character unit in the description text to obtain a first coding vector corresponding to the description text; performing position coding processing on the first coding vector to obtain a second coding vector; determining a first hidden layer vector corresponding to the second coding vector based on a first attention mechanism; performing residual connection and normalization processing on the second coding vector and the first hidden layer vector to obtain a second hidden layer vector; determining a third hidden layer vector corresponding to the second hidden layer vector based on a first activation function; performing residual connection and normalization processing on the second hidden layer vector and the third hidden layer vector to obtain a first coding feature corresponding to the description text;

[0292] Perform vector conversion processing on each character unit in the first language text to obtain a third encoding vector corresponding to the first language text; perform position encoding processing on the third encoding vector to obtain a fourth encoding vector; based on a second attention mechanism, determine a fourth hidden layer vector corresponding to the fourth encoding vector; perform residual connection and normalization processing on the fourth encoding vector and the fourth hidden layer vector to obtain a fifth hidden layer vector; based on a second activation function, determine a sixth hidden layer vector corresponding to the fifth hidden layer vector; perform residual connection and normalization processing on the fifth hidden layer vector and the sixth hidden layer vector to obtain a second encoding feature corresponding to the first language text.

[0293] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the image generation method, which includes: determining the first language text corresponding to the description text according to the acquired description text and the preset language mapping set, wherein the description text contains the second language text; performing encoding processing on the description text and the first language text respectively to obtain the first encoding feature corresponding to the description text and the second encoding feature corresponding to the first language text; performing image generation processing on the first encoding feature and the second encoding feature to obtain the target image corresponding to the description text.

[0294] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0295] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image generation method provided by the above-mentioned methods, and the method includes: determining a first language text corresponding to the description text based on the acquired description text and a preset language mapping set, and the description text contains a second language text; encoding the description text and the first language text respectively to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text; performing image generation processing on the first encoding feature and the second encoding feature to obtain a target image corresponding to the description text.

[0296] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the image generation method provided by the above-mentioned methods, the method comprising: determining a first language text corresponding to the description text based on an acquired description text and a preset language mapping set, wherein the description text contains a second language text; encoding the description text and the first language text respectively to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text; performing image generation processing on the first encoding feature and the second encoding feature to obtain a target image corresponding to the description text.

[0297] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0298] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0299] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image generation method, characterized in that: include: Determining, according to the acquired description text and a preset language mapping set, a first language text corresponding to the description text, wherein the description text includes a second language text; Encoding the description text and the first language text respectively to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text; The first coding feature and the second coding feature are processed for image generation to obtain a target image corresponding to the description text.

2. The image generation method according to claim 1, characterized in that: The performing image generation processing on the first coding feature and the second coding feature to obtain a target image corresponding to the description text includes: Concatenate the first coding feature and the second coding feature to obtain a third coding feature; Inputting the third coding feature into the target denoising model for denoising to obtain a fourth coding feature; The fourth coding feature is decoded to generate a target image corresponding to the description text.

3. The image generation method according to claim 2, characterized in that: Before inputting the third coding feature into the target denoising model for denoising to obtain the fourth coding feature, the method further includes: A target denoising model is determined from a plurality of initial denoising models according to the description text and / or the first encoding feature.

4. The image generation method according to claim 3, characterized in that: Determining a target denoising model from a plurality of initial denoising models according to the description text and / or the first coding feature comprises: Perform keyword detection on the description text to obtain a detection result, and determine a target denoising model from multiple initial denoising models according to the detection result; or, Determine the classification score of each of the initial denoising models according to the first coding feature; determine the target denoising model from the multiple initial denoising models according to each of the classification scores; or, Perform keyword detection on the description text to obtain a detection result; if the detection result is that the keyword is detected, determine a target denoising model from multiple initial denoising models based on the keyword; if the detection result is that the keyword is not detected, determine the classification score of each of the initial denoising models based on the first coding feature, and determine the target denoising model from the multiple initial denoising models based on each of the classification scores.

5. The image generation method according to claim 4, characterized in that: Determining the classification score of each of the initial denoising models according to the first coding feature includes: Obtaining a style identifier of each of the initial denoising models; According to each of the style identifiers and the first coding feature, a classification score of each of the initial denoising models is determined.

6. The image generation method according to claim 5, characterized in that: Determining the classification score of each of the initial denoising models according to each of the style identifiers and the first coding feature includes: Calculating the matching degree between each of the style identifiers and the first coding feature; Adding the matching degrees to obtain a total matching degree; For each of the initial denoising models, a ratio of the matching degree corresponding to the initial denoising model to the total matching degree is determined as the classification score of the initial denoising model.

7. The image generation method according to any one of claims 4 to 6, characterized in that: Determining a target denoising model from a plurality of initial denoising models according to each of the classification scores comprises: determining a highest category score and a second highest category score among the category scores; When the difference between the highest classification score and the second highest classification score is greater than a set threshold, determining the initial denoising model corresponding to the highest classification score as the target denoising model; When the difference between the highest classification score and the second highest classification score is less than or equal to a set threshold, the universal denoising model in each of the initial denoising models is determined as the target denoising model.

8. The image generation method according to claim 2, characterized in that: The decoding process of the fourth coding feature to generate a target image corresponding to the description text includes: Calculating feature distribution according to the fourth coding feature; Performing sampling processing on the feature distribution to obtain sampling features; The sampled features are decoded to obtain a target image corresponding to the description text.

9. The image generation method according to claim 1, characterized in that: The language mapping set includes a first language vocabulary and a second language vocabulary corresponding to at least one initial vocabulary; The determining, based on the acquired description text and a preset language mapping set, a first language text corresponding to the description text includes: Matching each second language vocabulary in the description text with the second language vocabulary corresponding to each of the initial vocabulary to determine the target vocabulary in each of the initial vocabulary; The first language words corresponding to the target words are concatenated to obtain the first language text corresponding to the description text.

10. The image generation method according to claim 1, characterized in that: The encoding process is performed on the description text and the first language text respectively to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text, including: Performing vector conversion processing on each character unit in the description text to obtain a first coding vector corresponding to the description text; performing position coding processing on the first coding vector to obtain a second coding vector; determining a first hidden layer vector corresponding to the second coding vector based on a first attention mechanism; performing residual connection and normalization processing on the second coding vector and the first hidden layer vector to obtain a second hidden layer vector; determining a third hidden layer vector corresponding to the second hidden layer vector based on a first activation function; performing residual connection and normalization processing on the second hidden layer vector and the third hidden layer vector to obtain a first coding feature corresponding to the description text; Perform vector conversion processing on each character unit in the first language text to obtain a third encoding vector corresponding to the first language text; perform position encoding processing on the third encoding vector to obtain a fourth encoding vector; based on a second attention mechanism, determine a fourth hidden layer vector corresponding to the fourth encoding vector; perform residual connection and normalization processing on the fourth encoding vector and the fourth hidden layer vector to obtain a fifth hidden layer vector; based on a second activation function, determine a sixth hidden layer vector corresponding to the fifth hidden layer vector; perform residual connection and normalization processing on the fifth hidden layer vector and the sixth hidden layer vector to obtain a second encoding feature corresponding to the first language text.

11. An image generating device, characterized in that: include: A first determination module is configured to determine a first language text corresponding to the description text according to the acquired description text and a preset language mapping set, wherein the description text includes a second language text; an encoding module configured to encode the description text and the first language text respectively to obtain a first encoding feature corresponding to the description text and a second encoding feature corresponding to the first language text; The image generation module is configured to perform image generation processing on the first coding feature and the second coding feature to obtain a target image corresponding to the description text.