Image generation method and device, storage medium and electronic equipment
By acquiring the initial text and the target reference image, determining the intermediate features and target text of the reference image, and generating a mask image, the problems of training time and resource waste in the prior art are solved, and the image of the target character image is quickly generated.
Patent Information
- Application Number
- CN202510382575.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-25
AI Technical Summary
When generating images of specific character images, the prior art requires adding image and text sample pairs of corresponding character images to the training data set, resulting in a long time-consuming and waste of resources.
By obtaining the initial text and the target reference image, determining the intermediate features and target text of the reference image, calling the target text production diagram model to generate the text reference image and performing role segmentation, combining the character area mask image to generate the target generated image.
Without training the target text and graphics model, quickly generate images that follow the initial text and maintain the target character image, saving training time and model parameter storage consumption, and avoiding resource waste.
Smart Images

Figure CN120374764A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to an image generation method, apparatus, storage medium, and electronic device. Background Art
[0002] Currently, image generation technology has been widely applied in various fields, such as the generation of anime characters, etc. However, when generating a specific character image, related technologies need to add image and text sample pairs of the corresponding character image to the training dataset to train the text-to-image model, etc., resulting in a long training process and resource waste. Based on this, there is currently no good solution for how to conveniently generate a target generated image of a target character image. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide an image generation method, apparatus, storage medium, and electronic device to solve problems such as the long time-consuming of model training by using images of characteristic characters in related technologies; that is, embodiments of the present invention can generate a target generated image that follows the initial text and maintains the target character image indicated by the target reference image without training the existing target text-to-image model. By obtaining the initial text and the target reference image, to achieve the convenient generation of the target generated image of the target character image, that is, embodiments of the present invention can generate the target generated image containing the target character without retraining the target text-to-image model or increasing model parameters, etc., which can effectively save training time-consuming and storage consumption of model parameters, thereby effectively avoiding resource waste.
[0004] According to an aspect of an embodiment of the present invention, there is provided an image generation method, the method comprising:
[0005] Obtain an initial text and a target reference image;
[0006] Determine the intermediate features of the reference image corresponding to the target reference image, and determine the target text corresponding to the initial text;
[0007] Call the target text-to-image model to generate a text reference image of the target text; and perform role segmentation on the text reference image to obtain a generated role region mask image of the text reference image;
[0008] Call the target text-to-image model, and based on the intermediate features of the reference image and the generated role region mask image, generate a target generated image of the target text.
[0009] According to another aspect of an embodiment of the present invention, there is provided an image generation apparatus, the apparatus comprising:
[0010] An obtaining unit, configured to obtain an initial text and a target reference image;
[0011] A processing unit, configured to determine intermediate reference graph features corresponding to the target reference image and determine a target text corresponding to the initial text;
[0012] The processing unit is further configured to call a target text-to-image generation model to generate a text reference image of the target text; and perform role segmentation on the text reference image to obtain a generated role area mask image of the text reference image;
[0013] The processing unit is further configured to call the target text-to-image generation model to generate a target generated image of the target text based on the intermediate reference graph features and the generated role area mask image.
[0014] According to another aspect of the embodiments of the present invention, there is provided an electronic device, which includes a processor and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method mentioned above.
[0015] According to another aspect of the embodiments of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method mentioned above.
[0016] In the embodiments of the present invention, after obtaining the initial text and the target reference image, the intermediate reference graph features corresponding to the target reference image can be determined, and the target text corresponding to the initial text can be determined. Based on this, the target text-to-image generation model can be called to generate a text reference image of the target text; and role segmentation can be performed on the text reference image to obtain a generated role area mask image of the text reference image. Further, the target text-to-image generation model can be called to generate a target generated image of the target text based on the intermediate reference graph features and the generated role area mask image. It can be seen that in the embodiments of the present invention, without training the existing target text-to-image generation model, by obtaining the initial text and the target reference image, a target generated image that follows the initial text and maintains the target role image indicated by the target reference image can be generated, so as to conveniently generate a target generated image of the target role image; that is to say, in the embodiments of the present invention, without retraining the target text-to-image generation model and without increasing model parameters, etc., a target generated image including the target role can be generated, which can effectively save training time-consuming and model parameter storage consumption, thereby effectively avoiding resource waste. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In the following description of the exemplary embodiments with reference to the accompanying drawings, more details, features and advantages of the present invention are disclosed. In the drawings:
[0018] Figure 1Shows a schematic flowchart of an image generation method according to an exemplary embodiment of the present invention;
[0019] Figure 2 Shows a schematic diagram of a role segmentation according to an exemplary embodiment of the present invention;
[0020] Figure 3 Shows a schematic diagram of determining intermediate features of a layer to be determined according to an exemplary embodiment of the present invention;
[0021] Figure 4 Shows a schematic flowchart of another image generation method according to an exemplary embodiment of the present invention;
[0022] Figure 5 Shows a schematic flowchart of yet another image generation method according to an exemplary embodiment of the present invention;
[0023] Figure 6 Shows a schematic block diagram of an image generation apparatus according to an exemplary embodiment of the present invention;
[0024] Figure 7 Shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present invention. Detailed implementation manners
[0025] The embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0026] It should be understood that the various steps recited in the method embodiments of the present invention can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.
[0027] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0028] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0030] It should be noted that the execution subject of the image generation method provided in the embodiments of the present invention can be one or more electronic devices, and the embodiments of the present invention do not make any limitations in this regard; among them, the electronic device can be a terminal (i.e., a client) or a server. Then, when the execution subject includes multiple electronic devices, and at least one terminal and at least one server are included in the multiple electronic devices, the image generation method provided in the embodiments of the present invention can be jointly executed by the terminal and the server. Correspondingly, the terminal mentioned herein may include, but is not limited to: smart phones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, and so on. The server mentioned herein can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, and so on.
[0031] Based on the above description, the embodiments of the present invention propose an image generation method, which can be executed by the above-mentioned electronic device (terminal or server); or, this image generation method can be jointly executed by the terminal and the server. For the convenience of description, hereinafter, the case where the electronic device executes this image generation method will be taken as an example for illustration; as Figure 1 shown, this image generation method may include the following steps S101 - S104:
[0032] S101, obtain the initial text and the target reference image.
[0033] Optionally, the initial text can be any text (a text can also be called a text prompt, a text prompt word or text data, etc.), which is not limited in the embodiment of the present invention; wherein, the initial text can be a description of the target generated image (i.e., the image to be generated). Optionally, the target reference image can be any image, which is not limited in the embodiment of the present invention; wherein, the target reference image can be used to indicate the target character image, that is, the image of the character in the target reference image can be the target character image, and the character that conforms to the target character image can be called the target character, that is, the target reference image can be any image including the target character, and so on. Optionally, the target character can be any character, that is, the target character image can be any character image, which is not limited in the embodiment of the present invention; illustratively, the target character image can be any cartoon character image, in which case the target character can be any cartoon character, and so on.
[0034] In the embodiment of the present invention, the initial text and the target reference image may be obtained in the following ways, but not limited to:
[0035] The first acquisition method: Multiple texts and reference images corresponding to each of the multiple texts can be stored in the electronic device's own storage space. In this case, the electronic device can select an initial text (such as randomly selected or selected in sequence, etc.) and a reference image corresponding to the initial text from the multiple texts, and use the reference image corresponding to the initial text as the target reference image to obtain the initial text and the target reference image.
[0036] The second acquisition method: the electronic device can obtain a data download link, and use the text downloaded based on the data download link as the initial text, and use the image downloaded based on the data download link as the target reference image, so as to obtain the initial text and the target reference image.
[0037] The third acquisition method: the electronic device may have a data input interface, that is, a displayable data input interface. In this case, the user may perform a data input operation on the data input interface, and the electronic device may respond to the data input operation detected on the data input interface to use the text input by the data input operation as the initial text, and use the image input by the data input operation as the target reference image, and so on.
[0038] S102, determining the intermediate features of the reference image corresponding to the target reference image, and determining the target text corresponding to the initial text.
[0039] Optionally, when determining the intermediate features of the reference image corresponding to the target reference image, the electronic device may perform character segmentation on the target reference image to obtain a reference character region mask image of the target reference image, such as Figure 2As shown; that is, the electronic device can use image segmentation technology to extract the image area where the character is located from the target reference image to obtain a reference character area mask image. Optionally, the image segmentation technology involved in the embodiments of the present invention can be a region-based segmentation method or an edge-based segmentation method, etc. That is, the embodiments of the present invention can adopt any image segmentation method for character segmentation, and the embodiments of the present invention do not limit this.
[0040] Correspondingly, the electronic device can determine a character extraction image of the target reference image based on the reference character area mask image. The character extraction image can refer to a clean character image after removing the background. Specifically, the electronic device can superimpose the target reference image and the reference character area mask image to obtain the character extraction image of the target reference image. That is, the electronic device can use the pixel values of the pixel points located on the character area indicated by the reference character area mask image (i.e., the character area in the target reference image) in the target reference image (i.e., the pixel points with the first pixel value in the reference character area mask image) as the pixel values of the corresponding pixel points in the character extraction image, and the pixel values of the pixel points outside the character area indicated by the reference character area mask image (i.e., the pixel points with the second pixel value in the reference character area mask image) in the character extraction image can be specified pixel values (such as (0, 0, 0), etc.). Among them, the character area indicated by the reference character area mask image can be used to indicate the character area in the target reference image, that is, it can be used to indicate the area where the character is located in the target reference image. Optionally, both the first pixel value and the second pixel value can be set according to experience or according to actual needs, and the embodiments of the present invention do not limit this. Based on this, the embodiments of the present invention can use the first pixel value to indicate the pixel points located on the character area in the target reference image, and use the second pixel value to indicate the pixel points not located on the character area in the target reference image. Exemplarily, the first pixel value can be 255, and the second pixel value can be 0, etc. Optionally, the specified pixel value can be set according to experience or according to actual needs, and the embodiments of the present invention do not limit this.
[0041] Further, the electronic device can perform image encoding on the role extraction image to obtain the role extraction image features of the role extraction image; optionally, the target text-to-image model can include, but is not limited to, at least one of the following: Text Encoder (text encoding) module, Image Encoder (image encoding) module, Denoiser (noise reduction) module, Image Decoder (image decoding) module, etc.; the embodiments of the present invention do not limit this; based on this, the electronic device can call the image encoding module in the target text-to-image model to perform image encoding on the role extraction image to obtain the role extraction image features of the role extraction image. Optionally, the target text-to-image model can be any pre-trained text-to-image model, and the embodiments of the present invention do not limit the specific training process of the target text-to-image model. For example, the target text-to-image model can be obtained by training a text set and the training role images corresponding to each training text in the training text set, etc. Optionally, the target text-to-image model can be any text-to-image diffusion model; optionally, any module in the target text-to-image model can be a network structure, and the embodiments of the present invention do not limit the specific network structure of any module in the target text-to-image model.
[0042] Based on this, the electronic device can determine a noise-added reference sequence, and based on the noise-added reference sequence and the role extraction image features, determine the reference intermediate features at each specified noise reduction step among the N noise reduction steps to achieve determining the reference intermediate features of the target reference image, where N is a positive integer; among them, the reference intermediate features of the reference image can include the reference intermediate features at each specified noise reduction step, and the reference intermediate features at one noise reduction step can include the reference layer intermediate features of each self-attention layer in the noise reduction module of the target text-to-image model at the corresponding noise reduction step (that is, the reference intermediate features at one specified noise reduction step can include the reference layer intermediate features of each self-attention layer in the noise reduction module at the corresponding specified noise reduction step), that is to say, the reference intermediate features of the reference image can include the reference layer intermediate features of each self-attention layer at each specified noise reduction step. Optionally, one layer intermediate feature can include an intermediate key vector (i.e., a K (Key) vector, also referred to as a K feature) and an intermediate value vector (i.e., a V (Value) vector, also referred to as a V feature). Optionally, the number of reference layer intermediate features of one self-attention layer at one specified noise reduction step can be one or more, that is, the number of reference intermediate key vectors and reference intermediate value vectors of one self-attention layer at one specified noise reduction step can be one or more, and the embodiments of the present invention do not limit this.
[0043] Optionally, the N denoising steps (which may also be referred to as time steps) may include at least one specified denoising step. That is to say, each specified denoising step among the N denoising steps may be each specified denoising step among at least one specified denoising step included in the N denoising steps. Optionally, the at least one specified denoising step may include the N denoising steps. That is to say, each denoising step among the N denoising steps may be taken as a specified denoising step; or, the at least one specified denoising step may include the denoising steps between the P-th denoising step and the Q-th denoising step among the N denoising steps (i.e., the denoising steps between [P, Q]), where P and Q may be positive integers less than or equal to N, and P is less than Q. That is to say, each denoising step between the P-th denoising step and the Q-th denoising step may be taken as a specified denoising step; or, the denoising steps at odd positions (such as the 1st denoising step, etc.) may be taken as a specified denoising step; or, the denoising steps at even positions (such as the 2nd denoising step, etc.) may be taken as a specified denoising step, and so on; the embodiments of the present invention do not limit this.
[0044] Optionally, the noise addition reference sequence may include N noise values, and the first noise value among the N noise values is greater than the second noise value. The first noise value and the second noise value may be any noise values among the N noise values, and the first noise value is before the second noise value. Optionally, the cumulative multiplication result among the N noise values may be close to 0; and / or, the first noise value in the noise addition reference sequence may be close to 1, and so on. Optionally, the noise addition reference sequence may be set according to experience or according to actual requirements. The embodiments of the present invention do not limit this; exemplarily, the noise addition reference sequence may be [0.9991, 0.9951, 0.9910, 0.9870, 0.9829, 0.9788, 0.9748, 0.9707, 0.9666, 0.9626, 0.9585, 0.9544, 0.9504, 0.9463, 0.9423, 0.9382, 0.9341, 0.9301, 0.9260, 0.9219, 0.9179, 0.9138, 0.9097, 0.9057, 0.9016, 0.8975, 0.8935, 0.8894, 0.8853, 0.8813, 0.8772, 0.8732, 0.8691, 0.8650, 0.8610, 0.8569, 0.8528, 0.8488, 0.8447, 0.8406, 0.8366, 0.8325, 0.8284, 0.8244, 0.8203, 0.8163, 0.8122, 0.8081, 0.8041, 0.8000], and so on.
[0045] Optionally, when determining the reference intermediate feature of each specified noise reduction step among the N noise reduction steps based on the noisy reference sequence and the role to extract image features, for the nth noise reduction step among the N noise reduction steps, if the nth noise reduction step is a specified noise reduction step, the electronic device can determine the first n noise values from the noisy reference sequence, where n ∈ [1, N]; and determine the noisy image feature at the nth noise reduction step based on the first n noise values and the role to extract image features; further, the reference intermediate feature of each self-attention layer at the nth noise reduction step can be determined based on the noisy image feature at the nth noise reduction step, so as to determine the reference intermediate feature of the noise reduction step at the nth noise reduction step, that is, to determine the reference intermediate feature of a specified noise reduction step, and thus to determine the reference intermediate feature of each specified noise reduction step. In other words, for any specified noise reduction step among at least one specified noise reduction step, the electronic device can determine the reference intermediate feature of the noise reduction step at any specified noise reduction step based on the noisy reference sequence and the role to extract image features. Optionally, when the nth noise reduction step is one of the at least one specified noise reduction steps, it can be determined that the nth noise reduction step is a specified noise reduction step; when the nth noise reduction step is not one of the at least one specified noise reduction steps, it can be determined that the nth noise reduction step is not a specified noise reduction step; by way of example, assuming that at least one specified noise reduction step includes the noise reduction steps within [P, Q] among the N noise reduction steps, then when n ∈ [P, Q], it can be determined that the nth noise reduction step is one of the at least one specified noise reduction steps, and so on.
[0046] Optionally, when determining the noisy image feature at the nth noise reduction step based on the first n noise values and the role to extract image features, the electronic device can use Equation 1.1 to determine the noisy image feature at the nth noise reduction step:
[0047]
[0048] where x n can represent the noisy image feature at the nth noise reduction step, and ε can represent the random noise sampled from the standard normal distribution to implement the process of adding Gaussian noise; optionally, can be the product of the first n noise values among the noisy reference sequence α1, α2, …, α N . Optionally, the random noise at different noise reduction steps can be the same or different, and the embodiments of the present invention do not make any limitation thereto.
[0049] Optionally, when determining the intermediate reference layer features of each self-attention layer at the nth noise reduction step based on the noise-added image features at the nth noise reduction step, the electronic device may input the noise-added image features at the nth noise reduction step into the noise reduction module in the target text-to-image model (i.e., input into the Denoiser), so as to determine the intermediate reference layer features of each self-attention layer at the nth noise reduction step. Optionally, the noise reduction module may include, but is not limited to, at least one of the following: at least one convolutional layer and a transformer layer (an encoder-decoder architecture). Optionally, the transformer layer may include at least one self-attention layer (i.e., a Self-attention layer) and at least one cross-attention layer (i.e., a cross-attention layer), that is to say, the noise reduction module may include at least one self-attention layer and at least one cross-attention layer. Based on this, the electronic device may input the noise-added image features at the nth noise reduction step into the noise reduction module to obtain the intermediate features of the undetermined layer of each self-attention layer in the noise reduction module at the nth noise reduction step, and may determine the intermediate reference layer features of each self-attention layer at the nth noise reduction step based on the intermediate features of the undetermined layer of each self-attention layer at the nth noise reduction step. An intermediate feature of the undetermined layer may include an intermediate key vector and an intermediate value vector. Then, the intermediate features of the undetermined layer of an attention layer at a noise reduction step may include the key vector and the value vector of each pixel point of the corresponding attention layer at the corresponding noise reduction step. Each pixel point at this time may refer to each pixel point in an image feature. Optionally, the electronic device may only input the noise-added image features at the nth noise reduction step into the noise reduction module. At this time, the Q (Query) vector in a cross-attention layer may perform cross-attention with an empty string; or, the electronic device may also input the text encoding features of the initial text or the text encoding features of the target text into the noise reduction module to call the noise reduction module to determine the intermediate features of the undetermined layer of each self-attention layer at the nth noise reduction step based on the text encoding features of the initial text or the target text, and the noise-added image features at the nth noise reduction step, and so on; the embodiments of the present invention do not limit this.
[0050] It can be seen that the noisy image features at the nth denoising step will be input into the Denoiser, and the intermediate features of the Self-attention layer of the Denoiser will be saved; that is to say, the Self-attenion layer contains three linear layers, which extract the Q, K, and V vectors from the input image features to calculate Attention (attention), so as to use the K and V vectors (also called K and V features) representing the key and value under each self-attention layer as the pending layer intermediate features of each self-attention layer at the nth denoising step. That is, the pending layer intermediate features of a self-attention layer at the nth denoising step can include the K vector and the V vector of the corresponding self-attention at the nth denoising step, so as to record the pending layer intermediate features of each self-attention layer at the nth denoising step, as Figure 3 shown. Optionally, an image feature may include pixel feature vectors of each pixel point among multiple pixel points, and a pixel feature vector can determine a K vector and a V vector (that is, a pixel point can correspond to a K vector and a V vector under a self-attention layer), etc. Then, the pending layer intermediate features of a self-attention layer at the nth denoising step can include the K vectors and V vectors corresponding to each pixel point of the corresponding self-attention at the nth denoising step, that is, it can include the pending layer intermediate features corresponding to each pixel point of the corresponding self-attention at the nth denoising step, and a pending layer intermediate feature can include a K vector and a V vector. Optionally, before an image feature is processed by an attention layer, the corresponding image feature can be represented as a matrix composed of pixel feature vectors of each pixel point in the corresponding image feature, that is, a row of data in this matrix can represent the pixel feature vector of a pixel point. That is to say, an image feature can be represented in the form of a matrix; it should be noted that the specific implementation of the representation method of the image feature in the embodiment of the present invention is not limited. That is to say, it can be represented in a row-first manner (at this time, the matrix used to represent the image feature can successively include the pixel feature vectors of the pixel points in the first row and first column, the first row and second column, the first row and third column, …, the pixel feature vectors of the pixel points in the second row and first column, the second row and second column, …) or in a column-first manner (at this time, the matrix used to represent the image feature can successively include the pixel feature vectors of the pixel points in the first row and first column, the second row and first column, the third row and first column, …, the pixel feature vectors of the pixel points in the first row and second column, the second row and second column, …), and so on.
[0051] Optionally, when determining the reference layer intermediate features of each self-attention layer at the nth noise reduction step based on the pending layer intermediate features of each self-attention layer at the nth noise reduction step, for any self-attention layer in the noise reduction module (i.e., any self-attention layer in at least one self-attention layer), the electronic device can determine the reference layer intermediate features of any self-attention layer at the nth noise reduction step based on the pending layer intermediate features of any self-attention layer at the nth noise reduction step and the reference role area mask image. Based on this, the electronic device can limit the range of the reference layer intermediate features through the reference role area mask image. At this time, only the K and V vectors corresponding to the positions shown as the role foreground in the reference role area mask image (i.e., the pixel points located on the role area indicated by the reference role area mask image) will be saved to obtain the reference layer intermediate features of any self-attention layer at the nth noise reduction step. That is to say, the pending layer intermediate features corresponding to the pixel points located on the role area indicated by the reference role area mask image of any self-attention layer at the nth noise reduction step can be used as the reference layer intermediate features corresponding to the corresponding pixel points of any self-attention layer at the nth noise reduction step to obtain the reference layer intermediate features corresponding to the pixel points located on the role area indicated by the reference role area mask image of any self-attention layer at the nth noise reduction step. The reference layer intermediate features of any self-attention layer at the nth noise reduction step can include the reference layer intermediate features corresponding to the pixel points located on the role area indicated by the reference role area mask image of any self-attention layer at the nth noise reduction step. Optionally, if the feature size of the self-attention layer in the internal input of the Denoiser is different from the size of the reference role area mask image (that is, the size of the image features input to a self-attention layer is different from the size of the reference role area mask image), the size of the reference role area mask image can be scaled to align with the image feature size of the input self-attention layer; it should be noted that the embodiments of the present invention do not limit the scaling method of the reference role area mask image. For example, the pixel values corresponding to each pixel point in the corresponding image features in the reference role area mask image can be determined by any interpolation method, etc.; based on this, the reference role area mask image for limiting the range can be the reference role area mask image aligned with the corresponding image features.Based on this, when determining the intermediate feature of the reference layer of any self-attention layer at the nth noise reduction step based on the to-be-determined intermediate feature of any self-attention layer at the nth noise reduction step and the reference role area mask image, if the size of the input image feature of any self-attention layer at the nth noise reduction step is different from the size of the reference role area mask image, the reference role area mask image can be scaled according to the size of the input image feature of any self-attention layer at the nth noise reduction step to obtain the current reference role area mask image; if the size of the input image feature of any self-attention layer at the nth noise reduction step is the same as the size of the reference role area mask image, the reference role area mask image can be used as the current reference role area mask image, that is, the reference role area mask image does not need to be scaled; thus, the intermediate feature of the reference layer of any self-attention layer at the nth noise reduction step can be determined based on the to-be-determined intermediate feature of any self-attention layer at the nth noise reduction step and the current reference role area mask image.
[0052] Based on this, when determining the intermediate feature of the reference layer of any self-attention layer at the nth noise reduction step based on the intermediate feature of the undetermined layer and the current reference role area mask image at the nth noise reduction step of any self-attention layer, for the intermediate feature of the undetermined layer corresponding to any pixel point of any self-attention layer at the nth noise reduction step, the electronic device can determine whether any pixel point of any self-attention layer at the nth noise reduction step is located on the role area indicated by the current reference role area mask image (such as whether the pixel value of the pixel point in the current reference role area mask image is the first pixel value). If any pixel point is located on the role area indicated by the current reference role area mask image, the intermediate feature of the undetermined layer corresponding to any pixel point of any self-attention layer at the nth noise reduction step can be used as an intermediate feature of the reference layer of any self-attention layer at the nth noise reduction step, that is, as the intermediate feature of the reference layer corresponding to any pixel point of any self-attention layer at the nth noise reduction step. If any pixel point is not located on the role area indicated by the current reference role area mask image, the intermediate feature of the undetermined layer corresponding to any pixel point of any self-attention layer at the nth noise reduction step may not be used as an intermediate feature of the reference layer of any self-attention layer at the nth noise reduction step. In other words, the electronic device can use the intermediate features of the undetermined layer corresponding to each reference pixel point among at least one reference pixel point of any self-attention layer at the nth noise reduction step as an intermediate feature of the reference layer of any self-attention layer at the nth noise reduction step (that is, as the intermediate feature of the reference layer corresponding to the corresponding reference pixel point of any self-attention layer at the nth noise reduction step) to save the intermediate features of the reference layer corresponding to each reference pixel point of any self-attention layer at the nth noise reduction step. Among them, at least one reference pixel point may include each pixel point located on the role area indicated by the current reference role area mask image among all pixel points of any self-attention layer at the nth noise reduction step. It can be seen that the intermediate feature of the reference layer of any self-attention layer at the nth noise reduction step may include the intermediate features of the reference layer corresponding to each reference pixel point of any self-attention layer at the nth noise reduction step, that is, may include the intermediate key vectors (i.e., K vectors) and intermediate value vectors (i.e., V vectors) corresponding to each reference pixel point of any self-attention layer at the nth noise reduction step.
[0053] Furthermore, when determining the target text corresponding to the initial text, the electronic device may call a character feature label extraction model to perform label extraction on the character extraction image to obtain at least one text label corresponding to the target reference image, wherein the at least one text label includes a text label of the character (i.e., the character image) in the target reference image; thereby combining (i.e., splicing) the initial text and at least one text label to obtain the target text corresponding to the initial text. Among them, the role feature label extraction model belongs to a type of image attribute recognition model, and can obtain a series of text labels describing the role in the image; optionally, the role feature label extraction model may include at least one role feature label recognition model (such as a classification model, etc.), that is, it may include at least one pre-trained role feature label recognition model, that is, at least one role feature label recognition model may be obtained by model training based on the attribute recognition training image set and at least one training text label corresponding to each attribute recognition training image in the attribute recognition training image set, and a role feature label recognition model may be used to predict one or more text labels corresponding to an image; illustratively, when an attribute recognition training image is an animation character image (that is, the character in an attribute recognition training image is an animation character), the role feature label extraction model may be a dedicated attribute recognition for animation character images, at which time the character in the target reference image may be an animation character, and the initial text may be a text used to describe the animation character, etc. It should be noted that the embodiment of the present invention does not limit the specific model structure and model training process of at least one role feature label recognition model. Optionally, at least one text tag may include but is not limited to at least one of the following: hair color, hair length, type of head ornaments, eye color, eye size, clothing color, and shoe style, etc.; this is not limited in this embodiment of the present invention.
[0054] Based on this, the target text may include the initial text and at least one text tag, so that the initial text and the at least one text tag may be combined into a longer text prompt word.
[0055] S103, calling the target text-generating graph model to generate a text reference image of the target text; and performing character segmentation on the text reference image to obtain a generated character region mask image of the text reference image.
[0056] In an embodiment of the present invention, an electronic device may input a target text into a target text-to-image model to generate a text reference image of the target text through the target text-to-image model; that is to say, at this time, image generation guided only by text (i.e., only inputting the target text to guide the generation of the text reference image) can be performed through the target text-to-image model. The image generation process guided only by text can use the existing target text-to-image model, select a random seed (an integer, which can be used to generate Gaussian noise), and then use the denoiser to perform iterative noise reduction, and finally obtain the text reference image through the image decoding module; among them, the seed determines all the randomness involved in the model when generating pictures. Optionally, the random seed when generating the target generated image described below can be the same as the random seed when generating the text reference image, that is, the Gaussian noise when generating the target generated image can be the same as the Gaussian noise when generating the text reference image, and so on.
[0057] It should be noted that the specific implementation manner of performing role segmentation on the text reference image to obtain the generated role region mask image of the text reference image can be the same as the specific implementation manner of performing role segmentation on the target reference image to obtain the reference role region mask image of the target reference image, and the embodiments of the present invention will not be elaborated herein.
[0058] S104, call the target text-to-image model, and generate a target generated image of the target text based on the intermediate features of the reference image and the generated role region mask image.
[0059] In an embodiment of the present invention, an electronic device may input a target text into a target text-to-image model to output a target generated image through the target text-to-image model based on the intermediate features of the reference image and the generated role region mask image. Among them, the generated role region mask image can be used to constrain the spatial range of the injection of the intermediate features of the reference image.
[0060] In an embodiment of the present invention, after obtaining the initial text and the target reference image, the intermediate features of the reference image corresponding to the target reference image are determined, and the target text corresponding to the initial text is determined. Based on this, the target text-to-image generation model can be called to generate a text reference image of the target text; and the text reference image is subjected to role segmentation to obtain a generated role region mask image of the text reference image. Further, the target text-to-image generation model can be called to generate a target generated image of the target text based on the intermediate features of the reference image and the generated role region mask image. It can be seen that in the embodiment of the present invention, without training the existing target text-to-image generation model, the target generated image that follows the initial text and maintains the target role image indicated by the target reference image can be generated through the obtained initial text and target reference image, so as to conveniently generate the target generated image of the target role image; that is to say, in the embodiment of the present invention, without retraining the target text-to-image generation model and without increasing model parameters, etc., the target generated image including the target role can be generated, which can effectively save training time and model parameter storage consumption, thereby effectively avoiding resource waste.
[0061] Based on the above description, another image generation method is also proposed in an embodiment of the present invention. Correspondingly, this image generation method can be executed by the above-mentioned electronic device (terminal or server); or, this image generation method can be jointly executed by the terminal and the server. For the convenience of description, in the following, it will be described by taking the electronic device executing this image generation method as an example; please refer to Figure 4 , this image generation method may include the following steps S401-S407:
[0062] S401, obtain the initial text and the target reference image.
[0063] S402, determine the intermediate features of the reference image corresponding to the target reference image, and determine the target text corresponding to the initial text.
[0064] S403, call the target text-to-image generation model to generate a text reference image of the target text; and perform role segmentation on the text reference image to obtain a generated role region mask image of the text reference image.
[0065] Among them, the above-mentioned intermediate features of the reference image include the reference intermediate features of each specified denoising step among N denoising steps, and the reference intermediate features of one denoising step include the reference layer intermediate features of each self-attention layer in the denoising module of the target text-to-image generation model at the corresponding denoising step.
[0066] S404, call the text encoding module in the target text-to-image generation model to perform text encoding on the target text to obtain the text encoding features of the target text.
[0067] S405. Traverse each of the N noise reduction steps in order from the last to the first, and use the currently traversed noise reduction step as the current noise reduction step.
[0068] In an embodiment of the present invention, the electronic device may traverse each of the N noise reduction steps in the order of the Nth noise reduction step, the (N - 1)th noise reduction step, the (N - 2)th noise reduction step,..., the 1st noise reduction step, so as to traverse each of the N noise reduction steps in order from the last to the first.
[0069] S406. Determine the generated image features at the current noise reduction step; if the current noise reduction step is a specified noise reduction step, call the noise reduction module in the target text-to-image model, and based on the generated image features at the current noise reduction step, the noise reduction step reference intermediate features at the current noise reduction step, the text encoding features, and the generated character region mask image, determine the generated image features at the previous noise reduction step of the current noise reduction step; if the current noise reduction step is not a specified noise reduction step, call the noise reduction module in the target text-to-image model, and based on the generated image features at the current noise reduction step and the text encoding features, determine the generated image features at the previous noise reduction step of the current noise reduction step.
[0070] In an embodiment of the present invention, the noise reduction module in the target text-to-image generation model includes at least one self-attention layer and at least one cross-attention layer. Based on this, when calling the noise reduction module in the target text-to-image generation model to determine the generated image feature at the previous noise reduction step of the current noise reduction step based on the generated image feature at the current noise reduction step, the noise reduction step reference intermediate feature at the current noise reduction step, the text encoding feature (herein referring to the text encoding feature of the target text), and the generated character region mask image, the electronic device can input the generated image feature and the text encoding feature at the current noise reduction step into the noise reduction module in the target text-to-image generation model to sequentially call each attention layer in the noise reduction module and use the currently called attention layer as the current attention layer; if the current attention layer is a self-attention layer, the original layer intermediate feature of the current attention layer at the current noise reduction step can be determined; and based on the reference layer intermediate feature of the current attention layer at the current noise reduction step, the original layer intermediate feature of the current attention layer at the current noise reduction step, and the generated character region mask image, the output image feature of the current attention layer at the current noise reduction step can be determined. Optionally, the input data of an attention layer at the current noise reduction step may include, but is not limited to, at least one of the following: the output image feature of the previous attention layer of the corresponding attention layer at the current noise reduction step, the text encoding feature (i.e., the text encoding feature of the target text), and the generated image feature at the current noise reduction step, etc., and the embodiments of the present invention do not limit this; that is, the input data of a self-attention layer at the current noise reduction step may include the generated image feature at the current noise reduction step, or may include the output image feature of the previous attention layer of the corresponding self-attention layer at the current noise reduction step, or may include the output image feature of the previous network of the corresponding self-attention layer at the current noise reduction step, etc.; the embodiments of the present invention do not limit this. Among them, an image feature can be represented in the form of a matrix (i.e., one row of data can represent the pixel feature vector of a pixel point in an image feature).
[0071] Optionally, the original layer intermediate feature of the current attention layer at the current noise reduction step can be determined based on the input data of the current attention layer at the current noise reduction step. That is, the input data of the current attention layer at the current noise reduction step can be input into the current attention layer to determine the original layer intermediate feature of the current attention layer at the current noise reduction step through the current attention layer.
[0072] Optionally, an intermediate layer feature includes an intermediate key vector and an intermediate value vector. Based on this, when determining the output image feature of the current attention layer at the current noise reduction step based on the reference layer intermediate feature of the current attention layer at the current noise reduction step, the original layer intermediate feature of the current attention layer at the current noise reduction step, and the generated character region mask image, the electronic device can concatenate the reference intermediate key vector of the current attention layer at the current noise reduction step and the original intermediate key vector of the current attention layer at the current noise reduction step to obtain the key vector concatenation result of the current attention layer at the current noise reduction step; and concatenate the reference intermediate value vector of the current attention layer at the current noise reduction step and the original intermediate value vector of the current attention layer at the current noise reduction step to obtain the value vector concatenation result of the current attention layer at the current noise reduction step. Among them, the concatenation of K and V occurs in the Token (representation) dimension, and each token represents each point position in the image feature, that is, it can represent the K vector or V vector corresponding to a pixel point in the image feature, etc. Exemplarily, the embodiment of the present invention may adopt K original to represent the original intermediate key vector of the current attention layer at the current noise reduction step (i.e., the original K vector, that is, the K vector obtained by inputting the input data into the current attention layer), and adopt K reference to represent the reference intermediate key vector of the current attention layer at the current noise reduction step. Then, the key vector concatenation result of the current attention layer at the current noise reduction step may be K concat =[K original , K reference , and so on.
[0073] Furthermore, the electronic device can determine the attention mask based on the generated character region mask image. That is to say, the embodiment of the present invention can use the attention mask mechanism to perform spatial range constraint on the additionally concatenated K vector and V vector; optionally, the electronic device can determine the attention mask according to Formula 2.1:
[0074]
[0075] wherein, M can represent the attention mask, and T Ko can represent the number of K original (that is, the number of K vectors in K original ), and F can represent the pixel points on the character region indicated by the generated character region mask image (that is, it can represent the character foreground point positions indicated by the generated character region mask image). Assume that the number of tokens of Q is T Q , and the number of K reference is T Kr , then the size of M can be T Q ×(T Ko +T Kr ); that is to say, the first TKo The columns can all be 0, representing Q and K original The attention scores (i.e., attention scores) of are not affected (here it means that the interaction with the original K and V vectors is not affected), and for the last T Kr columns of M, if the position of Q (i.e., the position of any pixel in Q) is on the role area indicated by the generated role area mask image (i.e., on the role foreground position indicated by the generated role area mask image, i.e., i ∈ F), it can be 0, that is, the attention scores are not affected (i.e., the pixels on the role area indicated by the generated role area mask image will interact with the additionally injected K and V (i.e., the reference K vector and the reference V vector)), otherwise it is set to negative infinity. Based on this, the embodiments of the present invention can only make the pixels on the role foreground position indicated by the generated role area mask image interact with the additionally injected K and V to absorb additional role features, reduce the interference to the background part of the generation result, and enable the background to be generated following the text prompt; that is to say, after injecting the intermediate features of the reference image, the role area (i.e., the role foreground area) indicated by the generated role area mask image has the opportunity to absorb the intermediate features of the reference image from the target reference image, that is, the role foreground area indicated by the text reference image has the opportunity to absorb the intermediate features of the reference image from the target reference image, which is of great benefit to maintaining the graphic consistency of the background part.
[0076] Optionally, if the size of Q is different from the size of the generated role area mask map, the size of the generated role area mask map can be scaled to align with the size of Q; it should be noted that the embodiments of the present invention do not limit the specific implementation manner of scaling the generated role area mask map. For example, the pixel values corresponding to each pixel point of the Q vector in the generated role area mask map can be determined by any interpolation method, etc.; then correspondingly, the role area indicated by the generated role area mask map can be the role area indicated by the scaled generated role area mask map, etc.
[0077] Based on this, the electronic device can determine the output image features of the current attention layer at the current noise reduction step based on the key vector splicing result, the value vector splicing result, and the attention mask. Optionally, the electronic device can use Equation 2.2 to determine the output image features of the current attention layer at the current noise reduction step (which can also be expressed as Attention, i.e., attention):
[0078]
[0079] Among them, V concatIt can represent the result of value vector concatenation. Q can represent the Q vector of the current attention layer at the current noise reduction step (i.e., it can include the Q vectors corresponding to each pixel point of the current attention layer at the current noise reduction step). The transpose multiplication of Q and K can obtain the attention scores.
[0080] It should be understood that if the current attention layer is a cross-attention layer, the current attention layer can be called to determine the output image features of the current attention layer at the current noise reduction step. That is to say, based on the text encoding features, the output image features of the current attention layer at the current noise reduction step can be determined. That is, at this time, the input data of the current attention layer at the current noise reduction step can also include text encoding features.
[0081] Then correspondingly, when calling the noise reduction module in the target text-to-image generation model to determine the generated image features at the previous noise reduction step of the current noise reduction step based on the generated image features and text encoding features at the current noise reduction step, the electronic device can directly input the generated image features and text encoding features at the current noise reduction step into the noise reduction module, so as to output the generated image features at the previous noise reduction step of the current noise reduction step through the noise reduction module. That is to say, at this time, the K and V for calculating the attention can be the original intermediate key vector of the current attention layer at the current noise reduction step and the original intermediate value vector of the current attention layer at the current noise reduction step, so as not to splice the original layer intermediate features, and so on.
[0082] Based on this, after calling each attention layer in the noise reduction module, the electronic device can determine the generated image features at the previous noise reduction step of the current noise reduction step based on the output image features of the last attention layer in the noise reduction module at the current noise reduction step. Optionally, the electronic device can use the output image features of the last attention layer at the current noise reduction step as the generated image features at the previous noise reduction step of the current noise reduction step. Or, after the last attention layer, the noise reduction module further includes at least one network layer (such as a convolutional layer, etc.). At this time, the generated image features at the previous noise reduction step of the current noise reduction step can also be determined through the at least one network layer based on the output image features of the last attention layer at the current noise reduction step, and so on. The embodiments of the present invention do not limit this.
[0083] It should be noted that when at least one specified denoising step includes N denoising steps (i.e., each of the N denoising steps is a specified denoising step), the electronic device may not determine whether the current denoising step is a specified denoising step, but directly trigger the execution of the above-mentioned denoising module in the target text-to-image generation model, and determine the generated image feature at the previous denoising step of the current denoising step based on the generated image feature at the current denoising step, the denoising step reference intermediate feature at the current denoising step, the text encoding feature, and the generated character region mask image; correspondingly, at this time, the above-mentioned determination of the first n noise values from the noise addition reference sequence can be directly triggered, without the need to determine whether the nth denoising step is a specified denoising step, and so on.
[0084] S407, after traversing each of the N denoising steps, obtain the target generated image feature, and use the target generated image feature to generate the target generated image of the target text.
[0085] In the embodiment of the present invention, the electronic device may call the image decoding module in the target text-to-image generation model to generate the target generated image of the target text based on the target generated image feature; that is to say, the target generated image feature can be input into the image decoding module to generate the target generated image.
[0086] In summary, the embodiment of the present invention can guide the generation of the target generated image by the extracted reference graph intermediate feature and the generated character region mask image, that is to say, the reference graph intermediate feature can be injected, and the range where the reference graph intermediate feature is absorbed can be restricted by the generated character region mask image, so that the character foreground position will interact with the additionally injected reference graph intermediate feature, as Figure 5 shown.
[0087] In an embodiment of the present invention, after obtaining the initial text and the target reference image, the intermediate reference features corresponding to the target reference image can be determined, and the target text corresponding to the initial text can be determined. Based on this, the target text-to-image generation model can be called to generate a text reference image of the target text; and the text reference image can be subjected to role segmentation to obtain a generated role region mask image of the text reference image. Among them, the above intermediate reference features of the reference image include the intermediate reference features of each specified denoising step in N denoising steps, and the intermediate reference features of a denoising step include the reference layer intermediate features of each self-attention layer in the denoising module of the target text-to-image generation model at the corresponding denoising step. Correspondingly, the text encoding module in the target text-to-image generation model can be called to perform text encoding on the target text to obtain the text encoding features of the target text; each denoising step in N denoising steps can be traversed in sequence from back to front, and the currently traversed denoising step is used as the current denoising step. Based on this, the generated image features at the current denoising step can be determined; if the current denoising step is a specified denoising step, the denoising module in the target text-to-image generation model can be called, and based on the generated image features at the current denoising step, the intermediate reference features of the denoising step at the current denoising step, the text encoding features, and the generated role region mask image, the generated image features at the previous denoising step of the current denoising step can be determined; if the current denoising step is not a specified denoising step, the denoising module in the target text-to-image generation model can be called, and based on the generated image features at the current denoising step and the text encoding features, the generated image features at the previous denoising step of the current denoising step can be determined. After traversing each denoising step in N denoising steps, the target generated image features can be obtained, and the target generated image of the target text can be generated using the target generated image features.It can be seen that in the embodiments of the present invention, the intermediate feature of the noise reduction step and the generated role area mask image can be determined for each noise reduction step, and the generated image features corresponding to the corresponding noise reduction steps can be determined in sequence, so as to obtain the target generated image features, and the target generated image of the target text can be generated. Therefore, the generation of the target generated image can be guided by referring to the intermediate features in the figure, that is, the generation of the target generated image can be guided by the target reference image, and then a new image (i.e., the target generated image) that not only follows the target text (the target text is determined by the initial text, that is, also follows the initial text) but also is consistent with the role image in the target reference image (i.e., the target role image indicated by the target reference image) can be obtained. The characters in the new image and the characters in the target reference image can have different expressions, postures, and backgrounds, such as being applicable to scenarios such as comic production, novel illustration, and AI (Artificial Intelligence) character dialogue illustration. Based on this, in the embodiments of the present invention, when generating a specified role (i.e., an image with a role image as the target role image, such as a specified anime character), without retraining the model or increasing the model parameters, changes in aspects such as character expression, posture, and image background can be achieved, effectively saving the usage cost of high-performance GPUs (Graphics Processing Unit), and achieving cost reduction and efficiency improvement.
[0088] Based on the description of the related embodiments of the above image generation method, embodiments of the present invention also propose an image generation device. This image generation device can be a computer program (including program code) running in an electronic device; as Figure 6 shown, the image generation device may include an acquisition unit 601 and a processing unit 602. The image generation device can execute Figure 1 or Figure 4 the image generation method shown, that is, the image generation device can run the above units:
[0089] The acquisition unit 601 is used to acquire the initial text and the target reference image;
[0090] The processing unit 602 is used to determine the intermediate feature of the reference image corresponding to the target reference image and determine the target text corresponding to the initial text;
[0091] The processing unit 602 is further used to call the target text-to-image generation model to generate a text reference image of the target text; and perform role segmentation on the text reference image to obtain a generated role area mask image of the text reference image;
[0092] The processing unit 602 is further used to call the target text-to-image generation model to generate a target generated image of the target text based on the intermediate feature of the reference image and the generated role area mask image.
[0093] In one implementation, when the processing unit 602 determines the reference graph intermediate features corresponding to the target reference image, it may specifically be used for:
[0094] Perform role segmentation on the target reference image to obtain a reference role region mask image of the target reference image; and based on the reference role region mask image, determine a role extraction image of the target reference image;
[0095] Perform image encoding on the role extraction image to obtain role extraction image features of the role extraction image;
[0096] Determine a noisy reference sequence, and based on the noisy reference sequence and the role extraction image features, determine the denoising step reference intermediate features at each specified denoising step among N denoising steps to implement determining the reference graph intermediate features corresponding to the target reference image, where N is a positive integer; wherein the reference graph intermediate features include the denoising step reference intermediate features at each specified denoising step, and the denoising step reference intermediate features at one denoising step include the reference layer intermediate features of each self-attention layer in the denoising module of the target text-to-image model at the corresponding denoising step.
[0097] In another implementation, when the processing unit 602 determines the denoising step reference intermediate features at each specified denoising step among N denoising steps based on the noisy reference sequence and the role extraction image features, it may specifically be used for:
[0098] For the nth denoising step among the N denoising steps, if the nth denoising step is a specified denoising step, determine the first n noise values from the noisy reference sequence, where n ∈ [1, N];
[0099] Based on the first n noise values and the role extraction image features, determine the noisy image features at the nth denoising step;
[0100] Based on the noisy image features at the nth denoising step, determine the reference layer intermediate features of each self-attention layer at the nth denoising step to implement determining the denoising step reference intermediate features at the nth denoising step.
[0101] In another implementation, when the processing unit 602 determines the target text corresponding to the initial text, it may specifically be used for:
[0102] Invoke a role feature label extraction model to perform label extraction on the role extraction image to obtain at least one text label corresponding to the target reference image, where the at least one text label includes the text label of the role in the target reference image;
[0103] Combine the initial text and the at least one text label to obtain the target text corresponding to the initial text.
[0104] In another embodiment, the intermediate features of the reference graph include the intermediate reference features of the noise reduction steps at each specified noise reduction step among the N noise reduction steps. The intermediate reference features of the noise reduction steps at one noise reduction step include the intermediate reference features of each self-attention layer in the noise reduction module of the target text-to-image generation model at the corresponding noise reduction step. When the processing unit 602 calls the target text-to-image generation model and generates the target generated image of the target text based on the intermediate features of the reference graph and the generated role region mask image, it can be specifically used for:
[0105] Call the text encoding module in the target text-to-image generation model to perform text encoding on the target text to obtain the text encoding features of the target text;
[0106] Traverse each of the N noise reduction steps in reverse order, and use the currently traversed noise reduction step as the current noise reduction step;
[0107] Determine the generated image features at the current noise reduction step; if the current noise reduction step is a specified noise reduction step, call the noise reduction module in the target text-to-image generation model, and based on the generated image features at the current noise reduction step, the intermediate reference features of the noise reduction steps at the current noise reduction step, the text encoding features, and the generated role region mask image, determine the generated image features at the previous noise reduction step of the current noise reduction step; if the current noise reduction step is not a specified noise reduction step, call the noise reduction module in the target text-to-image generation model, and based on the generated image features at the current noise reduction step and the text encoding features, determine the generated image features at the previous noise reduction step of the current noise reduction step;
[0108] After traversing each of the N noise reduction steps, obtain the target generated image features, and use the target generated image features to generate the target generated image of the target text.
[0109] In another embodiment, the noise reduction module in the target text-to-image generation model includes at least one self-attention layer and at least one cross-attention layer; when the processing unit 602 calls the noise reduction module in the target text-to-image generation model and determines the generated image features at the previous noise reduction step of the current noise reduction step based on the generated image features at the current noise reduction step, the intermediate reference features of the noise reduction steps at the current noise reduction step, the text encoding features, and the generated role region mask image, it can be specifically used for:
[0110] Input the generated image features and the text encoding features at the current noise reduction step into the noise reduction module in the target text-to-image model to sequentially call each attention layer in the noise reduction module, and use the currently called attention layer as the current attention layer;
[0111] If the current attention layer is a self-attention layer, determine the original layer intermediate features of the current attention layer at the current noise reduction step; and based on the reference layer intermediate features of the current attention layer at the current noise reduction step, the original layer intermediate features of the current attention layer at the current noise reduction step, and the generated role region mask image, determine the output image features of the current attention layer at the current noise reduction step;
[0112] After calling each attention layer in the noise reduction module, based on the output image features of the last attention layer in the noise reduction module at the current noise reduction step, determine the generated image features at the previous noise reduction step of the current noise reduction step.
[0113] In another implementation, a layer intermediate feature includes an intermediate key vector and an intermediate value vector; when the processing unit 602 determines the output image features of the current attention layer at the current noise reduction step based on the reference layer intermediate features of the current attention layer at the current noise reduction step, the original layer intermediate features of the current attention layer at the current noise reduction step, and the generated role region mask image, it can specifically be used for:
[0114] Concatenate the reference intermediate key vector of the current attention layer at the current noise reduction step and the original intermediate key vector of the current attention layer at the current noise reduction step to obtain the key vector concatenation result of the current attention layer at the current noise reduction step; and concatenate the reference intermediate value vector of the current attention layer at the current noise reduction step and the original intermediate value vector of the current attention layer at the current noise reduction step to obtain the value vector concatenation result of the current attention layer at the current noise reduction step;
[0115] Based on the generated role region mask image, determine the attention mask;
[0116] Based on the key vector concatenation result, the value vector concatenation result, and the attention mask, determine the output image features of the current attention layer at the current noise reduction step.
[0117] According to an embodiment of the present invention, Figure 6Each unit in the image generation device shown can be separately or all combined into one or several other units to form, or some of them can be further split into multiple smaller units with more specific functions to form. This can achieve the same operations without affecting the realization of the technical effects of the embodiments of the present invention. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present invention, any image generation device may also include other units. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units.
[0118] According to another embodiment of the present invention, it can be achieved by running a computer program (including program code) that can execute the respective steps involved in the corresponding method shown in, for example, on a general electronic device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM). Figure 1 or Figure 4 shown in to construct an image generation device as shown in Figure 6 shown in, and to implement the image generation method of the embodiments of the present invention. The computer program can be recorded on, for example, a computer storage medium, loaded into the above-mentioned electronic device through the computer storage medium, and run therein.
[0119] After obtaining the initial text and the target reference image, the embodiments of the present invention can determine the intermediate reference features corresponding to the target reference image and determine the target text corresponding to the initial text. Based on this, the target text-to-image generation model can be called to generate a text reference image of the target text; and the text reference image can be subjected to role segmentation to obtain a generated role area mask image of the text reference image. Further, the target text-to-image generation model can be called to generate a target generated image of the target text based on the intermediate reference features of the reference image and the generated role area mask image. It can be seen that the embodiments of the present invention can generate a target generated image that follows the initial text and maintains the target role image indicated by the target reference image without training the existing target text-to-image generation model by obtaining the initial text and the target reference image, so as to conveniently generate a target generated image of the target role image; that is to say, the embodiments of the present invention can generate a target generated image including the target role without retraining the target text-to-image generation model or increasing model parameters, which can effectively save training time and storage consumption of model parameters, thereby effectively avoiding resource waste.
[0120] Based on the descriptions of the above method embodiments and apparatus embodiments, an exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiments of the present invention.
[0121] An exemplary embodiment of the present invention further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiments of the present invention.
[0122] An exemplary embodiment of the present invention further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiments of the present invention.
[0123] Referring to Figure 7 , a block diagram of an electronic device 700 that can be a server or a client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0124] As Figure 7 shown, the electronic device 700 includes a computing unit 701, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0125] Multiple components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, an output unit 707, a storage unit 708, and a communication unit 709. The input unit 706 can be any type of device capable of inputting information into the electronic device 700. The input unit 706 can receive input digital or character information and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 707 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 708 can include, but is not limited to, magnetic disks and optical discs. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0126] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above. For example, in some embodiments, the image generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. In some embodiments, the computing unit 701 can be configured to execute the image generation method in any other suitable manner (e.g., by means of firmware).
[0127] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0128] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0129] As used in the present invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) that can be used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that can be used to provide machine instructions and / or data to a programmable processor.
[0130] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0131] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of the communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0132] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other.
[0133] Also, it should be understood that the above-disclosed is only a preferred embodiment of the present invention, and of course, it cannot be used to limit the scope of the rights of the present invention. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. An image generation method, characterized in that, Including: Obtain the initial text and the target reference image; Determine the intermediate reference image features corresponding to the target reference image, and determine the target text corresponding to the initial text; Invoke the target text-to-image model to generate a text reference image of the target text; And perform role segmentation on the text reference image to obtain a generated role region mask image of the text reference image; Invoke the target text-to-image model, and based on the intermediate reference image features and the generated role region mask image, generate a target generated image of the target text.
2. The method according to claim 1, characterized in that, The determination of the intermediate reference image features corresponding to the target reference image includes: Perform role segmentation on the target reference image to obtain a reference role region mask image of the target reference image; and based on the reference role region mask image, determine a role extraction image of the target reference image; Perform image encoding on the role extraction image to obtain role extraction image features of the role extraction image; Determine a noise-added reference sequence, and based on the noise-added reference sequence and the role extraction image features, determine the reference intermediate features at each specified denoising step among N denoising steps to achieve the determination of the intermediate reference image features corresponding to the target reference image, where N is a positive integer; wherein, the intermediate reference image features include the reference intermediate features at each specified denoising step, and the reference intermediate features at one denoising step include the reference layer intermediate features of each self-attention layer in the denoising module of the target text-to-image model at the corresponding denoising step.
3. The method according to claim 2, characterized in that, The determination of the reference intermediate features at each specified denoising step among N denoising steps based on the noise-added reference sequence and the role extraction image features includes: For the nth denoising step among the N denoising steps, if the nth denoising step is a specified denoising step, then determine the first n noise values from the noise-added reference sequence, where n ∈ [1, N]; Based on the first n noise values and the role extraction image features, determine the noise-added image features at the nth denoising step; Based on the noise-added image features at the nth denoising step, determine the reference layer intermediate features of each self-attention layer at the nth denoising step to achieve the determination of the reference intermediate features at the nth denoising step.
4. The method according to claim 2, characterized in that, The determination of the target text corresponding to the initial text includes: Invoke a role feature label extraction model to perform label extraction on the role extraction image to obtain at least one text label corresponding to the target reference image, and the at least one text label includes the text label of the role in the target reference image; Combine the initial text and the at least one text label to obtain the target text corresponding to the initial text.
5. The method according to any one of claims 1-4, characterized in that, The intermediate features of the reference figure include the reference intermediate features of each specified noise reduction step in N noise reduction steps. The reference intermediate features of a noise reduction step include the reference layer intermediate features of each self-attention layer in the noise reduction module of the target text-to-image model at the corresponding noise reduction step; calling the target text-to-image model to generate the target generated image of the target text based on the intermediate features of the reference figure and the generated role region mask image includes: Calling the text encoding module in the target text-to-image model to perform text encoding on the target text to obtain the text encoding features of the target text; Traversing each noise reduction step in the N noise reduction steps in order from back to front, and taking the currently traversed noise reduction step as the current noise reduction step; Determining the generated image features at the current noise reduction step; if the current noise reduction step is a specified noise reduction step, calling the noise reduction module in the target text-to-image model, and based on the generated image features at the current noise reduction step, the reference intermediate features of the current noise reduction step, the text encoding features, and the generated role region mask image, determining the generated image features at the previous noise reduction step of the current noise reduction step; if the current noise reduction step is not a specified noise reduction step, calling the noise reduction module in the target text-to-image model, and based on the generated image features at the current noise reduction step and the text encoding features, determining the generated image features at the previous noise reduction step of the current noise reduction step; After traversing each noise reduction step in the N noise reduction steps, obtaining the target generated image features, and using the target generated image features to generate the target generated image of the target text.
6. The method according to claim 5, wherein The noise reduction module in the target text-to-image model includes at least one self-attention layer and at least one cross-attention layer; calling the noise reduction module in the target text-to-image model to determine the generated image features at the previous noise reduction step of the current noise reduction step based on the generated image features at the current noise reduction step, the reference intermediate features of the current noise reduction step, the text encoding features, and the generated role region mask image includes: Inputting the generated image features at the current noise reduction step and the text encoding features into the noise reduction module in the target text-to-image model to sequentially call each attention layer in the noise reduction module, and taking the currently called attention layer as the current attention layer; If the current attention layer is a self-attention layer, determining the original layer intermediate features of the current attention layer at the current noise reduction step; and based on the reference layer intermediate features of the current attention layer at the current noise reduction step, the original layer intermediate features of the current attention layer at the current noise reduction step, and the generated role region mask image, determining the output image features of the current attention layer at the current noise reduction step; After calling each attention layer in the noise reduction module, determining the generated image features at the previous noise reduction step of the current noise reduction step based on the output image features of the last attention layer in the noise reduction module at the current noise reduction step.
7. The method according to claim 6, wherein An intermediate feature of a layer includes an intermediate key vector and an intermediate value vector; determining the output image feature of the current attention layer at the current noise reduction step based on the reference layer intermediate feature of the current attention layer at the current noise reduction step, the original layer intermediate feature of the current attention layer at the current noise reduction step, and the generated role region mask image includes: Concatenating the reference intermediate key vector of the current attention layer at the current noise reduction step and the original intermediate key vector of the current attention layer at the current noise reduction step to obtain the key vector concatenation result of the current attention layer at the current noise reduction step; and concatenating the reference intermediate value vector of the current attention layer at the current noise reduction step and the original intermediate value vector of the current attention layer at the current noise reduction step to obtain the value vector concatenation result of the current attention layer at the current noise reduction step; Determining an attention mask based on the generated role region mask image; Determining the output image feature of the current attention layer at the current noise reduction step based on the key vector concatenation result, the value vector concatenation result, and the attention mask.
8. An image generation device, characterized in that, The apparatus includes: An acquisition unit for acquiring an initial text and a target reference image; A processing unit for determining the reference map intermediate feature corresponding to the target reference image and determining the target text corresponding to the initial text; The processing unit is further configured to call a target text-to-image generation model to generate a text reference image of the target text; and perform role segmentation on the text reference image to obtain a generated role region mask image of the text reference image; The processing unit is further configured to call the target text-to-image generation model to generate a target generated image of the target text based on the reference map intermediate feature and the generated role region mask image.
9. An electronic device, characterized in that, Comprising: A processor; And A memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-7.