Image generation method and apparatus, electronic device, and computer-readable medium
By combining the noise processing and face enhancement technology of image and text information, images that meet specific styles and face constraints are generated, solving the problem of poor generation results in the prior art and achieving more efficient image generation effects.
Patent Information
- Application Number
- PCT/CN2024/141417
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
The prior art is difficult to effectively combine image and text information to generate images that meet specific styles and face constraints, resulting in poor generation results.
By acquiring the image to be processed and its corresponding prompt text, the image generation process is performed using the noise addition processing results, face position representation data and text features, including noise addition and denoising processing at multiple time steps, combining face enhancement and non-face enhancement steps to generate images that meet the face and style constraints.
Improve the effect of image generation, so that the generated image better meets the constraints of the pending image and prompt text, especially the face and style requirements.
Smart Images

Figure CN2024141417_03072025_PF_FP_ABST
Abstract
Description
Image generation method, device, electronic device, and computer-readable medium
[0001] This application claims priority to Chinese patent application No. 202311812281.6 filed on December 26, 2023. The contents of the above-mentioned Chinese patent application are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The present disclosure relates to an image generation method, an apparatus, an electronic device, and a computer-readable medium. Background Art
[0003] For some application scenarios, these application scenarios may have the following requirements: image generation processing is performed based on images and text provided by the user to meet the user's image generation requirements, such as image generation requirements in a certain style, etc. For ease of understanding, the following examples are used to illustrate.
[0004] As an example, in some application scenarios, when a user provides a facial image and the text "Style 1", it may be necessary to generate an image with Style 1 based on the facial image so that the face described by the image with Style 1 is as close as possible to the face described by the facial image. Summary of the Invention
[0005] The present disclosure provides an image generation method, device, electronic device, and computer-readable medium, which are conducive to improving image generation effects.
[0006] In order to achieve the above objectives, the technical solutions provided by the present disclosure are as follows:
[0007] The present disclosure provides an image generation method, the method comprising:
[0008] Obtaining an image to be processed and a prompt text corresponding to the image to be processed;
[0009] Image generation processing is performed based on the noise processing result of the image to be processed, the facial position representation data of the image to be processed, and the text features of the prompt text to obtain a generated image, wherein the noise processing result includes noise processing results corresponding to multiple time steps.
[0010] In one possible implementation, the process of determining the generated image includes:
[0011] Denoising the noise data to be processed corresponding to the face enhancement time step according to the text features to obtain denoised data corresponding to the face enhancement time step; the noise data to be processed is determined based on the noise addition result of the image to be processed;
[0012] performing face adjustment processing on the denoised data corresponding to the face enhancement time step based on the face representation data corresponding to the face enhancement time step to obtain a denoising processing result corresponding to the face enhancement time step; the face representation data is obtained by performing face extraction processing on reference data corresponding to the face enhancement time step based on the face position representation data; the reference data is determined based on the image to be processed or the denoising result of the image to be processed;
[0013] The generated image is determined according to a denoising result corresponding to the face enhancement time step.
[0014] In a possible implementation manner, the noise processing result of the image to be processed includes the noise processing result corresponding to the face enhancement time step;
[0015] The reference data corresponding to the face enhancement time step is determined based on the noise addition result corresponding to the face enhancement time step.
[0016] In one possible implementation, the generated image is determined based on a denoising time step sequence; the denoising time step sequence includes the face enhancement time step;
[0017] If the arrangement position of the face enhancement time step in the denoising time step sequence is not the last arrangement position in the denoising time step sequence, the denoising time step sequence also includes a next denoising time step corresponding to the face enhancement time step, and the reference data corresponding to the face enhancement time step is determined based on the denoising processing result corresponding to the next denoising time step; the denoising processing result of the image to be processed includes the denoising processing result corresponding to the next denoising time step; the arrangement position of the face enhancement time step in the denoising time step sequence is adjacent to the arrangement position of the next denoising time step in the denoising time step sequence, and the arrangement position of the face enhancement time step in the denoising time step sequence is earlier than the arrangement position of the next denoising time step in the denoising time step sequence;
[0018] If the arrangement position of the face enhancement time step in the denoising time step sequence is the last arrangement position in the denoising time step sequence, the reference data corresponding to the face enhancement time step is determined based on the image to be processed.
[0019] In one possible implementation, the process of determining the denoising result corresponding to the face enhancement time step includes:
[0020] Obtaining face enhancement parameters corresponding to the face enhancement time step and non-face enhancement parameters corresponding to the face enhancement time step; the non-face enhancement parameters are determined based on the face enhancement parameters and the face position representation data;
[0021] determining first data according to a product of the face enhancement parameter and the face representation data corresponding to the face enhancement time step;
[0022] determining second data based on a product of the non-face enhancement parameter and the denoised data corresponding to the face enhancement time step;
[0023] The first data and the second data are added together to obtain a denoising result corresponding to the face enhancement time step.
[0024] In one possible implementation, the generated image is determined based on a denoising time step sequence; the denoising time step sequence includes the face enhancement time step;
[0025] If the face enhancement time step is located at the first position in the denoising time step sequence, the noise data to be processed corresponding to the face enhancement time step is determined based on the noise processing result corresponding to the face enhancement time step; and the noise processing result of the image to be processed includes the noise processing result corresponding to the face enhancement time step.
[0026] If the arrangement position of the face enhancement time step in the denoising time step sequence is not the first arrangement position in the denoising time step sequence, the denoising time step sequence also includes the previous denoising time step corresponding to the face enhancement time step, and the noise data to be processed corresponding to the face enhancement time step is determined based on the denoising processing result corresponding to the previous denoising time step; the denoising processing result corresponding to the previous denoising time step is obtained by processing the noise data to be processed corresponding to the previous denoising time step, and the noise data to be processed corresponding to the previous denoising time step is determined based on the denoising processing result of the image to be processed; the arrangement position of the face enhancement time step in the denoising time step sequence is adjacent to the arrangement position of the previous denoising time step in the denoising time step sequence, and the arrangement position of the face enhancement time step in the denoising time step sequence is later than the arrangement position of the previous denoising time step in the denoising time step sequence.
[0027] In one possible implementation, the process of determining the generated image includes a process of determining a denoising result corresponding to at least one face enhancement time step;
[0028] For any of the face enhancement time steps, the process of determining the denoising result corresponding to the face enhancement time step includes:
[0029] performing denoising on the noise data to be processed corresponding to the face enhancement time step according to the text features to obtain denoised data corresponding to the face enhancement time step; the noise data to be processed is determined based on a noise addition result of the image to be processed;
[0030] Based on the facial representation data corresponding to the face enhancement time step, face adjustment processing is performed on the denoised data corresponding to the face enhancement time step to obtain a denoising processing result corresponding to the face enhancement time step; the facial representation data is obtained by performing face extraction processing on the reference data corresponding to the face enhancement time step based on the facial position representation data; the reference data is determined based on the image to be processed or the denoising processing result of the image to be processed.
[0031] In one possible implementation, the generated image is determined based on a denoising time step sequence; and the at least one face enhancement time step includes each time step in the denoising time step sequence.
[0032] In one possible implementation, the process of determining the generated image further includes a process of determining a denoising result corresponding to at least one non-enhancement time step;
[0033] For any of the non-enhanced time steps, a process of determining a denoising result corresponding to the non-enhanced time step includes:
[0034] De-noising is performed on the noise data to be processed corresponding to the non-enhanced time step according to the text features to obtain a de-noising result corresponding to the non-enhanced time step.
[0035] In one possible implementation, the generated image is determined based on a denoising time step sequence; the at least one face enhancement time step includes a portion of the time steps in the denoising time step sequence; and the at least one non-enhancement time step includes the remaining time steps in the denoising time step sequence except the portion of the time steps.
[0036] In one possible implementation, the generated image is determined based on a denoising time step sequence; the denoising time step sequence includes the at least one face enhancement time step and the at least one non-enhancement time step; and in the denoising time step sequence, one or more non-enhancement time steps exist between any two face enhancement time steps.
[0037] In one possible implementation, the generated image is determined based on a denoising time step sequence;
[0038] The process of determining the at least one face enhancement time step and the at least one non-enhancement time step comprises:
[0039] Acquiring image generation requirement description information corresponding to the image to be processed; the image generation requirement description information includes face preservation degree representation data and / or style degree representation data;
[0040] Determining facial enhancement degree representation data based on the image generation requirement description information;
[0041] The at least one face enhancement time step and the at least one non-enhancement time step are determined from the denoised time step sequence according to the face enhancement degree characterization data.
[0042] In one possible implementation, the prompt text includes positive prompt content and negative prompt content;
[0043] The process of determining the text features of the prompt text includes:
[0044] Obtaining a feature vector of the positive prompt content and a feature vector of the negative prompt content;
[0045] Comparing the size of the feature vector of the positive prompt content with the size of the feature vector of the negative prompt content to obtain a comparison result;
[0046] The text feature of the prompt text is determined according to the comparison result, the feature vector of the positive prompt content, and the feature vector of the negative prompt content.
[0047] In one possible implementation, determining the text feature of the prompt text based on the comparison result, the feature vector of the positive prompt content, and the feature vector of the negative prompt content includes:
[0048] When the comparison result indicates that the size of the first feature vector is smaller than the size of the second feature vector, data expansion processing is performed on the first feature vector using part of the data in the first feature vector to obtain an expanded feature vector, the size of the expanded feature vector being equal to the size of the second feature vector; the first feature vector is the feature vector of the positive prompt content, and the second feature vector is the feature vector of the negative prompt content; or, the first feature vector is the feature vector of the negative prompt content, and the second feature vector is the feature vector of the positive prompt content;
[0049] The expanded feature vector and the second feature vector are concatenated to obtain text features of the prompt text.
[0050] In one possible implementation, the first feature vector includes semantic representation data of at least one candidate semantic unit;
[0051] The process of determining the partial data includes:
[0052] Determining a target semantic unit from the at least one candidate semantic unit based on the importance representation data of each candidate semantic unit, wherein the importance representation data of the target semantic unit is higher than the importance representation data of other candidate semantic units in the at least one candidate semantic unit except the target semantic unit;
[0053] The partial data is determined according to the semantic representation data of the target semantic unit.
[0054] The present disclosure provides an image generating device, comprising:
[0055] An acquiring unit, configured to acquire an image to be processed and a prompt text corresponding to the image to be processed;
[0056] A generation unit is used to perform image generation processing based on the noise processing result of the image to be processed, the facial position representation data of the image to be processed, and the text features of the prompt text to obtain a generated image, wherein the noise processing result includes the noise processing results corresponding to multiple time steps.
[0057] The present disclosure provides an electronic device, the device comprising: a processor and a memory;
[0058] The memory is used to store instructions or computer programs;
[0059] The processor is configured to execute the instructions or computer program in the memory so that the electronic device executes the image generation method provided by the present disclosure.
[0060] The present disclosure provides a computer-readable medium having instructions or a computer program stored therein. When the instructions or the computer program are executed on a device, the device executes the image generation method provided by the present disclosure.
[0061] The present disclosure provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, wherein the computer program contains program code for executing the image generation method provided by the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0063] FIG1 is a flow chart of an image generation method provided by an embodiment of the present disclosure;
[0064] FIG2 is a schematic diagram of an image generation process provided by an embodiment of the present disclosure;
[0065] FIG3 is a schematic structural diagram of an image generating device provided by an embodiment of the present disclosure; and
[0066] FIG4 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0067] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0068] To better understand the technical solutions provided by the present disclosure, the image generation method provided by the present disclosure is described below with reference to some accompanying figures. As shown in Figure 1, the image generation method provided by an embodiment of the present disclosure includes the following steps S1-S2. Figure 1 is a flow chart of an image generation method provided by an embodiment of the present disclosure.
[0069] S1: Obtain an image to be processed and a prompt text corresponding to the image to be processed.
[0070] The image to be processed refers to an image required to be referenced during image generation processing, such as image 1 shown in FIG2 , so that the image to be processed is used to provide facial constraints, so that the image finally generated satisfies the facial constraints.
[0071] Furthermore, the present disclosure does not limit the implementation of the above-mentioned image to be processed. For example, the image to be processed may be an image provided by a user via an input device. It should be noted that the present disclosure does not limit the implementation of the input device. For example, it may be implemented using any image acquisition device, such as a camera. For another example, it may be implemented using any information input device, such as a keyboard, mouse, or stylus.
[0072] Furthermore, for the image to be processed, the corresponding prompt text refers to text describing other constraints besides the facial constraint, which is referenced during image generation. For example, the prompt text may include the string "hair is blue." It should be noted that this disclosure does not limit the implementation of these other constraints. For example, in some application scenarios, these other constraints may include at least a style constraint.
[0073] Furthermore, in some application scenarios, in order to better improve the image generation effect, the present disclosure also provides a possible implementation of the above-mentioned prompt text, in which the prompt text can include positive prompt content and negative prompt content. The positive prompt content is used to constrain the state that the final generated image should present, such as blue hair, relatively high image quality, relatively harmonious relationship between the various parts of the image, and relatively natural light distribution in the image; the negative prompt content is used to constrain the state that the final generated image should avoid, such as relatively low image quality, poor coordination between the various parts of the image, and unnatural light distribution in the image.
[0074] It should be noted that the present disclosure does not limit the implementation method of the "positive prompt content" in the above paragraph. For example, the positive prompt content may include the positive prompt words shown in Figure 2. In addition, the present disclosure does not limit the method of obtaining the positive prompt content. For example, in some application scenarios, the positive prompt content may include a positive prompt text input by the user, so that the positive prompt content can express the positive constraints described by the positive prompt text, such as constraints such as hair being blue. For another example, in other application scenarios, the positive prompt content may include a positive prompt text input by the user and a pre-set default positive prompt word, so that the positive prompt content can not only express the positive constraints described by the positive prompt text, but also describe the positive constraints described by the default positive prompt word, such as constraints such as relatively high image quality, relatively coordinated between the various parts of the image, and relatively natural light distribution in the image. Among them, the positive constraint is used to describe the state that the final generated image should present.
[0075] It should also be noted that the present disclosure does not limit the implementation method of the above-mentioned negative prompt content. For example, the negative prompt content may include the negative prompt words shown in Figure 2. In addition, the present disclosure does not limit the method of obtaining the negative prompt content. For example, in some application scenarios, the negative prompt content may include a negative prompt text input by the user and at least one of the pre-set default negative prompt words, so that the negative prompt content can express the negative constraints described by the negative prompt text, such as the constraint that the hair cannot be red, and / or the negative constraints described by the default negative prompt words, such as the image quality is relatively low, the degree of coordination between the various parts of the image is not high, the light distribution in the image is unnatural, etc. Among them, the negative constraint is used to describe the state that the final generated image should avoid.
[0076] Based on the relevant content of S1 above, it can be seen that in some application scenarios, if you want to perform image generation processing, you need to first obtain the image to be processed and the prompt text corresponding to the image to be processed, so that the image to be processed is used to provide facial constraints, and the prompt text is used to provide one or more constraints other than the facial constraints, such as style constraints, etc., so that an image that meets these constraints can be generated based on these two data subsequently.
[0077] S2: performing image generation processing based on the noise processing result of the image to be processed, the facial position representation data of the image to be processed, and the text features of the prompt text to obtain a generated image, wherein the noise processing result includes noise processing results corresponding to multiple time steps.
[0078] Among them, the noise processing result of the image to be processed refers to the result obtained by performing noise processing on the image to be processed; and the present disclosure does not limit the implementation method of the noise processing result. For example, when the noise processing result of the image to be processed is determined based on a noise time step sequence, the noise processing result of the image to be processed may include the noise processing results corresponding to each time step in the noise time step sequence. It can be seen that under one possible implementation method, the noise processing result of the image to be processed includes the noise processing results corresponding to multiple time steps. It should be noted that the present disclosure does not limit the implementation method of the multiple time steps. For example, the multiple time steps may include at least two noise time steps described below.
[0079] The noise addition time step sequence refers to the time step sequence required for performing noise addition processing on the above-mentioned image to be processed; and the noise addition time step sequence includes at least two noise addition time steps. The noise addition time step refers to the time step required for use when performing noise addition processing. In addition, the present disclosure does not limit the noise addition time step sequence. For example, the noise addition time step sequence can be implemented using the time step sequence {1, 2, ..., N}, so that the noise addition time step sequence can include time step 1, time step 2, ... (and so on), and time step N, so that the noise addition time step sequence includes N noise addition time steps arranged in sequence. The time step n refers to the time step in the nth arrangement position in the noise addition time step sequence, n is a positive integer, n≤N, and N is a positive integer.
[0080] Based on the above two paragraphs, it can be seen that when the above noise time step sequence includes time step 1, time step 2, ... (and so on), and time step N, the above noise processing result of the image to be processed can include the noise processing result corresponding to time step 1, the noise processing result corresponding to time step 2, ... (and so on), and the noise processing result corresponding to time step N. Among them, the noise processing result corresponding to time step n is obtained by performing noise processing on the image to be processed using the noise processing process corresponding to time step n, so that the noise processing result corresponding to time step n can represent the noise processing result of the image to be processed at time step n. The noise processing process corresponding to time step n refers to the process required to use when performing noise processing at time step n; and the present disclosure does not limit the implementation method of the noise processing process corresponding to time step n. For example, the noise processing process corresponding to time step n can specifically be: performing noise processing on the image to be processed n times to obtain the noise processing result corresponding to time step n. For another example, the noise addition process corresponding to time step n can specifically be: adding the noise corresponding to time step n to the image to be processed, and obtaining the noise addition result corresponding to time step n. The noise corresponding to time step n refers to the noise required for performing noise addition processing at time step n; and the amount of noise carried by the noise corresponding to time step n is positively correlated with the position of time step n in the sequence of noise addition time steps. In addition, the present disclosure does not limit the implementation of the noise corresponding to time step n. For example, the noise corresponding to time step n can be pre-set. n is a positive integer, n≤N, and N is a positive integer.
[0081] Based on the above, in one possible implementation, for the time step at the nth position in the noise time step sequence, such as time step n, the noise processing result corresponding to time step n can be obtained by performing n noise processing on the image to be processed, so that the noise processing result corresponding to time step n is obtained by performing noise processing on the image to be processed using the position of time step n in the noise time step sequence as the number of noise processing times, so that the amount of noise carried by the noise processing result corresponding to time step n is positively correlated with the position of time step n in the noise time step sequence. Where n is a positive integer, n≤N, and N is a positive integer.
[0082] Based on the relevant content of the noise processing results of the image to be processed above, it can be known that the noise processing results of the image to be processed may include noise processing results corresponding to at least two noise time steps, so that the noise processing results of the image to be processed can describe the noise processing results at different time steps, so that the noise processing results of the image to be processed can better represent the noise situation of the image to be processed, so that the noise processing results of the image to be processed can be used to better perform image generation processing later.
[0083] In addition, for the above-mentioned image to be processed, the facial position representation data of the image to be processed is used to describe the position of the face in the image to be processed; and the present disclosure does not limit the implementation method of the facial position representation data. For example, the facial position representation data may include the position coordinates of the area where the face is located in the image to be processed. For another example, in order to better improve the representation effect of the facial position, the facial position representation data may be implemented using a facial mask image of the image to be processed. The facial mask image is used to describe the position of the facial area in the image to be processed in an image manner, so that the corresponding facial information can be extracted later by multiplying the facial mask image with the image to be processed or the noise processing result of the image to be processed; and the present disclosure does not limit the implementation method of the facial mask image. For example, when the image to be processed is image 1 shown in Figure 2, the facial mask image may be implemented using the facial mask image shown in Figure 2. It should be noted that the present disclosure does not limit the method for obtaining the facial position representation data. For example, it can be implemented using any existing or future method that can perform facial position determination processing, such as any facial mask image determination method.
[0084] In addition, for the prompt text corresponding to the image to be processed above, the text features of the prompt text are used to characterize the information carried by the prompt text, such as the constraint information described by the prompt text, etc.; and the present disclosure does not limit the acquisition process of the text features. For example, it can be implemented using any existing or future method that can extract features from a text, such as a text encoding method, etc.
[0085] In fact, in order to better improve the text representation effect, the present disclosure also provides a possible implementation method of the process of determining the text features of the above prompt text. Under this implementation method, when the prompt text includes positive prompt content and negative prompt content, the process of determining the text features of the prompt text may include the following steps 11-13.
[0086] Step 11: Obtain the feature vector of the positive prompt content and the feature vector of the negative prompt content.
[0087] The feature vector of the positive prompt content is used to represent the information carried by the positive prompt content, such as positive constraints. Moreover, the present disclosure does not limit the implementation of the feature vector of the positive prompt content. For example, when the positive prompt content includes the positive prompt word shown in Figure 2, the feature vector of the positive prompt content may include the positive prompt word vector shown in Figure 2. The positive prompt word vector refers to the word vector determined by the positive prompt word.
[0088] In addition, the present disclosure does not limit the method for obtaining the feature vector of the positive prompt content above. For example, it can be implemented using any existing or future text feature extraction method, such as a word vector extraction method.
[0089] The feature vector of the negative prompt content is used to represent the information carried by the negative prompt content, such as negative constraints. Moreover, the present disclosure does not limit the implementation of the feature vector of the negative prompt content. For example, when the negative prompt content includes the negative prompt word shown in Figure 2, the feature vector of the negative prompt content may include the negative prompt word vector shown in Figure 2. The negative prompt word vector refers to the word vector determined by the negative prompt word.
[0090] In addition, the present disclosure does not limit the method for obtaining the feature vector of the negative prompt content. For example, the method for obtaining the feature vector of the negative prompt content is similar to the method for obtaining the feature vector of the positive prompt content.
[0091] Step 12: Compare the size of the feature vector of the positive prompt content with the size of the feature vector of the negative prompt content to obtain a comparison result.
[0092] The size of the feature vector of the positive prompt content is used to describe the amount of data carried by the feature vector. Furthermore, this disclosure does not limit the size of the feature vector of the positive prompt content. For example, when the positive prompt content includes M characters, the size of the feature vector of the positive prompt content may be M×Q, where M is a positive integer and Q is a positive integer representing the number of column vectors in the feature vector of the positive prompt content.
[0093] The size of the feature vector of the negative prompt content is used to describe the amount of data carried by the feature vector; and the present disclosure does not limit the size of the feature vector of the negative prompt content. For example, when the negative prompt content includes K characters, the size of the feature vector of the negative prompt content can be K×Q. Where K is a positive integer and Q is a positive integer representing the number of column vectors in the feature vector of the negative prompt content.
[0094] The comparison result is used to represent the relative size between the size of the feature vector of the positive prompt content and the size of the feature vector of the negative prompt content.
[0095] Step 13: Determine the text features of the prompt text according to the comparison result, the feature vector of the positive prompt content, and the feature vector of the negative prompt content.
[0096] It should be noted that the present disclosure does not limit the implementation of step 13 above. For example, in some application scenarios, step 13 may specifically be: first, find the feature fusion algorithm corresponding to the comparison result above from a pre-set mapping relationship; then, according to this found feature fusion algorithm, perform feature fusion processing on the feature vector of the positive prompt content above and the feature vector of the negative prompt content above to obtain the text features of the prompt text above. The mapping relationship is used to record the method required for fusing two features with different relative size relationships; and the mapping relationship can be set in advance based on the actual application scenario.
[0097] In fact, in order to better improve the degree of constraint of the above prompt text on the image generation process, the present disclosure also provides a possible implementation of the above step 13. Under this implementation, the step 13 may specifically include the following steps 131-132.
[0098] Step 131: When the comparison result indicates that the size of the first feature vector is smaller than the size of the second feature vector, data expansion processing is performed on the first feature vector using a portion of the data in the first feature vector to obtain an expanded feature vector, so that the size of the expanded feature vector is equal to the size of the second feature vector. The first feature vector is the feature vector of the positive prompt content, and the second feature vector is the feature vector of the negative prompt content; alternatively, the first feature vector is the feature vector of the negative prompt content, and the second feature vector is the feature vector of the positive prompt content.
[0099] The first feature vector refers to a feature vector with a smaller size determined from the feature vector of the positive prompt content and the feature vector of the negative prompt content, so that the first feature vector can represent the feature vector that needs to be processed for data expansion. It can be seen that under one possible implementation, the process of determining the first feature vector may include: if the size of the feature vector of the positive prompt content is smaller than the size of the feature vector of the negative prompt content, then the feature vector of the positive prompt content may be determined as the first feature vector; however, if the size of the feature vector of the positive prompt content is larger than the size of the feature vector of the negative prompt content, then the feature vector of the negative prompt content may be determined as the first feature vector.
[0100] The second feature vector refers to the feature vector with the larger size determined from the feature vector of the positive prompt content and the feature vector of the negative prompt content. Thus, in one possible implementation, the process of determining the second feature vector may include: if the size of the feature vector of the positive prompt content is smaller than the size of the feature vector of the negative prompt content, then the feature vector of the negative prompt content may be determined as the second feature vector; however, if the size of the feature vector of the positive prompt content is larger than the size of the feature vector of the negative prompt content, then the feature vector of the positive prompt content may be determined as the second feature vector.
[0101] In addition, for the first eigenvector mentioned above, the partial data in the first eigenvector refers to the data existing in the first eigenvector and required for use when performing data expansion processing on the first eigenvector; and the present disclosure does not limit the determination process of the partial data. For example, when the size of the first eigenvector is H1×W, and the size of the second eigenvector mentioned above is H2×W, if D H =H2-H1, then D can be randomly extracted from the first eigenvector H row vectors as the partial data, so that the partial data can be used to perform data expansion processing on the first feature vector to obtain the expanded feature vector, so that the D HEach row vector in the D row vectors appears twice in the expanded feature vector, so that the expanded feature vector can double represent the D H The semantics represented by the row vectors, so that the expanded feature vector can further emphasize the semantics represented by the D on the premise of fully expressing the semantics represented by the first feature vector. H The semantics represented by the row vectors is improved. Wherein, H1 represents the number of row vectors in the first eigenvector. H2 represents the number of row vectors in the second eigenvector. W represents the number of column vectors in the first eigenvector or the number of column vectors in the second eigenvector.
[0102] Based on the content of the previous paragraph, it can be seen that in a possible implementation method, the above step 131 can be specifically as follows: when the above comparison result indicates that the size of the first eigenvector is smaller than the size of the second eigenvector, first calculate the difference between the number of row vectors of the second eigenvector and the number of row vectors of the first eigenvector; then randomly extract some row vectors from the first eigenvector so that the number of extracted row vectors is equal to the difference; then, randomly insert these extracted row vectors into the first eigenvector to obtain an expanded eigenvector, so that the number of row vectors of the expanded eigenvector is equal to the number of row vectors of the second eigenvector, so that the size of the expanded eigenvector is equal to the size of the second eigenvector, thereby realizing the size alignment processing of the eigenvectors. Among them, because there are some row vectors with the same data in the expanded feature vector, the expanded feature vector can better emphasize the semantics carried by these row vectors with the same data, so that the expanded feature vector can further emphasize the semantics represented by these extracted row vectors on the premise of fully representing the semantics represented by the first feature vector, and thus the expanded feature vector can better represent the constraint information represented by the first feature vector, so that the expanded feature vector can better constrain the image generation process, which is beneficial to improve the text represented by the first feature vector, such as positive prompt content or negative prompt content, in the image generation process. The constraint strength presented by the above prompt text in the image generation process is further improved, so that the final generated image can better meet the constraints described by the prompt text, which is beneficial to improve the image generation effect. It should be noted that the text represented by the first feature vector refers to the text used to generate the first feature vector. For example, if the first feature vector is the feature vector of the positive prompt content above, then the text represented by the first feature vector is the positive prompt content; if the first feature vector is the feature vector of the negative prompt content above, then the text represented by the first feature vector is the negative prompt content.
[0103] In fact, for a text content, such as the text represented by the first eigenvector above, the semantics carried by some characters in the text content are more important, while the semantics carried by other characters are less important, so that the semantic importance of different characters in the text content may be different. Based on this, in order to better improve the degree of constraint of the above prompt text on the image generation process, the present disclosure also provides a possible implementation method of the process of determining part of the data in the first eigenvector. Under this implementation method, when the first eigenvector includes semantic representation data of at least one candidate semantic unit, the process of determining part of the data in the first eigenvector may include the following steps 1311-1312.
[0104] Step 1311: Determine a target semantic unit from at least one candidate semantic unit based on the importance representation data of each candidate semantic unit, wherein the importance representation data of the target semantic unit is higher than the importance representation data of other candidate semantic units in the at least one candidate semantic unit except the target semantic unit.
[0105] The candidate semantic unit refers to a semantically-carrying unit existing in the text represented by the first feature vector above, so that at least one candidate semantic unit above may include some or all semantic units existing in the text represented by the first feature vector. It should be noted that the present disclosure does not limit the implementation method of the semantic unit. For example, for any semantic unit, the semantic unit can be a character, word, phrase, or short sentence.
[0106] Based on the content of the previous paragraph, it can be seen that if the first feature vector above is the feature vector of the positive prompt content above, then the text represented by the first feature vector is the positive prompt content, and the at least one candidate semantic unit above may include part or all of the semantic units existing in the positive prompt content; if the first feature vector is the feature vector of the negative prompt content above, then the text represented by the first feature vector is the negative prompt content, and the at least one candidate semantic unit may include part or all of the semantic units existing in the negative prompt content.
[0107] In addition, for the t-th candidate semantic unit, the importance representation data of the t-th candidate semantic unit is used to represent the importance of the t-th candidate semantic unit in the text represented by the first feature vector above, such as the semantic importance of the semantics described by the t-th candidate semantic unit in the text represented by the first feature vector above. Wherein, t is a positive integer, t≤T, T is a positive integer, and T represents the number of units in the at least one candidate semantic unit above.
[0108] In addition, the present disclosure does not limit the process of determining the importance representation data of the t-th candidate semantic unit mentioned above. For example, it can be implemented by any existing or future method that can determine the importance of a semantic unit, such as a method implemented by a pre-built semantic unit importance determination model or the term frequency-inverse document frequency index (TF-IDF). The semantic unit importance determination model is used to determine the importance of each semantic unit in the input text of the semantic unit importance determination model; and the present disclosure does not limit the implementation method of the semantic unit importance determination model.
[0109] Furthermore, for the t-th candidate semantic unit above, the semantic representation data of the t-th candidate semantic unit refers to the data present in the first feature vector above and used to represent the t-th candidate semantic unit, so that the semantic representation data of the t-th candidate semantic unit is used to represent the semantics described by the t-th candidate semantic unit. Wherein, t is a positive integer, t≤T, T is a positive integer, and T represents the number of units in the at least one candidate semantic unit above.
[0110] Based on the content of the above paragraph, it can be seen that for the text represented by the first feature vector above, if the text includes at least one candidate semantic unit, the first feature vector generated based on the text can include the semantic representation data of each candidate semantic unit, so that the first feature vector can represent the semantics described by these candidate semantic units.
[0111] The target semantic unit refers to a semantic unit that exists in the text represented by the first feature vector above and carries relatively important semantics; and the target semantic unit meets the following conditions: the importance representation data of the target semantic unit is higher than the importance representation data of other candidate semantic units other than the target semantic unit in the at least one candidate semantic unit above. It can be seen that the target semantic unit can be a candidate semantic unit that exists in the at least one candidate semantic unit and has higher importance representation data.
[0112] Based on the relevant content of step 1311 above, it can be known that for the text represented by the first feature vector above, such as positive prompt text or negative prompt text, the importance representation data of each semantic unit in the text can be calculated first; then, based on these importance representation data, semantic units carrying relatively important semantics are selected from the text as target semantic units, so that the semantic representation data of the target semantic unit can be used to perform data expansion processing on the first feature vector in the future.
[0113] Step 1312: Determine the above-context data based on the semantic representation data of the target semantic unit.
[0114] It should be noted that the present disclosure does not limit the implementation method of the above step 1312. For example, it can be specifically as follows: if the size of the semantic representation data of the target semantic unit is equal to the difference between the size of the feature vector of the above positive prompt content and the size of the feature vector of the above negative prompt content, then the semantic representation data of the target semantic unit is directly determined as the above partial data. For another example, when the number of the target semantic units is multiple, the step 1312 can be specifically as follows: based on the difference, a certain number of target semantic units are selected from these target semantic units so that the size and value of the semantic representation data of these selected target semantic units are equal to the difference, and these selected target semantic units are regarded as the above partial data.
[0115] It should also be noted that the present disclosure does not limit the selection method of the "certain number of target semantic units" in the above paragraph. For example, the "certain number of target semantic units" can be implemented by random selection. For another example, the selection process of the "certain number of target semantic units" can be specifically as follows: first, based on the importance representation data of multiple target semantic units, these target semantic units are sorted in a manner of gradually decreasing importance; then, based on the difference between the size of the feature vector of the positive prompt content above and the size of the feature vector of the negative prompt content above, one or more target semantic units that are arranged relatively high are selected from these target semantic units, so that the size and value of the semantic representation data of these selected target semantic units are equal to the difference.
[0116] Based on the relevant contents of steps 1311 to 1312 above, it can be known that for the first feature vector above, if the text represented by the first feature vector above includes at least one candidate semantic unit, then the first feature vector includes semantic representation data of the at least one candidate semantic unit, and part of the data in the first feature vector may include semantic representation data of a target semantic unit selected from the at least one candidate semantic unit. The importance representation data of the target semantic unit is higher than the importance representation data of other candidate semantic units in the at least one candidate semantic unit except the target semantic unit.
[0117] In addition, the present disclosure does not limit the implementation method of the above step of "using part of the data in the first feature vector to perform data expansion processing on the first feature vector to obtain an expanded feature vector". For example, in some application scenarios, such as scenarios where the arrangement order of row vectors does not affect the semantics represented by the feature vector, when the part of the data includes multiple row vectors, the step can specifically be: for any row vector in the part of the data, randomly insert the row vector into a row vector position in the first feature vector, so that the number of row vectors in the first feature vector after insertion is greater than the number of row vectors in the first feature vector before insertion.
[0118] Based on the relevant content of step 131 above, it can be known that after obtaining the above comparison result, if the comparison result indicates that the size of the first eigenvector is smaller than the size of the second eigenvector, it can be determined that the first eigenvector needs to be data expanded, so part of the data in the first eigenvector can be used to perform data expansion on the first eigenvector to obtain an expanded eigenvector, so that the size of the expanded eigenvector is equal to the size of the second eigenvector, thereby realizing the size alignment of the eigenvectors.
[0119] Step 132: Concatenate the expanded feature vector of the above text with the second feature vector to obtain the text features of the above prompt text.
[0120] It should be noted that the present disclosure does not limit the implementation method of the above step 132. For example, it can be implemented by using any existing or future method that can splice two features into one feature.
[0121] Based on the relevant contents of steps 131 to 132 above, it can be known that, in a possible implementation manner, after obtaining the comparison result between the size of the feature vector of the positive prompt content above and the size of the feature vector of the negative prompt content above, if the comparison result indicates that the size of the feature vector of the positive prompt content is larger than the size of the feature vector of the negative prompt content, then part of the data in the feature vector of the negative prompt content can be used to perform data expansion processing on the feature vector of the negative prompt content to obtain an expanded feature vector of the negative prompt content, so that the size of the expanded feature vector of the negative prompt content is equal to the size of the feature vector of the positive prompt content; then the expanded feature vector of the negative prompt content is spliced with the feature vector of the positive prompt content to obtain the text features of the prompt text above; if the If the comparison result indicates that the size of the feature vector of the positive prompt content is smaller than the size of the feature vector of the negative prompt content, then part of the data in the feature vector of the positive prompt content can be used to perform data expansion processing on the feature vector of the positive prompt content to obtain the expanded feature vector of the positive prompt content, so that the size of the expanded feature vector of the positive prompt content is equal to the size of the feature vector of the negative prompt content; then the expanded feature vector of the positive prompt content is spliced with the feature vector of the negative prompt content to obtain the text features of the prompt text; if the comparison result indicates that the size of the feature vector of the positive prompt content is equal to the size of the feature vector of the negative prompt content, then the feature vector of the positive prompt content is directly spliced with the feature vector of the negative prompt content to obtain the text features of the prompt text.
[0122] Based on the relevant content of steps 11 to 13 above, it can be known that in some application scenarios, for the prompt text corresponding to the above image to be processed, if the prompt text includes positive prompt content and negative prompt content, the feature vector of the positive prompt content and the feature vector of the negative prompt content can be obtained first; then these two feature vectors are aligned and spliced to obtain the text features of the prompt text, so that the text features can further emphasize certain semantics, such as some more important semantics, on the premise of fully expressing the semantics described by the prompt text, so that the text features can better express the constraints carried by the prompt text, which is beneficial to improving the constraint effect of the prompt text on the image generation process, thereby improving the image generation effect.
[0123] The generated image refers to an image generated based on the image to be processed and its corresponding prompt text. For example, if the image to be processed is Image 1 shown in Figure 2, and the prompt text corresponding to the image to be processed includes the positive prompt words and negative prompt words shown in Figure 2, the generated image can be implemented using Image 2 shown in Figure 2.
[0124] In addition, the present disclosure does not limit the implementation of S2 above. For example, in some application scenarios, S2 can be implemented with the aid of a pre-built image generator, so that the image generator can be used to perform image generation processing based on the noise processing result of the image to be processed, the facial position representation data of the image to be processed, and the text features of the prompt text, to obtain and output a generated image. In particular, because the image generator can better perform facial generation processing based on the facial position representation data, the generated image output by the image generator can better meet the facial constraints described by the image to be processed, thereby facilitating improved image generation effects.
[0125] For example, in some application scenarios, the generated image described above may be determined based on a denoising time step sequence. The denoising time step sequence refers to a time step sequence required for denoising during the image generation process; and the denoising time step sequence may include at least one denoising time step. The denoising time step refers to a time step required for denoising.
[0126] In addition, for the above denoised time step sequence, the denoised time step sequence can include some or all of the time steps in the above denoised time step sequence. For example, when the above denoised time step sequence includes time step 1, time step 2, ..., and time step N, the denoised time step sequence can be {S M , S M-1 , ..., S1} this time step sequence, so that the denoised time step sequence can include time step S M , time step S M-1 , ... (and so on), and time step S1; and the S M 、The S M-1 , ..., and S1 all belong to the data interval [1, N]. Among them, the time step S m It refers to the time step at the M-m+1th position in the denoised time step sequence, m≤M, M is a positive integer. It should be noted that the time steps in the denoised time step sequence are arranged in descending order, so that the time step S m The arrangement position in the denoising time step sequence is greater than the time step S m-1 The arrangement position in the denoising time step sequence is at the front, and the time step S m The value is greater than the time step S m-1 The value of S m ≤S M , the S m-1 ≥S1. For easier understanding, the following is an explanation with examples.
[0127] As an example, when the above noisy time step sequence is the time step sequence of {1, 2, ..., 100}, the noisy time step sequence may include time step 1, time step 2, ..., and time step 100, and the denoised time step sequence may be the time step sequence of {70, 54, 36, 25, 18}, so that the denoised time step sequence includes time step 70, time step 54, time step 36, time step 25 and time step 18.
[0128] In addition, the present disclosure is not limited to the image generation process implemented based on the denoising time step sequence. For example, it can be implemented using any existing or future method that can perform image generation processing based on the denoising time step sequence.
[0129] Furthermore, to further enhance image generation, some or all of the time steps in the denoising time step sequence can be used for face enhancement. This helps to enhance the constraint effect exerted by the facial information carried by the processed image on the image generation process. Based on this, the present disclosure also provides a possible implementation of S2 above. In this implementation, S2 may specifically include steps 21-23 below.
[0130] Step 21: De-noising the noise data to be processed corresponding to the face enhancement time step based on the above text features to obtain the denoised data corresponding to the face enhancement time step; the noise data to be processed is determined based on the noise processing result of the image to be processed above.
[0131] The face enhancement time step refers to the denoising time step that exists in the above denoising time step sequence and needs to be subjected to face enhancement processing.
[0132] In addition, for the face enhancement time step described above, the noise data to be processed corresponding to the face enhancement time step refers to the data that needs to be subjected to noise removal processing at the face enhancement time step; and the noise data to be processed is determined based on the noise addition processing result of the image to be processed described above.
[0133] In addition, the present disclosure does not limit the process of determining the noise data to be processed corresponding to the above face enhancement time step. For example, in one possible implementation, when the above generated image is determined based on a denoising time step sequence, and the denoising time step sequence includes the face enhancement time step, the process of obtaining the noise data to be processed corresponding to the face enhancement time step may include the following steps 211-212.
[0134] Step 211: If the face enhancement time step is located at the first position in the denoising time step sequence, then based on the noise processing result corresponding to the face enhancement time step, noise data to be processed corresponding to the face enhancement time step is determined. The noise processing result of the image to be processed includes the noise processing result corresponding to the face enhancement time step.
[0135] In this disclosure, for the above face enhancement time step, such as the above time step S M For example, if the face enhancement time step is in the first position in the denoising time step sequence, it can be determined that the face enhancement time step is the first time step in the denoising time step sequence, and thus it can be determined that the denoising process has not been performed. Therefore, the denoising process result corresponding to the face enhancement time step can be directly extracted from the denoising process result of the image to be processed above, such as the time step S M The corresponding noise processing result is obtained, and the noise processing result corresponding to the face enhancement time step is used as the noise data to be processed corresponding to the face enhancement time step, so that the denoising processing under the face enhancement time step can be completed based on the noise data to be processed and the noise data to be processed corresponding to other time steps in the denoising time step sequence can be determined.
[0136] Step 212: If the arrangement position of the face enhancement time step in the denoising time step sequence is not the first arrangement position in the denoising time step sequence, the denoising time step sequence further includes the previous denoising time step corresponding to the face enhancement time step, and based on the denoising processing result corresponding to the previous denoising time step, the noise data to be processed corresponding to the face enhancement time step is determined. The denoising processing result corresponding to the previous denoising time step is obtained by processing the noise data to be processed corresponding to the previous denoising time step, and the noise data to be processed corresponding to the previous denoising time step is determined based on the noise addition processing result of the image to be processed. The arrangement position of the face enhancement time step in the denoising time step sequence is adjacent to the arrangement position of the previous denoising time step in the denoising time step sequence, and the arrangement position of the face enhancement time step in the denoising time step sequence is later than the arrangement position of the previous denoising time step in the denoising time step sequence.
[0137] The previous denoising time step corresponding to the face enhancement time step mentioned above refers to the denoising time step that exists in the denoising time step sequence and is arranged adjacent to the arrangement position of the face enhancement time step and is arranged before the arrangement position of the face enhancement time step. For example, when the face enhancement time step is the time step S above m-1 , such as the time step 54 above, the previous denoising time step corresponding to the face enhancement time step can be the time step S abovem , such as time step 70 above.
[0138] Based on the above, we can see that the face enhancement time step and the previous denoising time step corresponding to it have the following two characteristics: ① The denoising time step sequence includes the face enhancement time step and the previous denoising time step corresponding to it. ② The position of the face enhancement time step in the denoising time step sequence is adjacent to the position of the previous denoising time step in the denoising time step sequence, and the position of the face enhancement time step in the denoising time step sequence is later than the position of the previous denoising time step in the denoising time step sequence.
[0139] In addition, for the previous denoising time step corresponding to the face enhancement time step above, such as time step S m For example, the denoising result corresponding to the previous denoising time step refers to the denoising correlation result obtained at the previous denoising time step, such as z shown in FIG2. M-1 Or by the z M-1 With facial representation data A M The fused data, etc.; and the denoising result corresponding to the previous denoising time step is obtained by processing the noise data to be processed corresponding to the previous denoising time step. The noise data to be processed corresponding to the previous denoising time step refers to the data that needs to be denoised at the previous denoising time step; and the noise data to be processed corresponding to the previous denoising time step is determined based on the noise processing result of the image to be processed above.
[0140] It should be noted that, for the previous denoising time step corresponding to the above face enhancement time step, the method for obtaining the to-be-processed noise data corresponding to the previous denoising time step is similar to the method for obtaining the to-be-processed noise data corresponding to the above face enhancement time step. For the sake of brevity, it will not be repeated here.
[0141] It should also be noted that for the previous denoising time step corresponding to the face enhancement time step above, such as time step S mFor example, if the previous denoising time step needs to be subjected to face enhancement processing, such as the previous denoising time step also belongs to the face enhancement time step, then the implementation method of the process of obtaining the denoising processing result corresponding to the previous denoising time step is similar to the implementation method of the process of obtaining the denoising processing result corresponding to the face enhancement time step provided in the present disclosure; however, if the previous denoising time step does not need to be subjected to face enhancement processing, such as the previous denoising time step belongs to the non-enhancement time step, then the process of obtaining the denoising processing result corresponding to the previous denoising time step can be: directly perform a denoising process on the noise data to be processed corresponding to the previous denoising time step, and obtain the denoising processing result corresponding to the previous denoising time step, as shown in FIG. M-1 The non-enhancement time step refers to the denoising time step that exists in the denoising time step sequence above and does not require face enhancement processing, so that the non-enhancement time step can represent the denoising time step that only requires one denoising process.
[0142] Based on the relevant content of step 212 above, it can be seen that for the above face enhancement time step, such as the above time step S M-1 For example, if the arrangement position of the face enhancement time step in the denoising time step sequence is not the first arrangement position in the denoising time step sequence, it can be determined that the face enhancement time step is not the first time step in the denoising time step sequence, so it can be determined that the process of determining the denoising processing result at at least one time step has been executed, and further it can be determined that the face enhancement time step can be used to process the denoising processing result at the previous denoising time step corresponding to the face enhancement time step. Therefore, the denoising processing result corresponding to the previous denoising time step can be directly determined as the noise data to be processed corresponding to the face enhancement time step.
[0143] Based on the relevant contents of steps 211 to 212 above, it can be seen that for the above denoised time step sequence, when the denoised time step sequence includes time step S M , time step S M-1 , ..., and time step S1, if the above face enhancement time step is the time step S M , then the noise data to be processed corresponding to the face enhancement time step may refer to the time step S M The corresponding noise processing result is shown in Figure 2. M ; If the face enhancement time step is the time step S M-1 , then the noise data to be processed corresponding to the face enhancement time step may refer to the time step S M The corresponding denoising result; if the face enhancement time step is the time step S M-2 , then the noise data to be processed corresponding to the face enhancement time step may refer to the time step S M-1 corresponding to the denoising result; ... (and so on); if the face enhancement time step is the time step S1, then the noise data to be processed corresponding to the face enhancement time step can refer to the denoising result corresponding to the time step S2. m The corresponding denoising result refers to the time step S m The denoising results determined under this condition; and if the time step S m Face enhancement processing is required, such as the time step S m belongs to the face enhancement time step, etc., then the time step S m The implementation of the process of determining the corresponding denoising result is similar to the implementation of the process of determining the denoising result corresponding to the face enhancement time step provided in the present disclosure; if the time step S m No face enhancement is required, such as the time step S m Belongs to non-enhanced time step, etc., then the time step S m The specific process of determining the corresponding denoising result can be as follows: directly performing the denoising operation on the time step S m The corresponding noise data to be processed is denoised once to obtain the time step S m The corresponding denoising result, m≤M, M is a positive integer.
[0144] Based on the above content regarding the noise data to be processed corresponding to the face enhancement time step, it can be seen that if the face enhancement time step is the first denoising time step in the denoising time step sequence, then the noise data to be processed corresponding to the face enhancement time step is the noise processing result corresponding to the first denoising time step; if the face enhancement time step is not the first denoising time step in the denoising time step sequence, then the noise data to be processed corresponding to the face enhancement time step is obtained by processing the noise processing result corresponding to the first denoising time step. It can be seen that in one possible implementation, when the noise processing result of the above image to be processed includes the noise processing result corresponding to the first denoising time step, the noise data to be processed corresponding to the face enhancement time step is determined based on the noise processing result corresponding to the first denoising time step. The first denoising time step refers to the time step that is at the first arrangement position in the denoising time step sequence.
[0145] In addition, for the face enhancement time step above, the denoised data corresponding to the face enhancement time step refers to the result obtained by performing a denoising process on the noise data to be processed corresponding to the face enhancement time step. For example, when the face enhancement time step is the time step S above M , and the noise data to be processed corresponding to the face enhancement time step is z as shown in Figure 2 M When , the denoised data corresponding to the face enhancement time step refers to the z shown in Figure 2 M-1It should be noted that the present disclosure does not limit the implementation method of the single denoising process. For example, it can be implemented using any existing or future method that can implement a single denoising process, such as the single denoising process implemented by a denoising module as shown in Figure 2. The denoising module is used to perform a single denoising process on the input data of the denoising module; and the present disclosure does not limit the implementation method of the denoising module.
[0146] Based on the relevant content of step 21 above, it can be known that for the above denoising time step sequence, if the denoising time step sequence includes a face enhancement time step, the noise data to be processed corresponding to the face enhancement time step can be denoised based on the above text features to obtain the denoised data corresponding to the face enhancement time step, so that the denoising result corresponding to the face enhancement time step can be obtained by subsequently performing face enhancement processing on the denoised data.
[0147] Step 22: Based on the facial representation data corresponding to the face enhancement time step above, face adjustment processing is performed on the denoised data corresponding to the face enhancement time step to obtain a denoising processing result corresponding to the face enhancement time step; the facial representation data is obtained by performing face extraction processing on the reference data corresponding to the face enhancement time step based on the facial position representation data of the image to be processed above; the reference data is determined based on the image to be processed or the denoising result of the image to be processed.
[0148] The facial representation data corresponding to the above face enhancement time step is used to describe the facial information required to be referenced when performing face enhancement processing on the denoised data corresponding to the face enhancement time step.
[0149] Furthermore, the facial representation data corresponding to the face enhancement time step described above is obtained by performing face extraction processing on the reference data corresponding to the face enhancement time step based on the facial position representation data of the image to be processed, so that the facial representation data can describe the facial information carried by the reference data. The reference data is used to provide facial information for the face enhancement processing at the face enhancement time step; and the reference data is determined based on the image to be processed or the result of a noise addition process on the image to be processed.
[0150] In addition, the present disclosure does not limit the method for obtaining the reference data corresponding to the above face enhancement time step. For example, in some application scenarios, the reference data corresponding to the face enhancement time step can be determined based on the above image to be processed, so that the reference data includes the image to be processed. In this way, the image to be processed can be used to provide facial information for the face enhancement processing at each face enhancement time step.
[0151] In addition, for the above-mentioned face enhancement time step, the denoised data corresponding to the face enhancement time step may still carry some noise information. Therefore, in order to better improve the face enhancement effect, the reference data corresponding to the face enhancement time step can be used to provide face information with noise for the face enhancement processing under the face enhancement time step. Based on this, the present disclosure also provides a possible implementation method of the process of obtaining the reference data corresponding to the above-mentioned face enhancement time step. Under this implementation method, when the noise processing result of the above-mentioned image to be processed includes the noise processing result corresponding to the face enhancement time step, the reference data corresponding to the face enhancement time step can be determined based on the noise processing result corresponding to the face enhancement time step, so that the reference data includes the noise processing result corresponding to the face enhancement time step. It can be seen that in one possible implementation method, when the face enhancement time step is the above-mentioned time step S m When the reference data corresponding to the face enhancement time step can be the time step S m The corresponding noise processing results are implemented. Among them, the time step S m The corresponding noise processing result is obtained by performing S on the image to be processed. m The noise is obtained by the second noise addition process, where m≤M, and M is a positive integer.
[0152] Based on the above content, we can know that for the above denoised time step sequence, when the denoised time step sequence includes time step S M , time step S M-1 , ..., and time step S1, if the above face enhancement time step is the time step S M , then the reference data corresponding to the face enhancement time step can be the time step S M The corresponding noise processing result; if the face enhancement time step is the time step S M-1 , then the reference data corresponding to the face enhancement time step can be the time step S M-1 The corresponding noise processing result; ... (and so on); if the face enhancement time step is the time step S1, then the reference data corresponding to the face enhancement time step can be the noise processing result corresponding to the time step S1.
[0153] Furthermore, in some application scenarios, to better improve the face enhancement effect, it is possible to ensure that the amount of noise carried by the reference data corresponding to the above face enhancement time step is as close as possible to the amount of noise carried by the denoised data corresponding to the above face enhancement time step. This helps avoid defects caused by excessive noise carried by the reference data corresponding to the face enhancement time step. Based on this, the present disclosure also provides a possible implementation method for obtaining the reference data corresponding to the above face enhancement time step. In this implementation method, when the above generated image is determined based on the denoising time step sequence, and the denoising time step sequence includes the face enhancement time step, the process of obtaining the reference data corresponding to the face enhancement time step may include the following steps 221-222.
[0154] Step 221: If the arrangement position of the above-mentioned face enhancement time step in the denoising time step sequence is not the last arrangement position in the denoising time step sequence, the denoising time step sequence also includes the next denoising time step corresponding to the face enhancement time step, and the reference data corresponding to the face enhancement time step is determined based on the denoising processing result corresponding to the next denoising time step. The denoising processing result of the above-mentioned image to be processed includes the denoising processing result corresponding to the next denoising time step. The arrangement position of the face enhancement time step in the denoising time step sequence is adjacent to the arrangement position of the next denoising time step in the denoising time step sequence, and the arrangement position of the face enhancement time step in the denoising time step sequence is earlier than the arrangement position of the next denoising time step in the denoising time step sequence.
[0155] The next denoising time step corresponding to the above face enhancement time step is the denoising time step that exists in the above denoising time step sequence and is arranged adjacent to the arrangement position of the face enhancement time step and is arranged later than the arrangement position of the face enhancement time step. For example, when the face enhancement time step is the above time step S m , such as the time step 70 above, the next denoising time step corresponding to the face enhancement time step can be the time step S above m-1 , such as time step 54 above.
[0156] Based on the above, we can see that the face enhancement time step and the next denoising time step corresponding to the face enhancement time step have the following two characteristics: ① The denoising time step sequence includes the face enhancement time step and the next denoising time step corresponding to the face enhancement time step. ② The position of the face enhancement time step in the denoising time step sequence is adjacent to the position of the next denoising time step in the denoising time step sequence, and the position of the face enhancement time step in the denoising time step sequence is earlier than the position of the next denoising time step in the denoising time step sequence.
[0157] In addition, for the next denoising time step corresponding to the above face enhancement time step, such as time step S m-1 For example, the denoising result corresponding to the next denoising time step refers to the result obtained by denoising the above image to be processed at the next denoising time step. m-1 When the denoising time step is completed, the denoising result corresponding to the next denoising time step can be obtained by performing S on the image to be processed. m-1 The noise is added.
[0158] Based on the relevant content of step 221 above, it can be known that for the above denoised time step sequence, when the denoised time step sequence includes time step S M , time step S M-1 , ... (and so on), and time step S1, if the above face enhancement time step is the time step S M , then the reference data corresponding to the face enhancement time step can be taken as the time step S M-1 The corresponding noise processing result is implemented; if the face enhancement time step is the time step S M-1 , then the reference data corresponding to the face enhancement time step can be taken as the time step S M-2The corresponding denoising processing result is implemented; ... (and so on); if the face enhancement time step is the time step S2, then the reference data corresponding to the face enhancement time step can be implemented using the denoising processing result corresponding to the time step S1. It can be seen that for the face enhancement time step, if the denoising time step sequence includes the face enhancement time step and the next denoising time step corresponding to the face enhancement time step, the denoising processing result corresponding to the next denoising time step can be determined as the reference data corresponding to the face enhancement time step, so that the amount of noise carried by the reference data is exactly consistent with the amount of noise that can be removed by all denoising processes after the face enhancement time step, so that the amount of noise carried by the reference data can be completely removed in all denoising processes after the face enhancement time step. This can effectively avoid the adverse effects caused by the excessive amount of noise carried by the reference data, thereby improving the image generation effect.
[0159] Step 222: If the arrangement position of the face enhancement time step in the denoising time step sequence is the last arrangement position in the denoising time step sequence, then the reference data corresponding to the face enhancement time step is determined based on the image to be processed.
[0160] In the present disclosure, for the above denoising time step sequence, if the denoising time step sequence includes a face enhancement time step, such as the above time step S1, and the arrangement position of the face enhancement time step in the denoising time step sequence is the last arrangement position in the denoising time step sequence, then it can be determined that the denoising process in the face enhancement time step is the last denoising process, and thus it can be determined that no denoising processing will be performed after the face enhancement time step. Therefore, in order to avoid noise interference, the above image to be processed can be directly determined as the reference data corresponding to the face enhancement time step, so that the reference data does not carry noise information, which can effectively avoid noise interference and thus help improve the image generation effect.
[0161] Based on the relevant contents of steps 221 to 222 above, it can be seen that for the above denoised time step sequence, when the denoised time step sequence includes time step S M , time step S M-1 , ... (and so on), and time step S1, if the above face enhancement time step is the time step S M , then the reference data corresponding to the face enhancement time step can be the time step S M-1 The corresponding noise processing results can be used to subsequently use the time step S M-1 The corresponding noise processing result for this time step S M The corresponding denoised data is processed for face enhancement to obtain the time step S MThe corresponding denoising result is so that the time step S M The facial noise carried by the corresponding denoising result can be completely removed after undergoing M-1 denoising processes; if the face enhancement time step is the time step S M-1 , then the reference data corresponding to the face enhancement time step can be the time step S M-2 The corresponding noise processing results can be used to subsequently use the time step S M-2 The corresponding noise processing result for this time step S M-1 The corresponding denoised data is processed for face enhancement to obtain the time step S M-1 The corresponding denoising result is so that the time step S M-1 The facial noise carried by the corresponding denoising result can be completely removed after undergoing M-2 denoising processes; ... (and so on); if the face enhancement time step is the time step S2, then the reference data corresponding to the face enhancement time step can be the denoising result corresponding to the time step S1, so that the face enhancement process can be performed on the denoised data corresponding to the time step S2 by using the denoising result corresponding to the time step S1, so as to obtain the denoising result corresponding to the time step S2, so that the facial noise carried by the denoising result corresponding to the time step S2 is completely removed. The noise can be completely removed after just one denoising process. If the face enhancement time step is the time step S1, the reference data corresponding to the face enhancement time step can be the above-mentioned image to be processed, so that the denoised data corresponding to the time step S1 can be subsequently subjected to face enhancement processing by using the image to be processed to obtain the denoising result corresponding to the time step S1, so that the denoising result corresponding to the time step S1 does not carry facial noise. In this way, facial noise interference can be effectively avoided, thereby improving the face generation effect, and further improving the image generation effect.
[0162] Based on the content of the above paragraph, it can be seen that for any face enhancement time step, the reference data corresponding to the face enhancement time step can be used to perform face enhancement processing on the denoised data corresponding to the face enhancement time step, and obtain the denoising processing result corresponding to the face enhancement time step, so that the facial information carried by the denoising processing result can better represent the facial constraints described by the image to be processed above. In addition, the present disclosure does not limit the implementation method of the face enhancement processing. For example, it can be specifically as follows: first, based on the facial position representation data of the image to be processed, face extraction processing is performed on the reference data corresponding to the face enhancement time step to obtain facial representation data corresponding to the face enhancement time step, so that the facial representation data can represent the facial information carried by the reference data; then, based on the facial representation data corresponding to the face enhancement time step, face adjustment processing is performed on the denoised data corresponding to the face enhancement time step to obtain the denoising processing result corresponding to the face enhancement time step.
[0163] It should be noted that the present disclosure does not limit the implementation method of the face extraction processing in the above paragraph. For example, when the facial position representation data of the image to be processed above is a face mask image of the image to be processed, the face extraction processing can specifically be: multiplying the face mask image with the reference data corresponding to the face enhancement time step to obtain the face representation data corresponding to the face enhancement time step.
[0164] It should also be noted that the present disclosure does not limit the implementation of the above face adjustment processing. For example, it may specifically include the following steps 223 to 226.
[0165] Step 223: Obtain the face enhancement parameters corresponding to the face enhancement time step and the non-face enhancement parameters corresponding to the face enhancement time step; the non-face enhancement parameters are determined based on the face enhancement parameters and the face position representation data of the image to be processed.
[0166] The face enhancement parameters corresponding to the above face enhancement time step are used to describe the degree of enhancement of the facial information required when performing face enhancement processing on the face enhancement time step.
[0167] In addition, the present disclosure does not limit the implementation method of the facial enhancement parameters corresponding to the above facial enhancement time steps. For example, in some application scenarios, the facial enhancement parameters corresponding to each facial enhancement time step may refer to the same pre-set fixed value to ensure that the degree of facial enhancement presented at each facial enhancement time step remains consistent.
[0168] In addition, in some application scenarios, for the denoised time step sequence above, the amount of noise carried by the reference data corresponding to the time step in the denoised time step sequence may be negatively correlated with the arrangement position of the time step in the denoised time step sequence, that is, the reference data corresponding to the time step with a relatively early arrangement position in the denoised time step sequence carries more noise, but the reference data corresponding to the time step with a relatively late arrangement position in the denoised time step sequence carries less noise. Based on this, it can be seen that in order to better reduce the interference caused by noise on the face enhancement effect, the present disclosure also provides a possible implementation method of the face enhancement parameters corresponding to the above-mentioned face enhancement time step. Under this implementation method, the face enhancement parameters corresponding to the face enhancement time step are positively correlated with the arrangement position of the face enhancement time step in the denoising time step sequence, so that the face enhancement time step with a relatively early arrangement position in the denoising time step sequence presents a lower degree of face enhancement, and the face enhancement time step with a relatively late arrangement position in the denoising time step sequence presents a higher degree of face enhancement. In this way, the interference caused by the large amount of noise carried by the reference data corresponding to some face enhancement time steps can be effectively avoided, thereby improving the image generation effect.
[0169] In fact, since the time steps in the above denoising time step sequence are arranged in descending order, in order to better reduce the interference of noise on the face enhancement effect, the present disclosure provides a possible implementation of the face enhancement parameters corresponding to the above face enhancement time step. Under this implementation, the face enhancement parameters corresponding to the face enhancement time step can be obtained by performing a power operation with the position of the face enhancement time step in the reverse result of the denoising time step sequence as the exponent, and the base of the power operation is in the open interval (0, 1), so that the face enhancement parameters corresponding to the face enhancement time step can be positively correlated with the position of the face enhancement time step in the denoising time step sequence. The reverse result of the denoising time step sequence is obtained by reversely arranging all the time steps in the denoising time step sequence. For example, if the denoising time step sequence is {S M , S M-1 , ..., S1}, then the reverse order of the denoised time step sequence can be {S1, ..., S M-1 , S M} this time step sequence.
[0170] Furthermore, the present disclosure does not limit the implementation of the base in the above paragraph. For example, in some application scenarios, the base can be 0.5, and when the denoising time step sequence above is {S M , S M-1 , ..., S1} this time step sequence, if the face enhancement time step is time step Sm , then the position of the face enhancement time step in the reverse order result of the denoising time step sequence is m, and the face enhancement parameter corresponding to the face enhancement time step can be implemented using the formula (1) shown below, so that the face enhancement parameter corresponding to the face enhancement time step is negatively correlated with the position of the face enhancement time step in the reverse order result of the denoising time step sequence, thereby making the face enhancement parameter corresponding to the face enhancement time step positively correlated with the position of the face enhancement time step in the denoising time step sequence, that is, if the face enhancement time step is positioned later in the denoising time step sequence, the face enhancement parameter corresponding to the face enhancement time step is larger, so that the facial information provided by the face enhancement time step has a greater impact on facial enhancement; if the face enhancement time step is positioned earlier in the denoising time step sequence, the face enhancement parameter corresponding to the face enhancement time step is smaller, so that the facial information provided by the face enhancement time step has a smaller impact on facial enhancement. α(S m )=c×0.5 (m-1) (1)
[0171] Where, α(S m ) represents the time step S m The corresponding face enhancement parameter; c represents a preset constant, and the present disclosure does not limit the implementation of c, for example, c = 1; 0.5 represents the base of the power operation; m represents the time step S m The position of the arrangement in the reverse order of the denoised time step sequence; (m-1) represents the exponent of the power operation.
[0172] In addition, for the face enhancement time step above, the non-face enhancement parameter corresponding to the face enhancement time step refers to the parameter required to be used when enhancing aspects other than the face, so that the non-face enhancement parameter is used to describe the degree of enhancement of the relevant information required when performing other enhancement processing on the face enhancement time step.
[0173] In addition, for the non-face enhancement parameters corresponding to the face enhancement time step above, the non-face enhancement parameters are determined based on the face enhancement parameters corresponding to the face enhancement time step and the face position representation data of the image to be processed above; and the present disclosure does not limit the determination process of the non-face enhancement parameters. For example, the non-face enhancement parameters can be implemented using the following formula (2). Time step S m Non-face enhancement parameter = 1-α(S m )×Mask (2)
[0174] Where, α(S m ) represents the time step S m The corresponding face enhancement parameters; Mask represents the facial position representation data of the above image to be processed, such as a face mask map.
[0175] Step 224: Determine first data according to the product of the face enhancement parameter corresponding to the face enhancement time step and the face representation data corresponding to the face enhancement time step.
[0176] The first data refers to the facial enhancement information involved in the facial enhancement time step described above; and the first data is determined based on the product between the facial enhancement parameter corresponding to the facial enhancement time step and the facial representation data corresponding to the facial enhancement time step, so that the first data includes the product.
[0177] Step 225: Determine second data based on the product of the non-face enhancement parameters corresponding to the face enhancement time step and the denoised data corresponding to the face enhancement time step.
[0178] The second data refers to other enhancement information other than the face involved in the above face enhancement time step; and the second data is determined based on the product between the non-face enhancement parameters corresponding to the above face enhancement time step and the denoised data corresponding to the face enhancement time step, so that the second data includes the product.
[0179] Step 226: Add the first data and the second data to obtain the denoising result corresponding to the above face enhancement time step.
[0180] Based on the relevant contents of steps 223 to 226 above, it can be seen that in some application scenarios, for the time step S in the denoising time step sequence above, m For example, if we need to m Perform face enhancement processing, then obtain the time step S m The denoised data and the time step S m After the reference data is obtained, the time step S m The denoised data and the time step S m The reference data is used, and the denoising result corresponding to the face enhancement time step is calculated using the following formula (3) and the above formula (1).
[0181] Time step S m The corresponding denoising result = z m-1 ×[1-α(S m )×Mask]+B m ×α(S m )×Mask(3)
[0182] Where z m-1 Indicates the time step S above m The denoised data; α(S m ) represents the time step S m Corresponding face enhancement parameters; B m Indicates the time step S m Reference data; Mask represents the facial position representation data of the above image to be processed, such as a facial mask map; A m =B m ×Mask indicates the time step S m Corresponding facial representation data; [1-α(S m )×Mask] represents the time step S m Corresponding non-face enhancement parameters.
[0183] Based on the relevant content of step 22 above, it can be known that for any face enhancement time step, after obtaining the denoised data corresponding to the face enhancement time step, face adjustment processing can be performed on the denoised data corresponding to the face enhancement time step based on the face representation data corresponding to the face enhancement time step to obtain the denoising processing result corresponding to the face enhancement time step, so that the denoising processing result carries the facial information described by the facial representation data, and the denoising processing result carries relevant information other than the face described by the denoised data, such as background information, so that the denoising processing result can better meet the facial constraints described by the image to be processed above, thereby facilitating improved image generation effects.
[0184] Step 23: Determine the generated image based on the denoising result corresponding to the above face enhancement time step.
[0185] It should be noted that the present disclosure does not limit the implementation of the above step 23. For ease of understanding, two examples are used for illustration below.
[0186] In Example 1, if the facial enhancement time step described above is located at the last position in the denoising time step sequence, then step 23 described above may specifically include decoding the denoising result corresponding to the facial enhancement time step to obtain a generated image. It should be noted that this disclosure does not limit the implementation of this decoding process. For example, it may be implemented using any existing or future method capable of decoding the results obtained from multiple rounds of denoising, such as the method implemented by the decoding processing module shown in FIG2 . The decoding processing module is used to decode the input data of the decoding processing module; and this disclosure does not limit the implementation of this decoding processing module.
[0187] Example 2: If the arrangement position of the above face enhancement time step in the denoising time step sequence is not the last arrangement position in the denoising time step sequence, then the denoising time step sequence also includes the next denoising time step corresponding to the face enhancement time step. Therefore, the above step 23 can specifically be: using the denoising processing result corresponding to the face enhancement time step as the noise data to be processed corresponding to the next denoising time step, and performing the process of determining the denoising processing result corresponding to the next denoising time step based on the noise data to be processed corresponding to the next denoising time step, so that the generated image can be determined based on the denoising processing result corresponding to the next denoising time step.
[0188] Based on the relevant content of steps 21 to 23 above, it can be seen that in some application scenarios, the generated image can be determined based on a sequence of denoising time steps; and the sequence of denoising time steps can include at least one face enhancement time step, so that the process of determining the generated image can include the process of determining the denoising result corresponding to at least one face enhancement time step. For any face enhancement time step, the implementation method of determining the denoising result corresponding to the face enhancement time step is similar to the implementation method of determining the denoising result corresponding to the face enhancement time step shown in steps 21 and 22 above. For the sake of brevity, it will not be repeated here.
[0189] In fact, in some application scenarios, in order to maximize the impact of the facial constraints described by the image to be processed above on the image generation process, the present disclosure also provides a possible implementation of the process for determining the generated image described above. In this implementation, the generated image can be determined based on a denoising time step sequence; and each time step in the denoising time step sequence is a face enhancement time step. Therefore, the process for determining the generated image can include determining the denoising processing results corresponding to multiple face enhancement time steps. The multiple face enhancement time steps include each time step in the denoising time step sequence.
[0190] To facilitate understanding of the above content, the following examples are provided for explanation.
[0191] As an example, when the above denoised time step sequence includes time step S M , time step S M-1 , ..., and at time step S1, the above determination process of generating an image may include the following steps 31 to 39.
[0192] Step 31: For the above time step S M The corresponding noise processing result is shown in Figure 2. M , perform denoising and get the time step S M The corresponding denoised data is shown in Figure 2.M-1 .
[0193] Step 32: Using the above time step S M The corresponding facial representation data, such as the facial representation data A shown in FIG2 M For this time step S M The corresponding denoised data is processed for face adjustment to obtain the time step S M The corresponding denoising result. Among them, the time step S M The corresponding facial representation data is based on the time step S M The corresponding reference data, such as the time step S M The corresponding noise processing result or the above time step S M-1 The corresponding noise processing results, etc. are determined.
[0194] Step 33: For the above time step S M The corresponding denoising results are denoised to obtain the above time step S M-1 The corresponding denoised data is shown in Figure 2. M-2 .
[0195] Step 34: Using the above time step S M-1 The corresponding facial representation data, such as the facial representation data A shown in FIG2 M-1 For this time step S M-1 The corresponding denoised data is processed for face adjustment to obtain the time step S M-1 The corresponding denoising result. Among them, the time step S M-1 The corresponding facial representation data is based on the time step S M-1 The corresponding reference data, such as the time step S M-1 The corresponding noise processing result or the above time step S M-2 The corresponding noise processing results, etc. are determined.
[0196] Step 35: For the above time step S M-1 The corresponding denoising results are denoised to obtain the above time step S M-2 The corresponding denoised data.
[0197] Step 36: Using the above time step S M-2 The corresponding facial representation data for this time step S M-2 The corresponding denoised data is processed for face adjustment to obtain the time step S M-2 The corresponding denoising result. Among them, the time step S M-2 The corresponding facial representation data is based on the time step S M-2 The corresponding reference data, such as the time step S M-2The corresponding noise processing result or the above time step S M-3 The corresponding noise processing results, etc. are determined.
[0198] ...(and so on)
[0199] Step 37: Perform denoising on the denoising result corresponding to the above time step S2 to obtain the denoised data corresponding to the above time step S1, such as z0 shown in FIG2 .
[0200] Step 38: Using the facial representation data corresponding to time step S1, perform facial adjustment processing on the denoised data corresponding to time step S1 to obtain a denoising result corresponding to time step S1. The facial representation data corresponding to time step S1 is determined based on reference data corresponding to time step S1, such as the denoising result corresponding to time step S1 or the image to be processed.
[0201] Step 39: Decode the denoising result corresponding to the above time step S1 to obtain the above generated image, such as image 2 shown in Figure 2.
[0202] Based on the relevant contents of steps 31 to 39 above, it can be seen that in one possible implementation, when the above generated image is determined based on a denoising time step sequence, and each time step in the denoising time step sequence is a face enhancement time step, the process of determining the generated image may include a denoising process and a face adjustment process corresponding to each time step, so that face enhancement processing can be performed at a corresponding intensity at each time step, so that the final generated image satisfies the face constraints described by the above image to be processed as much as possible, which is conducive to improving the image generation effect.
[0203] In fact, in some application scenarios, in order to better balance the influence of facial constraints and the influence of other constraints, the present disclosure also provides a possible implementation method of the above-mentioned process for determining the generated image. Under this implementation method, the process for determining the generated image can include a process for determining the denoising result corresponding to at least one face enhancement time step, and a process for determining the denoising result corresponding to at least one non-enhancement time step. For any non-enhancement time step, the process for determining the denoising result corresponding to the non-enhancement time step can specifically be: denoising the noise data to be processed corresponding to the non-enhancement time step based on the above-mentioned text features to obtain the denoising result corresponding to the non-enhancement time step. The noise data to be processed corresponding to the non-enhancement time step refers to the data that needs to be denoised at the non-enhancement time step; and the implementation method of the noise data to be processed corresponding to the non-enhancement time step is similar to the implementation method of the noise data to be processed corresponding to the face enhancement time step above. For the sake of brevity, it will not be repeated here.
[0204] To facilitate understanding of the above content, the following examples are provided for explanation.
[0205] As an example, when the above denoised time step sequence includes time step S M , time step S M-1 , ..., and time step S1, if the time step S M Belongs to the face enhancement time step, the time step S M-1 Belongs to the non-enhanced time step, the time step S M- 2 belongs to the face enhancement time step, the time step S M-3 belongs to a non-enhancement time step, ... (and so on), and the time step S1 belongs to a face enhancement time step, then the above determination process of generating an image may include the following steps 41 to 49.
[0206] Step 41: For the above time step S M The corresponding noise processing result is denoised to obtain the time step S M The corresponding denoised data.
[0207] Step 42: Using the above time step S M The corresponding facial representation data for this time step S M The corresponding denoised data is processed for face adjustment to obtain the time step S M The corresponding denoising result. Among them, the time step S M The corresponding facial representation data is based on the time step S M The corresponding reference data, such as the time step S M The corresponding noise processing result or the above time step S M-1The corresponding noise processing results, etc. are determined.
[0208] Step 43: For the above time step S M The corresponding denoising results are denoised to obtain the above time step S M-1 The corresponding denoising results.
[0209] Step 44: For the above time step S M-1 The corresponding denoising results are denoised to obtain the above time step S M-2 The corresponding denoised data.
[0210] Step 45: Using the above time step S M-2 The corresponding facial representation data for this time step S M-2 The corresponding denoised data is processed for face adjustment to obtain the time step S M-2 The corresponding denoising result. Among them, the time step S M-2 The corresponding facial representation data is based on the time step S M-2 The corresponding reference data, such as the time step S M-2 The corresponding noise processing result or the above time step S M-3 The corresponding noise processing results, etc. are determined.
[0211] Step 46: For the above time step S M-2 The corresponding denoising results are denoised to obtain the above time step S M-3 The corresponding denoising results.
[0212] ...(and so on)
[0213] Step 47: Perform denoising on the denoising result corresponding to the above time step S2 to obtain the denoised data corresponding to the above time step S1.
[0214] Step 48: Using the facial representation data corresponding to time step S1, perform facial adjustment processing on the denoised data corresponding to time step S1 to obtain a denoising result corresponding to time step S1. The facial representation data corresponding to time step S1 is determined based on reference data corresponding to time step S1, such as the denoising result corresponding to time step S1 or the image to be processed.
[0215] Step 49: Decode the denoising result corresponding to the above time step S1 to obtain the above generated image.
[0216] Based on the relevant contents of steps 41 to 49 above, it can be known that in one possible implementation manner, when the generated image above is determined based on a denoising time step sequence, some time steps in the denoising time step sequence are face enhancement time steps, and another part of the time steps in the denoising time step sequence are non-enhancement time steps, the process of determining the generated image can be implemented by skipping steps to perform face enhancement processing, which is beneficial for balancing the influence of facial constraints and the influence of other constraints, thereby facilitating the generation of an image that takes into account as much as possible the facial constraints and other constraints.
[0217] Based on the content of the above paragraph, it can be seen that in one possible implementation, when the above-mentioned generated image is determined based on a denoising time step sequence, and the determination process of the generated image may include a determination process of a denoising processing result corresponding to at least one face enhancement time step, and a determination process of a denoising processing result corresponding to at least one non-enhancement time step, the at least one face enhancement time step may include some time steps in the denoising time step sequence, and the at least one non-enhancement time step may include other remaining time steps in the denoising time step sequence except the some time steps.
[0218] In practice, when the process for determining the generated image described above is implemented by skipping face enhancement steps, the number of skipped steps varies in different scenarios, resulting in a different number of time steps between two adjacent face enhancement processes. Based on this, it can be seen that in one possible implementation, when the generated image described above is determined based on a denoising time step sequence, and the process for determining the generated image can include determining the denoising result corresponding to at least one face enhancement time step and determining the denoising result corresponding to at least one non-enhancement time step, the denoising time step sequence includes the at least one face enhancement time step and the at least one non-enhancement time step, and in the denoising time step sequence, one or more non-enhancement time steps exist between any two face enhancement time steps. To facilitate understanding, the following explanation is provided with two examples.
[0219] Example 1: In some application scenarios, when the denoising time step sequence above includes time step S M , time step S M-1 , ..., and time step S1, if the above determination process of generating an image is implemented by skipping one step to perform face enhancement processing, then the time step in the denoising time step sequence can have the following characteristics: the time step S M Belongs to the face enhancement time step, the time step S M-1 Belongs to non-enhanced time step, time step S M-2 Belongs to the face enhancement time step, time step S M-3belongs to a non-enhancement time step, ... (and so on), so that there is a non-enhancement time step between any two closest face enhancement time steps in the denoising time step sequence, and thus there is at least one non-enhancement time step between any two face enhancement time steps in the denoising time step sequence.
[0220] Example 2: In some application scenarios, when the denoising time step sequence above includes time step S M , time step S M-1 , ..., and time step S1, if the above determination process of generating an image is implemented by skipping two steps to perform face enhancement processing, then the time steps in the denoising time step sequence can have the following characteristics: the time step S M Belongs to the face enhancement time step, the time step S M-1 Belongs to non-enhanced time step, time step S M-2 Belongs to non-enhanced time step, time step S M-3 Belongs to the face enhancement time step, the time step S M-4 Belongs to non-enhanced time step, time step S M-5 belongs to a non-enhancement time step, ... (and so on), so that there are two non-enhancement time steps between any two closest face enhancement time steps in the denoising time step sequence, and thus there are at least two non-enhancement time steps between any two face enhancement time steps in the denoising time step sequence.
[0221] In fact, when the above-mentioned determination process of generating an image includes the process of determining the denoising processing result corresponding to at least one face enhancement time step and the process of determining the denoising processing result corresponding to at least one non-enhancement time step, the proportion of the face enhancement time step in the denoising time step sequence is positively correlated with the degree of influence of the face constraint on the image generation process. Therefore, in order to better meet the image generation requirements, the at least one face enhancement time step and the at least one non-enhancement time step can be determined based on the image generation requirements provided by the user.
[0222] Based on the content of the above paragraph, it can be seen that in one possible implementation manner, when the above determination process of the generated image includes the process of determining the denoising processing result corresponding to at least one face enhancement time step, and the process of determining the denoising processing result corresponding to at least one non-enhancement time step, if the generated image is determined based on the denoising time step sequence, then the at least one face enhancement time step and the at least one non-enhancement time step can be determined from the denoising time step sequence, and the determination process can specifically include the following steps 51 to 53.
[0223] Step 51: Obtain image generation requirement description information corresponding to the image to be processed; the image generation requirement description information includes face preservation degree representation data and / or style degree representation data.
[0224] The image generation requirement description information corresponding to the image to be processed is used to describe the influence requirements presented in terms of facial constraints and the influence requirements presented in terms of other constraints when image generation processing is performed based on the image to be processed.
[0225] In addition, the present disclosure does not limit the implementation of the image generation requirement description information corresponding to the image to be processed. For example, in some application scenarios, the image generation requirement description information may include face preservation degree representation data and / or style degree representation data. The face preservation degree representation data is used to represent the degree of influence required in terms of facial constraints during image generation processing. The style degree representation data is used to represent the degree of influence required in terms of style constraints during image generation processing.
[0226] Step 52: Determine facial enhancement degree representation data based on the above image generation requirement description information.
[0227] The facial enhancement degree representation data refers to the degree of facial enhancement required for image generation processing; and the present disclosure does not limit the implementation method of the facial enhancement degree representation data. For example, the facial enhancement degree representation data can be implemented using the number of facial enhancement times. The number of facial enhancement times refers to the number of times the facial enhancement process needs to be performed during image generation processing. For another example, the facial enhancement degree representation data can be implemented using the facial enhancement ratio. The facial enhancement ratio refers to the proportion of facial enhancement processes during image generation processing, such as the proportion of facial enhancement time steps in the denoising time step sequence mentioned above.
[0228] Step 53: Based on the facial enhancement degree characterization data, determine at least one face enhancement time step and at least one non-enhancement time step from the denoised time step sequence.
[0229] In the present disclosure, after obtaining the above-mentioned facial enhancement degree representation data, all time steps in the denoised time step sequence can be classified according to the facial enhancement degree representation data to obtain the above-mentioned at least one facial enhancement time step and the above-mentioned at least one non-enhancement time step, so that the facial enhancement degree described by the facial enhancement degree representation data can be satisfied subsequently by executing the facial enhancement process at the at least one facial enhancement time step, thereby facilitating meeting the user's image generation needs.
[0230] Based on the relevant contents of steps 51 to 53 above, it can be known that in some application scenarios, at least one face enhancement time step and at least one non-enhancement time step can be determined from the above denoising time step sequence based on the image generation requirement description information provided by the user, so that the degree of influence of the face constraints and the degree of influence of the style constraints described by the image generation requirement description information can be met by subsequently executing the determination process of the denoising processing results corresponding to these time steps, which is conducive to meeting the user's image generation requirements.
[0231] Based on the relevant contents of S1 to S2 above, it can be seen that for the image generation method provided in the embodiment of the present disclosure, the image to be processed and the prompt text corresponding to the image to be processed are first obtained; then, image generation processing is performed based on the noise processing result of the image to be processed, the facial position representation data of the image to be processed, and the text features of the prompt text to obtain a generated image, so that the generated image not only satisfies the facial constraints described by the image to be processed, but also satisfies certain constraints described by the prompt text, such as style constraints. In particular, because the generated image is generated based on the facial position representation data, the facial generation processing can be better performed based on the facial position representation data during the generation process of the generated image, so that the face described by the generated image is as close as possible to the face described by the image to be processed, thereby improving the image generation effect. Furthermore, because the noise processing result of the image to be processed includes noise processing results corresponding to multiple time steps, the generated image obtained based on the noise processing result of the image to be processed can better satisfy the facial constraints described by the image to be processed, thus improving the image generation effect.
[0232] In addition, the present disclosure does not limit the execution subject of the image generation method provided in the embodiments of the present disclosure. For example, the image generation method provided in the embodiments of the present disclosure can be applied to a terminal device or a server. For another example, the image generation method provided in the embodiments of the present disclosure can also be implemented with the help of a data interaction process between a terminal device and a server. The terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a standalone server, a cluster server, or a cloud server.
[0233] Based on the image generation method provided in the embodiments of the present disclosure, the embodiments of the present disclosure also provide an image generation device, which will be explained and illustrated below in conjunction with Figure 3. Figure 3 is a schematic diagram of the structure of the image generation device provided in the embodiments of the present disclosure. It should be noted that for technical details of the image generation device provided in the embodiments of the present disclosure, please refer to the relevant content of the image generation method above.
[0234] As shown in FIG3 , an image generating apparatus 300 provided in an embodiment of the present disclosure includes:
[0235] An acquiring unit 301 is configured to acquire an image to be processed and a prompt text corresponding to the image to be processed;
[0236] The generation unit 302 is used to perform image generation processing based on the noise processing result of the image to be processed, the facial position representation data of the image to be processed, and the text features of the prompt text to obtain a generated image, wherein the noise processing result includes the noise processing results corresponding to multiple time steps.
[0237] In a possible implementation, the generating unit 302 includes:
[0238] a denoising subunit, configured to perform denoising on the noise data to be processed corresponding to the face enhancement time step based on the text features, to obtain denoised data corresponding to the face enhancement time step; the noise data to be processed is determined based on a noise addition result of the image to be processed;
[0239] an enhancement subunit, configured to perform face adjustment processing on the denoised data corresponding to the face enhancement time step based on the face representation data corresponding to the face enhancement time step, to obtain a denoising result corresponding to the face enhancement time step; the face representation data is obtained by performing face extraction processing on reference data corresponding to the face enhancement time step based on the face position representation data; the reference data is determined based on the image to be processed or the denoising result of the image to be processed;
[0240] The determination subunit is configured to determine the generated image based on a denoising result corresponding to the face enhancement time step.
[0241] In one possible implementation, the noise processing result of the image to be processed includes the noise processing result corresponding to the face enhancement time step; and the reference data corresponding to the face enhancement time step is determined based on the noise processing result corresponding to the face enhancement time step.
[0242] In one possible implementation, the generated image is determined based on a denoising time step sequence; the denoising time step sequence includes the face enhancement time step;
[0243] If the arrangement position of the face enhancement time step in the denoising time step sequence is not the last arrangement position in the denoising time step sequence, the denoising time step sequence also includes a next denoising time step corresponding to the face enhancement time step, and the reference data corresponding to the face enhancement time step is determined based on the denoising processing result corresponding to the next denoising time step; the denoising processing result of the image to be processed includes the denoising processing result corresponding to the next denoising time step; the arrangement position of the face enhancement time step in the denoising time step sequence is adjacent to the arrangement position of the next denoising time step in the denoising time step sequence, and the arrangement position of the face enhancement time step in the denoising time step sequence is earlier than the arrangement position of the next denoising time step in the denoising time step sequence;
[0244] If the arrangement position of the face enhancement time step in the denoising time step sequence is the last arrangement position in the denoising time step sequence, the reference data corresponding to the face enhancement time step is determined based on the image to be processed.
[0245] In one possible implementation, the enhancement subunit is specifically configured to: obtain face enhancement parameters corresponding to the face enhancement time step and non-face enhancement parameters corresponding to the face enhancement time step; the non-face enhancement parameters are determined based on the face enhancement parameters and the face position representation data; determine first data based on the product of the face enhancement parameters and the face representation data corresponding to the face enhancement time step; determine second data based on the product of the non-face enhancement parameters and the denoised data corresponding to the face enhancement time step; and add the first data and the second data to obtain a denoising result corresponding to the face enhancement time step.
[0246] In one possible implementation, the generated image is determined based on a denoising time step sequence; the denoising time step sequence includes the face enhancement time step;
[0247] If the face enhancement time step is located at the first position in the denoising time step sequence, the noise data to be processed corresponding to the face enhancement time step is determined based on the noise processing result corresponding to the face enhancement time step; and the noise processing result of the image to be processed includes the noise processing result corresponding to the face enhancement time step.
[0248] If the arrangement position of the face enhancement time step in the denoising time step sequence is not the first arrangement position in the denoising time step sequence, the denoising time step sequence also includes the previous denoising time step corresponding to the face enhancement time step, and the noise data to be processed corresponding to the face enhancement time step is determined based on the denoising processing result corresponding to the previous denoising time step; the denoising processing result corresponding to the previous denoising time step is obtained by processing the noise data to be processed corresponding to the previous denoising time step, and the noise data to be processed corresponding to the previous denoising time step is determined based on the denoising processing result of the image to be processed; the arrangement position of the face enhancement time step in the denoising time step sequence is adjacent to the arrangement position of the previous denoising time step in the denoising time step sequence, and the arrangement position of the face enhancement time step in the denoising time step sequence is later than the arrangement position of the previous denoising time step in the denoising time step sequence.
[0249] In one possible implementation, the process of determining the generated image includes a process of determining a denoising result corresponding to at least one face enhancement time step;
[0250] For any of the face enhancement time steps, the process of determining the denoising result corresponding to the face enhancement time step includes: performing denoising on the noise data to be processed corresponding to the face enhancement time step based on the text features to obtain denoised data corresponding to the face enhancement time step; the noise data to be processed is determined based on the noise addition result of the image to be processed; performing face adjustment on the denoised data corresponding to the face enhancement time step based on the face representation data corresponding to the face enhancement time step to obtain the denoising result corresponding to the face enhancement time step; the face representation data is obtained by performing face extraction on the reference data corresponding to the face enhancement time step based on the face position representation data; and the reference data is determined based on the image to be processed or the noise addition result of the image to be processed.
[0251] In one possible implementation, the generated image is determined based on a denoising time step sequence; and the at least one face enhancement time step includes each time step in the denoising time step sequence.
[0252] In one possible implementation, the process of determining the generated image further includes a process of determining a denoising result corresponding to at least one non-enhancement time step;
[0253] For any of the non-enhanced time steps, the process of determining the denoising result corresponding to the non-enhanced time step includes: performing denoising on the noise data to be processed corresponding to the non-enhanced time step based on the text features to obtain the denoising result corresponding to the non-enhanced time step.
[0254] In one possible implementation, the generated image is determined based on a denoising time step sequence; the at least one face enhancement time step includes a portion of the time steps in the denoising time step sequence; and the at least one non-enhancement time step includes the remaining time steps in the denoising time step sequence except the portion of the time steps.
[0255] In one possible implementation, the generated image is determined based on a denoising time step sequence; the denoising time step sequence includes the at least one face enhancement time step and the at least one non-enhancement time step; and in the denoising time step sequence, one or more non-enhancement time steps exist between any two face enhancement time steps.
[0256] In one possible implementation, the generated image is determined based on a denoising time step sequence;
[0257] The process of determining the at least one face enhancement time step and the at least one non-enhancement time step includes: obtaining image generation requirement description information corresponding to the image to be processed; the image generation requirement description information includes face preservation degree representation data and / or style degree representation data; determining face enhancement degree representation data based on the image generation requirement description information; and determining the at least one face enhancement time step and the at least one non-enhancement time step from the denoising time step sequence based on the face enhancement degree representation data.
[0258] In one possible implementation, the prompt text includes positive prompt content and negative prompt content;
[0259] The process of determining the text features of the prompt text includes: obtaining the feature vector of the positive prompt content and the feature vector of the negative prompt content; comparing the size of the feature vector of the positive prompt content with the size of the feature vector of the negative prompt content to obtain a comparison result; and determining the text features of the prompt text based on the comparison result, the feature vector of the positive prompt content, and the feature vector of the negative prompt content.
[0260] In one possible implementation, the process of determining the text features of the prompt text includes: when the comparison result indicates that the size of the first feature vector is smaller than the size of the second feature vector, using part of the data in the first feature vector to perform data expansion processing on the first feature vector to obtain an expanded feature vector, and the size of the expanded feature vector is equal to the size of the second feature vector; the first feature vector is the feature vector of the positive prompt content, and the second feature vector is the feature vector of the negative prompt content; or, the first feature vector is the feature vector of the negative prompt content, and the second feature vector is the feature vector of the positive prompt content; the expanded feature vector and the second feature vector are spliced together to obtain the text features of the prompt text.
[0261] In one possible implementation, the first feature vector includes semantic representation data of at least one candidate semantic unit;
[0262] The process of determining the partial data includes: determining a target semantic unit from the at least one candidate semantic unit based on the importance representation data of each candidate semantic unit, the importance representation data of the target semantic unit being higher than the importance representation data of other candidate semantic units in the at least one candidate semantic unit except the target semantic unit; and determining the partial data based on the semantic representation data of the target semantic unit.
[0263] Based on the above-described related content of the image generation device 300, it can be seen that the image generation device 300 provided in the embodiment of the present disclosure first obtains a to-be-processed image and a corresponding prompt text for the to-be-processed image; then, image generation processing is performed based on the noise processing result of the to-be-processed image, the facial position representation data of the to-be-processed image, and the text features of the prompt text to obtain a generated image, so that the generated image not only satisfies the facial constraints described by the to-be-processed image, but also satisfies certain constraints described by the prompt text, such as style constraints. Because the generated image is generated based on the facial position representation data, the facial generation processing can be better performed based on the facial position representation data during the generation process of the generated image, so that the face described by the generated image is as close as possible to the face described by the to-be-processed image, thereby improving the image generation effect. Furthermore, because the noise processing result of the to-be-processed image includes noise processing results corresponding to multiple time steps, the generated image obtained based on the noise processing result of the to-be-processed image can better satisfy the facial constraints described by the to-be-processed image, thus improving the image generation effect.
[0264] In addition, an embodiment of the present disclosure also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the image generation method provided by the embodiment of the present disclosure.
[0265] Referring to FIG4 , a schematic diagram of the structure of an electronic device 400 suitable for implementing embodiments of the present disclosure is shown. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG4 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.
[0266] As shown in Figure 4, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of electronic device 400 are also stored in RAM 403. Processing device 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to bus 404.
[0267] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although FIG4 shows the electronic device 400 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0268] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0269] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0270] The embodiments of the present disclosure further provide a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the image generation method provided by the embodiments of the present disclosure.
[0271] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0272] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0273] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0274] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can perform the method.
[0275] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0276] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0277] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.
[0278] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0279] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0280] It should be noted that the various embodiments of this disclosure are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the descriptions of the systems or devices disclosed in the embodiments for similarities and differences between them. Since the systems or devices disclosed in the embodiments correspond to the methods disclosed in the embodiments, their descriptions are relatively simple, and reference can be made to the descriptions of the methods for any related details.
[0281] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0282] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0283] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0284] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present disclosure. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments shown herein, but is intended to be construed in the widest manner consistent with the principles and novel features disclosed herein.
Claims
1. An image generation method, comprising: Obtaining an image to be processed and a prompt text corresponding to the image to be processed; Performing image generation processing based on the noise addition processing result of the image to be processed, the facial position characterization data of the image to be processed, and the text features of the prompt text to obtain a generated image, where the noise addition processing result includes noise addition processing results corresponding to multiple time steps.
2. The method according to claim 1, wherein, The determination process of the generated image includes: Denosing the noise data to be processed corresponding to the facial enhancement time step according to the text features to obtain the denoised data corresponding to the facial enhancement time step; the noise data to be processed is determined according to the noise addition processing result of the image to be processed; Performing facial adjustment processing on the denoised data corresponding to the facial enhancement time step according to the facial characterization data corresponding to the facial enhancement time step to obtain the noise addition processing result corresponding to the facial enhancement time step; the facial characterization data is obtained by performing facial extraction processing on the reference data corresponding to the facial enhancement time step according to the facial position characterization data; the reference data is determined according to the image to be processed or the noise addition processing result of the image to be processed; Determining the generated image according to the noise addition processing result corresponding to the facial enhancement time step.
3. The method according to claim 2, wherein, The noise addition processing result of the image to be processed includes the noise addition processing result corresponding to the facial enhancement time step; The reference data corresponding to the facial enhancement time step is determined according to the noise addition processing result corresponding to the facial enhancement time step.
4. The method according to claim 2, wherein The generated image is determined according to a denoising time step sequence; the denoising time step sequence includes the facial enhancement time step; If the arrangement position of the facial enhancement time step in the denoising time step sequence is not the last arrangement position in the denoising time step sequence, the denoising time step sequence further includes the next denoising time step corresponding to the facial enhancement time step, and the reference data corresponding to the facial enhancement time step is determined according to the noise addition processing result corresponding to the next denoising time step; the noise addition processing result of the image to be processed includes the noise addition processing result corresponding to the next denoising time step; the arrangement position of the facial enhancement time step in the denoising time step sequence is adjacent to the arrangement position of the next denoising time step in the denoising time step sequence, and the arrangement position of the facial enhancement time step in the denoising time step sequence is before the arrangement position of the next denoising time step in the denoising time step sequence; If the arrangement position of the facial enhancement time step in the denoising time step sequence is the last arrangement position in the denoising time step sequence, the reference data corresponding to the facial enhancement time step is determined according to the image to be processed.
5. The method according to claim 2, wherein, The determination process of the noise addition processing result corresponding to the facial enhancement time step includes: Obtaining the facial enhancement parameter corresponding to the facial enhancement time step and the non-facial enhancement parameter corresponding to the facial enhancement time step; the non-facial enhancement parameter is determined according to the facial enhancement parameter and the facial position characterization data; Determine a first data based on the product between the facial enhancement parameter and the facial representation data corresponding to the facial enhancement time step; Determine a second data based on the product between the non-facial enhancement parameter and the denoised data corresponding to the facial enhancement time step; Add the first data and the second data to obtain the denoising processing result corresponding to the facial enhancement time step.
6. The method according to claim 2, wherein, The generated image is determined according to a denoising time step sequence; the denoising time step sequence includes the facial enhancement time step; If the arrangement position of the facial enhancement time step in the denoising time step sequence is the first arrangement position in the denoising time step sequence, the noise data to be processed corresponding to the facial enhancement time step is determined according to the noise addition processing result corresponding to the facial enhancement time step; the noise addition processing result of the image to be processed includes the noise addition processing result corresponding to the facial enhancement time step; If the arrangement position of the facial enhancement time step in the denoising time step sequence is not the first arrangement position in the denoising time step sequence, the denoising time step sequence further includes the previous denoising time step corresponding to the facial enhancement time step, and the noise data to be processed corresponding to the facial enhancement time step is determined according to the denoising processing result corresponding to the previous denoising time step; the denoising processing result corresponding to the previous denoising time step is obtained by processing the noise data to be processed corresponding to the previous denoising time step, and the noise data to be processed corresponding to the previous denoising time step is determined according to the noise addition processing result of the image to be processed; the arrangement position of the facial enhancement time step in the denoising time step sequence is adjacent to the arrangement position of the previous denoising time step in the denoising time step sequence, and the arrangement position of the facial enhancement time step in the denoising time step sequence is behind the arrangement position of the previous denoising time step in the denoising time step sequence.
7. The method according to claim 1, wherein The determination process of the generated image includes the determination process of the denoising processing result corresponding to at least one facial enhancement time step; For any one of the facial enhancement time steps, the determination process of the denoising processing result corresponding to this facial enhancement time step includes: Perform denoising processing on the noise data to be processed corresponding to this facial enhancement time step according to the text feature to obtain the denoised data corresponding to this facial enhancement time step; the noise data to be processed is determined according to the noise addition processing result of the image to be processed; Perform facial adjustment processing on the denoised data corresponding to this facial enhancement time step according to the facial representation data corresponding to this facial enhancement time step to obtain the denoising processing result corresponding to this facial enhancement time step; the facial representation data is obtained by performing facial extraction processing on the reference data corresponding to this facial enhancement time step according to the facial position representation data; the reference data is determined according to the image to be processed or the noise addition processing result of the image to be processed.
8. The method according to claim 7, wherein The generated image is determined according to a denoising time step sequence; The at least one face enhancement time step includes each time step in the denoising time step sequence.
9. The method according to claim 7, wherein, The determination process of the generated image further includes a determination process of the denoising processing results corresponding to at least one non-enhancement time step; For any one of the non-enhancement time steps, the determination process of the denoising processing result corresponding to this non-enhancement time step includes: Performing denoising processing on the to-be-processed noise data corresponding to this non-enhancement time step according to the text feature to obtain the denoising processing result corresponding to this non-enhancement time step.
10. The method according to claim 9, wherein, The generated image is determined according to the denoising time step sequence; The at least one face enhancement time step includes some time steps in the denoising time step sequence; The at least one non-enhancement time step includes the other remaining time steps in the denoising time step sequence except for the some time steps.
11. The method according to claim 9, wherein, The generated image is determined according to the denoising time step sequence; The denoising time step sequence includes the at least one face enhancement time step and the at least one non-enhancement time step; In the denoising time step sequence, there is one or more non-enhancement time steps between any two face enhancement time steps.
12. The method according to claim 9, wherein, The generated image is determined according to the denoising time step sequence; The determination process of the at least one face enhancement time step and the at least one non-enhancement time step includes: Obtaining the image generation requirement description information corresponding to the to-be-processed image; the image generation requirement description information includes face retention degree characterization data and / or style degree characterization data; Determining the face enhancement degree characterization data according to the image generation requirement description information; Determining the at least one face enhancement time step and the at least one non-enhancement time step from the denoising time step sequence according to the face enhancement degree characterization data.
13. The method according to any one of claims 1-12, wherein, The prompt text includes positive prompt content and negative prompt content; The determination process of the text feature of the prompt text includes: Obtaining the feature vector of the positive prompt content and the feature vector of the negative prompt content; Comparing the size of the feature vector of the positive prompt content with the size of the feature vector of the negative prompt content to obtain a comparison result; Determining the text feature of the prompt text according to the comparison result, the feature vector of the positive prompt content, and the feature vector of the negative prompt content.
14. The method according to claim 13, wherein, The determining the text feature of the prompt text according to the comparison result, the feature vector of the positive prompt content, and the feature vector of the negative prompt content includes: When the comparison result indicates that the size of the first feature vector is smaller than the size of the second feature vector, using some data in the first feature vector to perform data augmentation processing on the first feature vector to obtain an augmented feature vector, and the size of the augmented feature vector is equal to the size of the second feature vector; the first feature vector is the feature vector of the positive prompt content, and the second feature vector is the feature vector of the negative prompt content; or, the first feature vector is the feature vector of the negative prompt content, and the second feature vector is the feature vector of the positive prompt content; Concatenate the augmented feature vector with the second feature vector to obtain the text feature of the prompt text.
15. The method according to claim 14, wherein The first feature vector includes semantic representation data of at least one candidate semantic unit; The process of determining the partial data includes: Determine a target semantic unit from the at least one candidate semantic unit according to the importance representation data of each candidate semantic unit, and the importance representation data of the target semantic unit is higher than that of other candidate semantic units except the target semantic unit among the at least one candidate semantic unit; Determine the partial data according to the semantic representation data of the target semantic unit.
16. An image generation device, comprising: An acquisition unit configured to acquire a to-be-processed image and a prompt text corresponding to the to-be-processed image; A generation unit configured to perform image generation processing based on the noise addition processing result of the to-be-processed image, the facial position representation data of the to-be-processed image, and the text feature of the prompt text to obtain a generated image, and the noise addition processing result includes noise addition processing results corresponding to multiple time steps.
17. An electronic device, comprising a processor and a memory, wherein The memory is configured to store instructions or computer programs; The processor is configured to execute the instructions or the computer programs in the memory so that the electronic device executes the method according to any one of claims 1-15.
18. A computer-readable medium stores instructions or a computer program, wherein, When the instructions or the computer programs are running on the device, the device executes the method according to any one of claims 1-15.
Citation Information
Patent Citations
Decoupling efficient fine tuning method and product for main body oriented character generation image
CN116522870A
Image synthesis method and device, electronic equipment and storage medium
CN117036184A
Image generation model training method and device, equipment and storage medium
CN117218217A
Image and object inpainting with diffusion models
GB202314582D0
Denoising diffusion generative adversarial networks
US20230095092A1